Scale AI logo

Machine Learning Research Engineer (ML Systems)

Scale AI

On-siteSan Francisco, CAmid$201k–$251kPosted 8h ago

Job description

  • You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development

  • You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation

  • Build, profile and optimize our training and inference framework

  • Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation

  • Research and integrate state-of-the-art technologies to optimize our ML system

Benefits

  • Health & Wellbeing: Our holistic approach to supporting Scaliens includes comprehensive health coverage, dental and vision insurance, mental healthcare services, and more. PTO policies and accommodating schedules ensure you’ll get time off when you need it to relax and recharge. Note that our offerings may vary by region as we strive to respond to the unique needs of Scaliens around the globe.

  • Personal & Career Growth: Continuously learn and grow through annual learning & development stipend, attending leadership breakfasts, manager training, speaker series, and joining an ERG.

  • Building Scale Community: We welcome guests to our offices, and you can expect to see Scalien families and friends around. Join local happy hours, and accept invites to game nights, book clubs, and many other employee-led community events.

  • Parental Support: Balancing work and family is essential, and Scale understands the importance of having adequate leave policies in place to promote a healthy home and work life.- If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you!

  • Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc

  • Strong written and verbal communication skills and the ability to operate in a cross functional team environment

  • Experience with multi-node LLM training and inference

  • Strong excitement about system optimization

  • Experience with developing large-scale distributed ML systems

  • Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc