Crusoe Energy Systems logo

Senior Staff Software Engineer (AI Model LifeCycle)

Crusoe Energy Systems

On-siteSan Francisco, CAlead$238k–$318kPosted 9h ago

Job description

  • The Senior Staff Software Engineer for the Model LifeCycle team will play a crucial role in building a comprehensive managed platform for the entire application development lifecycle, with a specific focus on leveraging Machine Learning models, including Large Language Models (LLMs)

  • Manage fine-tuning systems for large foundation models (SFT, PEFT, LoRA, adapters), including multi-node orchestration, checkpointing, failure recovery, and cost-efficient scaling

  • Implement and maintain end-to-end training pipelines for Large Language Models

  • RFT and Reinforcement learning to the fine tuning and training sections

  • Distillation and reinforcement learning pipelines (e.g., preference optimization, policy optimization, reward modeling)

  • Dataset, model, and experiment management: versioning, lineage, evaluation, and reproducible fine-tuning at scale

Benefits

  • Health & wellbeing: Comprehensive health benefits designed to support your overall wellness

  • Time away: Paid time off for vacations, family bonding, and unexpected needs

  • 401(k) match: Build your financial future with our 401(k) matching program

  • Mental wellness: Resources and support for your emotional wellbeing and navigating life’s challenges- We’re looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services

  • If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe

  • 8+ years of industry experience leading and driving impactful projects in the AI Space

  • Passion for building cutting-edge AI products and solving challenging technical problems

  • Hands-on experience training, fine-tuning, and aligning LLMs using Reinforcement Learning and Reinforcement Fine-Tuning (RFT) techniques

  • Proactive and collaborative approach with the ability to work autonomously

  • Experience in Generative AI (Large Language Models, Multimodal)

  • Advanced degree in Computer Science, Engineering, or a related field

  • Performance optimizations on GPU systems and inference frameworks

  • Contributions to open-source AI projects such as vLLM or similar frameworks

  • Proficiency in Golang or Python for large-scale, production-level services and PyTorch