
Principal Software Engineer, ML & Distributed Systems
Microsoft
Job description
Overview
We are looking for a Principal Software Engineer to design, build, and operationalize machine learning systems and hyperscale distributed services that power Microsoft's AI experiences. This is a role for an engineer who is equally at home training and shipping ML models and building the large-scale distributed infrastructure that serves them — and who has the appetite to do both. You will own systems end to end: from model development and evaluation, through the serving and orchestration layers, to reliable operation at hyperscale.
MSN is a personalized content feed powering user experiences across Microsoft. Our mission is to empower every person on the planet to be informed, entertained, and inspired. With nearly 30 years of history, MSN has evolved into a premier content destination with high-quality content, AI-powered user-controlled personalization, and massive global reach. Over the past 4 years, AI and Machine Learning technologies have fueled massive growth, transforming MSN’s content moderation, personalization, and content entry points.Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Starting January 26, 2026, Microsoft AI (MAI) employees who live within a 50- mile commute of a designated Microsoft office in the U.S. or 25-mile commute of a non-U.S., country-specific location are expected to work from the office at least four days per week. This expectation is subject to local law and may vary by jurisdiction.
Responsibilities
-
Design, build, and operationalize scalable ML and deep learning models using containers and orchestration platforms (e.g., Kubernetes).
-
Develop and refine LLM prompt and fine-tuning strategies, build evaluation pipelines, and continuously optimize model quality, latency, and cost.
-
Architect, implement, and operate multi-tiered distributed services at hyperscale, with high availability, fault tolerance, and low latency.
-
Build model serving and inference infrastructure, including caching, batching, GPU capacity management, and A/B experimentation at scale.
-
Design scalable APIs, data pipelines, and feature/signal stores that ensure efficient, secure, and reliable data flow between ML systems and product surfaces.
-
Drive live-site excellence: instrumentation, monitoring, capacity planning, and incident response for ML-backed services.
-
Collaborate with applied scientists, data scientists, backend engineers, and product teams to translate requirements into production ML systems.
-
Participate in code reviews and architectural discussions, and mentor engineers across both ML and systems disciplines.
Qualifications
Required Qualifications:
-
Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR equivalent experience.
Preferred Qualifications:
-
Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
-
OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
-
OR equivalent experience.
-
-
Proven experience designing, developing, and operating multi-tiered distributed services at scale.
-
Hands-on experience building, deploying, and operating machine learning systems in production (model training, evaluation, and/or LLM-based applications).
-
Experience working through full product cycles, from initial design to final delivery.
-
Experience with LLM application patterns: prompt engineering, RAG, fine-tuning, and model evaluation frameworks.
-
Experience with ML infrastructure: distributed training, inference optimization, GPU capacity management, and orchestration platforms such as Kubernetes.
-
Experience with large-scale data systems (streaming, caching such as Redis, feature stores) and experimentation platforms.
-
Demonstrated ability to work across the ML/systems boundary and a strong desire to keep doing both.
#MicrosoftAI #Software Engineer #ML engineer
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.