ServiceNow logo

Principal Machine Learning Engineer

ServiceNow

HybridSanta Clara, us, Building A,B,C 2225 Lawson Lanelead$240k–$420kPosted 8h ago

Job description

  • AI Engineering and Delivery is the customer-obsessed engineering group building the agentic AI and enterprise-scale search systems that power Now Assist, AI Agents, and the AI-driven experiences our customers rely on every day. We build AI as foundational platform infrastructure — prioritizing robustness, performance, safety, and real-world customer impact at scale

  • Emerging tech is a small senior group inside AI Engineering and Delivery. We turn early bets on AI and emerging tech into strategic capability for our customers, our people, and ServiceNow. We are working to unlock features that will be helping our platform and products evolve in line with the Fast paced world of Agentic AI — prioritizing robustness, performance, safety, and real-world customer impact at scale

  • You will design, build, and help build out production-grade agentic AI systems embedded across ServiceNow’s platform — autonomous agents that reason over real enterprise data, take action across workflows, and stay safe at Fortune 500 scale

  • Agentic architecture. Design and ship multi-agent systems — orchestration, tool use, planning loops, memory, and failure recovery — that operate reliably in production, not in notebooks

  • Enterprise-grounded reasoning. Build agents that leverage ServiceNow’s data layer — CMDB, Workflow Data Fabric, and Knowledge Graph — to make decisions with context no frontier model has on its own

  • Trust, safety, and governance. Own the guardrails: observability, human-in-the-loop controls, and compliance infrastructure that make autonomous systems safe to deploy at scale

  • Retrieval and grounding. Work closely with our search team to ensure agents are grounded in accurate, low-latency retrieval — RAG pipelines, hybrid search, re-ranking, and evaluation — as a critical dependency of agentic quality

  • Model integration and evaluation. Integrate frontier models (Anthropic, Google, OpenAI) into the Sense → Decide → Act → Govern architecture; evaluate trade-offs across cost, latency, and capability for production use cases

  • Technical leadership and strong bias for action. Set the architectural patterns the group works from. Own the hard design calls, run the design reviews, and raise the bar on agentic design and production AI discipline across engineers and principals

  • Designing scalable and robust architectures that will support at scale deployment across hyperscalers and our own infrastructure

  • Work on emerging model capabilities and applying them to real world customer problems on a short timeline

Benefits

  • Generous family leave

  • Matched donations

  • Annual learning stipends

  • Flexible PTO

  • Competitive retirement plan

  • Paid volunteer time- Production-grade Python. Systems language (Go, Java, or C++) is a plus

  • 9+ years of software engineering with strong fundamentals in data structures, algorithms, and distributed systems

  • Track record of technical leadership: architecture ownership, code quality bar-raising, and mentoring engineers on production AI practices

  • Hands-on depth designing, shipping, and operating agentic systems in production — multi-agent orchestration, tool calling, planning loops, memory, and failure recovery. Not prototypes

  • Familiarity with RAG and retrieval patterns in production — vector stores, hybrid search, and retrieval evaluation metrics

  • Exposure to LLM fine-tuning or inference optimization in production

  • Nice to Have

  • Published work or open-source contributions in agentic systems or retrieval

  • Working experience with frontier AI SDKs (Anthropic, Google, or OpenAI) — prompt engineering, structured outputs, and model evaluation in production settings

  • Deeper specialization in search and retrieval at scale or MLOps/model observability