
Software Engineer - AI Eval & Automation
ServiceNow
Job description
Job Description:
It all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today โ ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500ยฎ. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.
Benefit options available through Magnit Global, depending on contract factors and upon meeting requirements.
Role Overview
Help build and scale the tooling we use to measure how well AI-powered software development tools actually perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.
Key Responsibilities
-
Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
-
Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
-
Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
-
Support execution-based benchmarking across quality, productivity, and efficiency measures, including cost and latency.
-
Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
-
Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.
Required Skills & Experience
-
Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
-
Proficient in Python, and comfortable in at least one of Java, JavaScript, or a similar language.
-
Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
-
Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
-
Some familiarity with how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.
-
Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
Preferred Experience
-
Hands-on work with AI-powered coding tools and agentic applications, such as Claude Code, Devin, or OpenCode.
-
Experience designing benchmarks or evaluations for software systems, especially execution based grading that verifies against tests.
-
Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them against human raters.
-
Experience with build-system-aware test selection, such as Bazel or mapping changed files to the tests that cover them.
-
Experience building reproducible test environments and managing versioned evaluation datasets.
-
Comfortable writing up methodology and results for engineering leadership.
Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Los Angeles Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, meet client expectations, standards, and accompanying requirements, and safeguard business operations and company reputation.