Software Engineer – AI Model Evaluation
Office Staff Now
Job description
Software Engineer – AI Model Evaluation (Technical Advisor)
Remote (United States) | W2 Contract | $150/hour | 6-Month Contract (Possibility of Extension)
Shape the Future of AI — From the Inside
Our client is a leading AI safety and research company dedicated to building AI systems that are reliable, beneficial, and genuinely understandable. Right now, they're focused on a critical question: how well do today's most advanced coding agents actually perform on real, production-grade software engineering work? This is your opportunity to help answer that question directly, working alongside one of the most respected names in AI research.
As a Software Engineer – Technical Advisor, you won't be writing product code from scratch. Instead, you'll serve as a high-bar technical authority — reviewing model-generated pull requests, dissecting full agent session logs, and building rigorous test problems that expose exactly where and why frontier AI models succeed or fail at real engineering tasks. This role demands sharp code review instincts, deep production experience, and the ability to put rigorous technical reasoning into writing.
What You'll Do
-
Audit model-generated pull requests against real production repositories, documenting every issue with severity ratings and detailed technical rationale
-
Evaluate full coding-agent session logs to determine what the model investigated, verified, assumed, or missed
-
Design and build container-based (Docker) technical benchmarks used to test AI models
-
Write clear, original technical rationales explaining why code fails — all written work must be self-authored, without AI writing tools
-
Collaborate asynchronously with AI researchers to share findings and help refine evaluation criteria
What You Bring
-
8+ years of production engineering experience (exceptions considered only for truly exceptional profiles)
-
Background as a Senior, Staff, or Principal Software Engineer, Tech Lead, or Open-Source Maintainer
-
Experience working in production codebases with a strict code review culture — startup, big tech, or open source
-
Polyglot adaptability: strong Python and TypeScript experience is common, but comfort picking up unfamiliar languages weekly is a must
-
Solid command of Docker, git, and CLI to reproduce, isolate, and debug issues locally
-
Cross-layer fluency across backend, frontend, APIs, data, testing, or developer tooling
-
Excellent written communication — able to clearly explain why code fails, not just how to fix it
This Role Is Not a Fit If You...
-
Are an Engineering Manager or Director who no longer writes or reviews code regularly
-
Have experience limited to low-code, no-code, or single-framework MVPs without production code review exposure
-
Work primarily as a QA tester or non-technical prompt engineer
-
Lack hands-on Docker and terminal/CLI experience
-
Are unwilling to complete a CodeSignal screening or write original, self-authored technical rationale
Schedule
30–40 hours/week, mostly asynchronous. You'll set your own schedule locally (not restricted to Pacific time), but must be reachable during US business hours for periodic syncs.
Why This Role
This is a rare chance to work at the cutting edge of AI evaluation with one of the most influential companies in the space — using your deep engineering expertise to directly shape how the next generation of AI coding tools is built, tested, and trusted.
Pay: $150.00 per hour
Expected hours: 30.0 – 40.0 per week
Work Location: Remote