
Staff Software Engineer, AI Developer Productivity — Agent Platform & Evaluation
UNAVAILABLE
Job description
Role Summary
Rivian Autonomy is building an Applied General AI team to make AI a dependable part of how hundreds of engineers develop, test, and ship software. We are seeking a Staff Software Engineer to own and evolve the agent platform behind that effort, spanning code generation and review, debugging, CI triage, operational support, and knowledge retrieval across our cloud, data, simulation, and vehicle-software ecosystem.
Autonomy operates one of Rivian’s largest engineering data platforms, including petabyte-scale sensor data and ML training pipelines. This creates an unusually rich environment for agent systems: complex real-world workflows, valuable telemetry signals, and outcomes that can be objectively tested.
You will pair a production-grade agent platform with a rigorous evaluation loop so that every expansion of agent autonomy is supported by measured results on real Rivian work, not demos or anecdotes. Success means reducing the quality-adjusted effort required to complete important workflows while maintaining explicit security, reliability, software-quality, and developer-experience guardrails.
As an early member of the team, you will help define its technical direction, operating model, and future hiring. You will build on an existing in-house multi-agent system with real users and partners across all of Rivian Autonomy and beyond.
Responsibilities
Agent platform
-
Architect and build the core runtime for autonomous agents operating on Rivian systems, including orchestration, isolated execution, tool and skill frameworks, durable state and context management, model routing, policy enforcement, and end-to-end tracing.
-
Establish an execution and permission model with isolated workspaces, short-lived task-scoped credentials, human-in-the-loop approval gates, network and data-access controls, and complete audit trails.
-
Deliver capabilities across the engineering lifecycle, including code generation and review, debugging, test and CI failure attribution, documentation and knowledge retrieval, and operational triage, integrated with Slack, GitLab, Kubernetes, AWS, Databricks, and adjacent systems.
-
Own the reliability of the platform and the lifecycle of long-running agent work, including recovery, cancellation, resource controls, and human escalation.
Evaluation and continuous improvement
-
Instrument agent workflows to capture traces, tests, diffs, review dispositions, task outcomes, and what engineers keep, modify, or reject. Build evaluation sets from representative Rivian engineering tasks and calibrate model-based grading against human judgment.
-
Define repeatable, per-workflow success metrics; quantify run-to-run variance; and automatically detect regressions as models, prompts, context, and tools evolve.
-
Design controlled rollouts that measure changes in engineering effort, cycle time, rework, quality, reliability, and developer experience. Measurements will evaluate tools and workflows, not individual engineer performance.
-
Drive improvement through systematic experiments over prompts, context construction, tools, models, and inference budgets. Use the results to determine where agents earn greater autonomy and where they should be constrained.
Technical strategy and adoption
-
Define the technical strategy and roadmap for AI developer productivity across Autonomy, including build-versus-buy decisions, platform boundaries, security standards, and prioritization of the workflows with the greatest measurable impact.
-
Work directly with engineers to identify high-friction workflows and understand where agents fail in practice. Turn those findings into improvements to context, tools, skills, interfaces, documentation, and enablement.
-
Track credible advances in LLMs and agent systems and validate promising techniques on Rivian tasks before adopting them. Guide model tiers, reasoning depth, sampling, and verification strategies within explicit latency, reliability, and spend budgets.
-
Lead architecture across organizational boundaries, communicate recommendations to engineering leadership, and mentor engineers building on the platform.
Qualifications
-
6+ years of software engineering experience or equivalent demonstrated experience and impact, including substantial backend or distributed-systems work across cloud infrastructure, service design, storage, queuing, or secure execution.
-
Staff-level technical leadership: you identify the problems worth solving, shape strategy across teams, make pragmatic tradeoffs, and drive ambiguous initiatives from evidence to production.
-
Hands-on experience building and operating LLM applications, agent systems, or closely related developer infrastructure beyond prototypes, including tool use, orchestration, retrieval, context management, and production failure modes.
-
Demonstrated rigor in evaluating ML, LLM, or other nondeterministic systems. You have designed metrics or experiments for a real product, understand statistical variance, and can defend or challenge whether a measured improvement is real.
-
Strong understanding of the cost, latency, quality, and reliability tradeoffs that drive agent-system design.
-
Strong programming skills in Python and experience with at least one additional relevant language such as Go, Rust, C++, or TypeScript.
-
Strong communication and developer empathy, with a track record of building platforms or tools that engineers adopt and trust.
-
Self-directed in ambiguous problem spaces and comfortable defining scope where no established playbook exists.
Bonus Points
-
Experience building evaluation harnesses, benchmark or task suites, or calibrated model-graded evaluations.
-
Production experimentation experience, including A/B testing, causal inference, or offline-to-online metric correlation.
-
Experience with MCP or other agent-interoperability protocols, plugin and skill frameworks, LLM gateways, or model-routing layers.
-
Background in developer experience and tooling design for engineers.
-
Experience with AWS, Kubernetes, GitLab-based CI/CD, and large
Pay Disclosure
The salary range for this role is $206,500-$258,100 for San Francisco Bay Area based applicants. This is the lowest to highest salary we in good faith believe we would pay for this role at the time of this posting. An employee’s position within the salary range will be based on several factors including, but not limited to, specific competencies, relevant education, qualifications, certifications, experience, skills, geographic location, shift, and organizational needs. We offer a comprehensive package of benefits for full-time and part-time employees, their spouse or domestic partner, and children up to age 26, including but not limited to paid vacation, paid sick leave, and a competitive portfolio of insurance benefits including life, medical, dental, vision, short-term disability insurance, and long-term disability insurance to eligible employees. You may also have the opportunity to participate in Rivian’s 401(k) Plan and Employee Stock Purchase Program if you meet certain eligibility requirements. Full-time employee coverage is effective on their first day of employment. Part-time employee coverage is effective the first of the month following 90 days of employment. More information about benefits is available at rivianbenefits.com.