ForgeApply · Job listing
Senior/Staff Software Engineer, AI Agent Infrastructure
Nuro
Apply in about a minute — without sacrificing quality.
ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.
About this role
Who We Are
Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides.
Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles.
With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected.
Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors.
About the Team
Frontier models are fungible. Any team can rent the same intelligence we can, and the model we build on today will be replaced within a month. What is not fungible is the infrastructure that decides whether an autonomous system's output can be trusted — evaluation, verification, and the discipline to gate on evidence instead of impressions. Nuro has spent a decade building exactly that discipline for a robot that drives on public roads, and this team turns it inward: we build the platform that lets AI agents operate autonomously inside Nuro's own engineering organization, under the same standard of proof we apply to the vehicle.
Our mandate is to amplify the output of every engineer and researcher at Nuro by 100x. Not a better IDE, not a faster build — a change in what a single person can attempt. That number is a target, not a claim, and reaching it depends on one thing above all: autonomous work has to be trustworthy enough to run unattended. So our central ambition is to build the most rigorous closed-loop evaluation system for AI work anywhere. Leverage follows from trust, and trust follows from measurement.
We operate as a startup inside a company that has already shipped a hard thing. Small team, no established playbook, direct access to compute and to the systems we are automating. You will work directly with engineering leadership and the CEO, and the decisions you make will be yours to make rather than yours to implement.
About the Role
We already operate a substantial agent system in production — a fleet of agents with an extensive library of skills and plugins, integrated into the tools our engineers use daily. This role is about what it takes to make that system trustworthy, autonomous, and an order of magnitude more capable. Three things sit at the center of it.
Closed-loop evaluation. Our ambition is to build the most rigorous evaluation system for AI work anywhere — closed-loop, meaning every agent action produces a measurable outcome that feeds back into whether that agent is trusted to act again. Acceptance, revert, and override rates per workflow. Statistical honesty about whether a difference is real. Regression detection that fires before a human notices. Everything else on this team depends on this being right, and almost nobody has built it well.
Agent platform. The runtime that makes autonomous agents safe to run against real systems: orchestration, sandboxing and isolation, tool and skill frameworks, memory, identity and permissioning, and the gateway and observability layer underneath. Agents that touch production code and production infrastructure need containment and auditability before they need capability.
Autoresearch infrastructure. The automation of the research loop itself: agents that read the current state of a model and its metrics, form a hypothesis, launch an experiment, evaluate the result honestly, and either propose a change or discard the idea and move on. At Nuro that loop runs against the training pipelines behind the driving model — real experiments, real compute budgets, real metrics that determine whether a behavior ships. The hard parts are trusting the measurement, surviving experiments that take days, spending finite research compute wisely, and producing proposals a skeptical researcher can audit and reject.
Alongside this, the team builds agent-powered tooling across the engineering lifecycle — code generation, review, debugging, test and CI failure attribution, knowledge retrieval, triage. There is also appetite on this team for post-training our own models where an internal workload justifies it, and the engineer in this role would be central to that work.
About the Work What You Might Own in Your First Two Quarters
• Build the closed-loop measurement layer that tells us, per workflow, whether agent output is accepted, reverted, or overridden — and use it to decide where autonomy expands and where it gets pulled back.
• Take the autoresearch loop from assisted to unattended for a bounded class of experiments, including the eval and confidence machinery required to run it without a human in the loop.
• Design the isolation and permissioning model that lets agents act on production repositories and infrastructure with an auditable record of what they did and why.
About You
• 5+ years of software engineering experience (or 4+ with a Master's) in computer science, engineering, or equivalent practical experience. Staff-level candidates should bring correspondingly deeper scope and ownership.
• Deep, current taste in LLM research. You understand how a model is trained from scratch — data, tokenization, architecture, pretraining dynamics, the full post-training stack of supervised fine-tuning, preference optimization, and RL — and you can reason about what a training decision does to model behavior. You follow the literature because you want to, not because it is on a roadmap.
• You
Salary insight
This posting doesn't disclose pay. Across 6,425 San Francisco jobs with disclosed salaries on ForgeApply, the median is $204k.
See full Software Engineer salary data for San Francisco →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Ready to apply to Nuro?
Apply in about a minuteSimilar jobs
- Senior / Staff Software Engineer, AI Engineering — Suno · Boston
- Senior Staff Software Engineer (AI) — Junipersquare · Remote
- Senior Software Engineer, AI Infrastructure — Thealleninstitute · Seattle, WA
- Senior Software Engineer, Backend (AI Agent) — Cresta · Remote
- Senior/Staff Product Manager, AI Agents — January · New York City
- Senior Software Engineer - AI Infrastructure — Armada · Bellevue Office, Sunset Corporate Campus
- Senior Software Engineer, AI Infrastructure — Handshake · San Francisco, CA
- Staff Software Engineer, Core AI Infrastructure — Coinbase · Remote
More like this: Software Engineer Jobs · Software Engineer Jobs in San Francisco · More jobs at Nuro · Browse all jobs