ForgeApply · Job listing
RL Environments Engineer
Bespokelabs
See all 10 open roles at Bespokelabs →
Tailor your resume for this Bespokelabs job in about a minute.
ForgeApply rewrites your resume for this exact posting, then autofills the application on Bespokelabs's site with it. You review everything before it's sent. Free trial, no card required.
About this role
ABOUT BESPOKE LABS
Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.
Recently, we curated Open Thoughts https://open-thoughts.ai/, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B https://www.bespokelabs.ai/blog/bespoke-minichart-7b and Bespoke-MiniCheck https://www.bespokelabs.ai/bespoke-minicheck, and taught https://www.bespokelabs.ai/blog/improving-multi-turn-tool-use-with-reinforcement-learning agents to do multi-turn tool-calling with reinforcement learning.
Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.
ABOUT THE ROLE
You will own coding environments end to end. You choose what world to build, design the tasks inside it, build the grading, run frontier models against it, and harden it until the only way to pass is to actually do the work.
This is a delivery role. We care most about whether you have done this before and can point to what came out of it. If you have built agentic coding environments or tasks at real volume and can tell us how many and how hard they were, we want to talk.
WHAT YOU'LL DO
Build high-fidelity coding worlds around real codebases, with the conventions, dependencies, tooling, and accumulated mess that real software has.
Choose which environments are worth building. A strong coding environment hits several marks:
- Targets work where frontier models measurably struggle
- Exercises real engineering, meaning navigation, diagnosis, sequencing, and design, and not just writing a function
- Rests on a codebase with enough history and structure that shortcuts do not survive
- Has a clear pass condition that a reviewer would agree with
- Produces many varied tasks from a single world rather than one
Design tasks across the full lifecycle. Prompt, environment, grader, running frontier models, failure analysis, and iteration, until each task is rigorous, fair, and hard to game.
Build grading and sandboxed execution that is deterministic and cannot be gamed. Assume the model will try to pass without doing the work, and close the door before it finds it.
Remove whatever is slowing the team down. Build the internal tooling that makes everyone around you faster.
Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce.
WHAT WE'RE LOOKING FOR
A record of shipped volume. You have built agentic coding tasks or environments and can show us how many you personally drove and what they cost to produce.
Experience scaling that output through automation rather than through more people doing more manual work.
Strong software engineering fundamentals and fluency in several languages that holds up in production code.
Real experience with production software. Large codebases, build systems, testing, deployment, on-call, and root cause analysis. You know what real engineering work feels like because you have done it.
An adversarial mindset. You look at a grader and ask how a model would cheat it, and then you fix that.
A clear sense of what frontier coding agents can and cannot do, and where they cut corners.
Ownership. You build, debug, and ship without much supervision.
YOU MAY BE A GOOD FIT IF YOU ALSO
- Have worked on RL training systems, post-training, verifiers, or tool-use harnesses
- Come from developer tooling, CI/CD, sandboxes, or code execution infrastructure
- Have built large-scale automated test generation, fuzzing harnesses, or benchmark suites, which is close cousin work even if it was never called an RL environment
- Have contributed to a public agentic benchmark such as Terminal-Bench
- Have open-source work that other people depend on
WHAT WE OFFER
- Location: Mountain View, CA (Onsite)
- Base Salary: $250,000 – $300,000 USD / year
- Additional Comp: 25% performance-based bonus + equity
Benefits & Perks
- Health, dental, and vision coverage
- 401(k)
- Daily onsite lunch provided
- Visa sponsorship and relocation support available
- Direct impact on how the industry trains and evaluates agents
We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.
Salary insight
The midpoint of this range ($275k) is about 38% above the median disclosed salary for San Francisco roles listed on ForgeApply ($200k across 8,498 jobs).
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Bespokelabs role before you apply.
Tailor my resume for this jobSimilar jobs
- Software Engineer - RL Environments — Afterquery · San Francisco
- Senior Software Engineer, RL Environments — Pareto AI · San Francisco Office, California
- Forward Deployed Engineer, RL Environments — Labelbox · San Francisco Bay Area
- Dynamic Environments Engineer — Stokespacetechnologies · Kent, Washington
- Fullstack Engineer, RL Environments — Mercor · San Francisco
- Research Engineer, Real Environments — Mercor · San Francisco
- Sr Dynamic Environments Engineer — Relativity · Long Beach, California, United States
- Member of Technical Staff, Post-Training, RL Environments — Mirendil · San Francisco
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)