ForgeApply
Try it free

ForgeApply · Job listing

Research Engineer

Tessera Labs

San Jose Office (HQ), US$200k – $300konsite

See all 13 open roles at Tessera Labs

Tailor your resume for this Tessera Labs job in about a minute.

ForgeApply rewrites your resume for this exact posting, then autofills the application on Tessera Labs's site with it. You review everything before it's sent. Free trial, no card required.

About this role

ABOUT TESSERA LABS

Tessera Labs is a new category of enterprise software: an AI platform that changes how the world's largest companies run.

Every large enterprise carries the same weight — decades of accumulated process, data, and code that no longer match the business it has become. Changing any of it is a program measured in years and hundreds of millions of dollars, staffed by armies of consultants, and it fails more often than anyone admits. Most companies have quietly accepted this as the cost of being large.

We don't. Tessera is a transformation engine: a governed, multi-agent platform that understands an enterprise's process, data, and code as one connected system and changes it in weeks rather than years. We're vendor-agnostic by design — SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft — and tied to none of them.

Two things make this hard, and they're the reason the research is interesting. Governance: every action is logged, traceable, and reversible, because our customers are regulated and these are the systems that close their books. And generality: the platform has to work on landscapes it has never seen, at companies whose complexity is genuinely unique to them.

We sell a product, not a service. Our people are here to make the product successful, not the other way around — which is also why research here is a durable investment rather than a line item on an engagement.

We raised a $60M Series A led by Andreessen Horowitz, with Foundation Capital, Myriad Venture Partners, and Osage University Partners participating.

ABOUT THE ROLE

Building this platform has two halves — the systems that surround the model, and the model itself. This role owns the second, at scale.

Frontier models have never seen most of what we work on. Proprietary dialects, customer-specific configuration, semantics that exist in no public corpus, and tasks that run forty steps before anything tells you whether you were right. Closing that gap with post-training, environments, and evaluation is a research problem, and Research Engineers are the people who make it happen at scale.

You build the training, environment, evaluation, and inference machinery that turns a hypothesis about agent behavior into a measured result, and then into a model that ships. You'll work closely with Research Scientists and own experiments of your own within weeks. The division isn't seniority or idea ownership — Research Engineering owns the machinery and is accountable for it working at scale; Research Science owns the agenda and is accountable for the result being true.

One property makes this an unusually good RL setting: much of our task space is verifiable. A transformation either produces a system that builds, passes the customer's regression suite, and behaves equivalently, or it doesn't. That's a real reward signal rather than a preference model, and building the environments that make it cheap and trustworthy is a large part of this job.

We post-train open-weight models on rented clusters. We're compute-constrained relative to a frontier lab and we buy more when a result justifies it — worth knowing up front.

WHAT YOU'LL DO

- Build and scale the post-training stack: SFT, preference optimization, and reinforcement learning for long-horizon tool use, transformation, and reconciliation over enterprise systems. RL is the center of gravity of this role, not a side interest.

- Build the memory and context machinery that long-horizon agents run on: what an agent retains across a forty-step run, how it's structured, retrieved, compacted, and revised — and how you train a model to use it well rather than bolting it on at inference.

- Build the representation layer that agents reason over — ontologies and knowledge graphs derived from real enterprise systems — and the pipelines that construct, validate, and keep them current.

- Design and implement data generation and curation pipelines — synthetic landscapes, transformation traces, tool-call trajectories, curriculum infrastructure — that teach models to operate systems no public model has seen.

- Build the RL environments: sandboxed landscapes and execution-and-verification harnesses where a change can be applied, checked, and scored automatically, backed by design-partner traces where synthetic data won't do.

- Build and run the offline eval harness for long-horizon agentic behavior — trajectory-level scoring, task suites, and the infrastructure that makes a result reproducible six weeks later. You own the harness; the methodology it implements is a shared argument with Research Science.

- Run experiments end to end — design, launch, debug, analyze — and be honest about which effects are real and which are noise.

- Optimize training and inference throughput: kernels, parallelism strategies, memory, batching, serving. Long context matters here more than most places; a single enterprise artifact can eat a context window.

- Take a training result from "the eval moved" to "it's serving traffic" — quantization, serving configuration, rollback path.

- Establish standards for reproducibility, experiment tracking, and result hygiene, so findings survive contact with the next person who builds on them.

REPRESENTATIVE PROJECTS

- Building a synthetic landscape generator that produces enterprise systems with realistic complexity and coupling, then measuring what training on it actually buys on held-out real customer environments.

- Standing up a distributed RL loop where reward comes from build-and-regression outcomes, and catching that the environment was leaking target state into the observation — the policy was scoring 0.9 by reading the answer rather than doing the task.

- Post-training a mid-size open-weight model to match frontier-model accuracy on our core transformation tasks at a fraction of the serving cost.

- Halving end-to-end latency on a multi-agent run: prefix KV-cache reuse across sub-agents so we stop re-prefilling the

Salary insight

The midpoint of this range ($250k) is about 25% above the median disclosed salary for San Francisco roles listed on ForgeApply ($200k across 8,659 jobs).

Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.

Tailor your resume for this Tessera Labs role before you apply.

Tailor my resume for this job

Similar jobs

Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)