ForgeApply
Try it free

ForgeApply · Job listing

Member of Technical Staff, Performance & Capacity

Psi

USonsite

See all 17 open roles at Psi

Tailor your resume for this Psi job in about a minute.

ForgeApply rewrites your resume for this exact posting, then autofills the application on Psi's site with it. You review everything before it's sent. Free trial, no card required.

About this role

OVERVIEW

Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking engineers to build platform infrastructure at the intersection of computational science, AI systems, and software engineering.

Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.

The last century's golden age of physics gave us transistors, lasers, and nuclear energy. We believe artificial superintelligence will unlock the next one. We're creating the infrastructure to industrialize scientific discovery and usher in this new era.

We have one product: new physics, at scale.

We are seeking a Member of Technical Staff, Performance & Capacity to answer two questions honestly: how much compute do we actually need, and how well are we using what we have. Your job is to make both answers measurements rather than estimates, and then be held to them. Then make the same fleet produce more, quarter after quarter.

ROLE AND RESPONSIBILITIES

- Measure the fleet instead of estimating it. Utilization, goodput, and training efficiency, taken on real workloads, so that capacity decisions rest on numbers someone actually observed. Where the platform wastes capacity, you find it and you close it.

- Own the capacity model. Turn a research workload into a defensible node count with stated assumptions and error bars, then defend it when real money is committed against it. What we ask providers for comes out of that model. You own the model and the measurements feeding it.

- Own the workload mix over time: what the fleet runs at a given hour, the preemptible fraction as a measured quantity, and checkpointing economics. A workload that can yield on demand is cheaper to run, and worth money in a negotiation. Both of those depend on someone having measured the fraction.

- Keep the data path fast enough to matter: parallel file systems and data locality feeding GPUs during training, interconnect behavior at multi-node scale. Find where the time actually goes, not where the profiler summary says it goes.

WHAT WE'RE LOOKING FOR

- Five or more years with GPU and large-scale compute workloads, including real multi-node experience: distributed training performance, interconnect and collective-communication behavior, and the patience to find where scale breaks down. Single-box optimization is not this job.

- You know GPU performance characteristics at the system level: which workloads justify an H100 and which a B200, how to optimize across clusters rather than within one, and where the bottleneck actually sits. Systems view first, the weeds when the numbers demand it.

- You understand AI training and inference performance deeply. On training: where step time goes, what MFU means and why it disappoints, how communication hides behind compute or fails to. On inference: why decode is bandwidth-bound, what batching buys, and what actually sets the latency floor. You can find the bottleneck in either regime and say what it costs.

- You have owned a capacity decision with money attached and been held to the number. You can walk us through the assumption that turned out wrong and what you changed.

- Parallel file systems and their effect on training throughput. InfiniBand or RDMA-class networking behavior. Fluent, not aware.

- You can explain a performance result to a researcher without making them learn the internals. Your numbers have to be usable by people who will never open a profiler.

NICE TO HAVE

- Time on a data center floor: thermal and power envelopes, physical failure domains, capacity planning against real hardware rather than a console.

- Energy- or preemption-aware scheduling. You have treated power, time of day, or interruptible capacity as real variables in where work lands.

- Capacity sourcing across cloud and specialist GPU providers, and the economics of moving workloads between them.

- Background in scientific computing or HPC environments, where efficiency was measured because someone paid for the machine.

HOW WE WORK

We hold a high technical bar and give people full ownership of their work, from spec to ship to on-call. We write contracts before logic, test against real systems instead of mocks, and favor simple designs that ship over clever ones that do not. Our development process is AI-native: we work with agentic coding tools daily, write specs that are legible to humans and agents alike, and lead with leverage.

LOCATION AND COMPENSATION

This role is based in Boston. We will consider remote candidates on a case-by-case basis. We offer competitive compensation including salary, benefits, and meaningful early-stage equity. We evaluate on technical breadth, systems thinking, scientific curiosity, and shipping velocity. We are an equal opportunity employer and value diverse perspectives in building platforms for AI-driven discovery.

Salary insight

This posting doesn't disclose pay. Across 1,915 Boston jobs with disclosed salaries on ForgeApply, the median is $160k.

Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.

Tailor your resume for this Psi role before you apply.

Tailor my resume for this job

Similar jobs

Free ATS checker · No Salary on the Job Posting? How to Find the Number Before You Interview