ForgeApply · Job listing
ML Platform Engineer
Hadrian-automation
Apply in about a minute — without sacrificing quality.
ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.
About this role
HADRIAN - MANUFACTURING THE FUTURE
Hadrian is building autonomous factories to reindustrialize America. By combining AI, advanced software, robotics, and full-stack manufacturing, we help aerospace and defense companies build rockets, satellites, aircraft, ships, and other mission-critical systems up to 10x faster and at significantly lower cost.
Following our $1.37B Series D at a $7.87B valuation, Hadrian is rapidly expanding our manufacturing footprint, launching new capabilities across welding, casting, forging, electronics, additive manufacturing, and more, while scaling our Factory-as-a-Service platform to transform how critical products are built.
Backed by leading investors including JPMorgan Chase, Valor Equity Partners, Andreessen Horowitz, Founders Fund, 137 Ventures, Lux Capital, T. Rowe Price, and Morgan Stanley, we’re building the future of American manufacturing—and looking for exceptional people to help make it happen.
If you’re ready to take on the most challenging and rewarding work of your career while helping create American manufacturing jobs for generations to come, you’re exactly who we’re looking for.
THE ROLE
This is an ML infrastructure role at the core of Hadrian’s technology stack. While Data Science, Operations Research, Vision, and Document AI teams build models, you will own the platform that ensures these models remain reliable, effective, and secure in production. You’ll standardize our deployment patterns built around MLflow, Dagster, ECR, FastAPI, and EKS, making them the backbone for packaging, evaluating, releasing, serving, monitoring, and rolling back models across Hadrian’s automated factories.
WHAT YOU’LL DO
- Build the production platform that enables Hadrian’s factories to safely depend on models for drawing extraction, cycle-time prediction, forecasting, scheduling, and more—with measurable performance and fast rollback.
- Develop shared batch and online serving for tabular, vision, document-AI, scheduling, graph, and embedding workloads, targeting clear SLAs for latency, availability, and isolation.
- Create repeatable release and evaluation processes featuring automated tests, reproducible artifacts, lineage, shadow deployments, canaries, and A/B tests.
- Own online feature serving and maintain contract integrity with offline feature tables; proactively detect and address training-serving skew, feature drift, bad data, and model degradation.
- Build operational tooling for telemetry, incident response, autoscaling, resource and GPU management, cost attribution, and secure model routing.
- Develop APIs, SDKs, reusable templates, and documentation that teams can adopt without requiring close support from platform engineers.
WHAT WE’RE LOOKING FOR
- Track record building and operating production ML infrastructure across multiple models or inference workloads.
- Strong production-level Python and SQL skills, including typing, testing, packaging, API design, and building observability features.
- Hands-on experience with Kubernetes, containers, and handling distributed-system failure modes such as retries, partial failures, idempotence, and resource isolation.
- Engineering background with model registries, feature systems, batch/real-time inference, experiment tracking, or model CI/CD workflows.
- Practical judgment around latency, throughput, availability, multi-tenancy, autoscaling, and infrastructure cost optimizations.
- Ability to build stable interfaces and collaborate closely with engineering and scientific stakeholders.
WHAT WILL SET YOU APART
- Experience implementing feature stores (Feast, Tecton, or internal systems).
- Production work with Ray Serve, KServe, Triton, BentoML, SageMaker, Vertex AI, or custom gRPC inference services.
- Experience serving and evaluating vision, document-understanding, embedding, or generative pipelines.
- Expertise in GPU inference optimization, multi-model serving, edge inference, or Go/Rust performance-sensitive AI services.
- Background in regulated environments or open-source contributions to ML infrastructure projects (MLflow, Feast, KServe, Ray).
COMPENSATION
Salary range: $170,000 – $300,000
This is the lowest to highest salary we reasonably and in good faith believe we would pay for this role at the time of posting. We may ultimately pay more or less than the posted range, and the range may be modified in the future. An employee's pay position within the salary range will depend on several factors, including relevant education, qualifications, certifications, experience, skills, geographic location, performance, and business needs.
BENEFITS
- Medical, dental, vision, and life insurance
- 401(k)
- Flexible vacation policy
ITAR REQUIREMENTS
To conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR), you must be a U.S. citizen, lawful permanent resident, protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State.
EQUAL OPPORTUNITY EMPLOYER
Hadrian provides equal employment opportunity to all applicants and employees. We do not unlawfully discriminate based on race, color, religion, sex, gender identity, gender expression, national origin, ancestry, citizenship, age, disability, medical condition, veteran status, marital status, sexual orientation, genetic information, or any other protected characteristic. Reasonable accommodations are available for qualified individuals with disabilities.
BENEFITS FOR FULL-TIME EMPLOYEES
- Medical, dental, vision, and life insurance plans for employees
- 401k
- Relocation support may be provided for certain situations, based on business need.
- Flexible vacation policy
- Equity
ITAR REQUIREMENTS
To conform to U.S. Government space technology export regulations, including the International Traffic in
Salary insight
The midpoint of this range ($235k) is about 52% above the median disclosed salary for Los Angeles roles listed on ForgeApply ($155k across 1,184 jobs).
See full DevOps / SRE salary data for Los Angeles →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Ready to apply to Hadrian-automation?
Apply in about a minuteSimilar jobs
- OS Platform Engineer — K2spacecorporation · Los Angeles, CA
- Principal Platform Engineer — First-resonance · Los Angeles
- Principal Platform Engineer — Clarityinnovates · Herndon, VA
- Senior Platform Engineer — Veeva · Remote
- Senior Platform Engineer — Gridcare · Remote
- Senior Platform Engineer — Bestow · Remote, US
- Senior Platform Engineer — Socket · Remote
- Senior Platform Engineer — Humeai · New York, New York, United States
More like this: DevOps & SRE Jobs · DevOps & SRE Jobs in Los Angeles · More jobs at Hadrian-automation · Browse all jobs