ForgeApply
Try it free

ForgeApply · Job listing

Staff Engineer - ML Infra / MLOps

Quince

Palo Alto, California, US$218konsite

Apply in about a minute — without sacrificing quality.

ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.

About this role

ABOUT QUINCE

Founded in 2018, Quince was built to challenge the idea that nice things have to cost a lot. Our mission is simple: to make really high quality essentials for really low prices, produced fairly and sustainably. We believe everyone deserves exceptional craftsmanship and timeless design without the traditional markups. Quince is a direct-to-consumer (DTC) model that cuts out middlemen and leverages just-in-time manufacturing to minimize waste and maximize value.

Quince is a tech company disrupting the retail industry by putting AI, analytics and automation at the center of everything we do. Our unwavering commitment to excellence and company values guide our teams and actions:

• Customer First : We prioritize customer satisfaction in every decision.

• High Quality : True quality means premium materials and rigorous production standards you can feel good about.

• Essential Design : We focus on timeless, functional essentials instead of chasing trends.

• Always a Better Deal : Innovation and transparency ensure value for both customers and partners.

• Social & Environmental Responsibility : We commit to sustainable materials, ethical production, and fair wages.

Quince partners with world-class manufacturers across the globe and serves millions of customers. With strong investor backing and a focus on sustainable growth, we are a company that is rapidly scaling while maintaining a commitment to quality, simplicity, and radical price transparency.

OUR TEAM AND SUCCESS

At Quince, you will be part of a high-performing team that is redefining what quality, value, and sustainability mean in modern retail. We are a destination for builders, innovators, and operators to come together and challenge the status quo. Our collective ambition is bold. We are creating an entirely new category and customer experience – one that democratizes luxury and provides high quality products at radically low prices. That mission demands a world-class team committed to excellence.

If you are motivated by impact, growth, and purpose, you will find a strong sense of belonging at Quince.

THE ROLE

Staff Engineer for ML Infra / MLOps

We are seeking a Staff ML Engineer t o join our growing team.

The ideal candidate is a deeply technical ML infrastructure engineer who combines hands-on mastery with system-level thinking. You have built and operated production-grade ML systems at scale — from distributed training pipelines and feature stores to high-throughput inference serving — and you take pride in engineering platforms that other engineers love to use. You don’t just build for today’s requirements; you design for extensibility, observability, and resilience.

You are the kind of engineer who gravitates toward the hardest problems — whether that’s optimizing GPU utilization at the tail of the cost curve, designing a zero-downtime model deployment system, or defining the architectural patterns that will define how Quince industrializes AI at scale. You operate with high autonomy, hold yourself to exceptional standards, and elevate the engineers around you through code reviews, technical mentorship, and by setting a bar for what great looks like.

Responsibilities

• Architect the ML Infrastructure Foundation: Own the end-to-end technical design of Quince’s ML platform — including model training, serving, feature pipelines, and monitoring — ensuring it is modular, scalable, and built for long-term extensibility.

• Build the “Paved Road” for Production: Design and implement the core developer experience for Quince’s Data Scientists and AI Researchers, enabling them to move from “idea to production” with minimal friction and maximum reliability.

• Drive Technical Excellence Across the Stack: Set and uphold engineering standards in CI/CD for ML, Infrastructure as Code (IaC), model versioning, experiment tracking, and deployment strategies (blue-green, canary) — and build the tooling that makes those standards the path of least resistance.

• Own High-Impact System Design Decisions: Lead the technical evaluation and selection of core platform components — from inference runtimes and feature stores to orchestration frameworks — with a clear-eyed view of build vs. buy tradeoffs.

• Optimize Compute Performance & Cost: Design and implement GPU utilization optimizations, model batching strategies, and cloud cost controls to maximize performance per dollar across training and inference workloads.

• Ensure Production Scalability & Reliability: Architect ML serving infrastructure that gracefully handles traffic surges, seasonal spikes, and model version transitions, with robust monitoring, alerting, and automated recovery.

• Mentor and Elevate the Engineering Team: Provide deep technical mentorship to junior and mid-level engineers through design reviews, code reviews, and pairing sessions — raising the collective technical bar without adding process overhead.

• Champion Operational Excellence: Lead root-cause analyses (RCAs) for production failures and drive systemic, permanent fixes over reactive patches. Model a culture of rigorous on-call discipline and accountability.

Qualifications:

• 8+ years of industry experience, with at least 4+ years of focused, hands-on work in ML Infrastructure, MLOps, or large-scale Data Platform engineering.

• Proven track record of designing and building MLOps platforms that support the full model lifecycle — from data ingestion and distributed training to real-time inference and model governance.

• Deep expertise in cloud-native infrastructure (preferably AWS), Kubernetes (EKS), Docker, and Infrastructure as Code tools (Terraform/Pulumi).

• Hands-on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker, with strong opinions on building a cohesive, high-leverage developer experience.

• Expertise in building Feature Stores and high-throughput data pipelines (Spark, Flink, Kafka), with a strong understanding of training/serving skew and data c

Salary insight

The midpoint of this range ($218k) is about 8% above the median disclosed salary for San Francisco roles listed on ForgeApply ($203k across 6,354 jobs).

Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.

Ready to apply to Quince?

Apply in about a minute

Similar jobs

More like this: More jobs at Quince · Browse all jobs