ForgeApply · Job listing
Manager, Machine Learning Engineering
Gofundme
See all 31 open roles at Gofundme →
Tailor your resume for this Gofundme job in about a minute.
ForgeApply rewrites your resume for this exact posting, then autofills the application on Gofundme's site with it. You review everything before it's sent. Free trial, no card required.
About this role
Want to help us help others? We’re hiring!
GoFundMe is the world’s most powerful community for good, dedicated to helping people help each other. By uniting individuals and nonprofits in one place, GoFundMe makes it easy and safe for people to ask for help and support causes – for themselves and each other. Together, our community has raised more than $40 billion since 2010.
Join GoFundMe as our next Manager, Machine Learning Engineering (ML and AI Operations). In this role, you will lead the team responsible for the infrastructure, pipelines, and operational rigor that keep GoFundMe's machine learning and AI systems reliable, scalable, and safe in production. This role requires strong technical judgment across the ML lifecycle (data → training → online inference → monitoring), a strong understanding of how to enable AI applications to operate safely at scale, and a proven ability to build and lead a high performance team that operates production ML/AI systems with the same rigor as core infrastructure.
Candidates considered for this role will be located in the San Francisco Bay Area. There will be an in-office requirement of 3x a week.
The Job
• Own the reliability, scalability, and operational health of ML/AI production systems across GoFundMe, including training pipelines, feature stores, model serving, and monitoring/observability infrastructure.
• Lead, hire, and grow a team of ML/AI operations engineers, setting technical direction through design reviews, architecture decisions, and shared best practices for production ML and AI systems.
• Partner with data science and ML engineering teams to streamline the path from model development to production deployment, including CI/CD for ML, model packaging, versioning, and rollback strategies.
• Establish ML operational excellence org-wide by driving standards for model observability (latency, errors, drift, calibration, business KPI deltas), automated retraining triggers, and incident response playbooks.
• Build and mature on-call processes, SLOs/SLAs, and postmortem practices for ML/AI systems, treating model incidents with the same discipline as production infrastructure incidents.
• Drive operational strategy for GoFundMe's generative AI systems alongside traditional ML, balancing innovation velocity with safety, compliance, cost, and reliability.
• Collaborate cross-functionally with Product, Engineering, Design, and Legal/Privacy stakeholders to translate business goals into team priorities and measurable operational outcomes.
• Manage vendor and platform relationships (e.g., cloud ML platforms, LLM providers) and make build-vs-buy calls that balance cost, control, and speed.
• Report on team health, system reliability metrics, and operational risk to senior engineering leadership.
• Employ a diverse set of tools and platforms, including Python, AWS, Databricks, Docker, Kubernetes, Terraform, Snowflake, and GitHub, to guide your team in developing, deploying, and maintaining scalable and robust machine learning systems.
You
• 7+ years of hands-on experience building and shipping production machine learning systems, with demonstrated ownership of backend services and ML pipelines in a high-availability environment.
• 1-3+ years of experience directly managing engineers, ideally in an MLOps, ML platform, or infrastructure context, with a track record of hiring and developing strong teams.
• Strong proficiency in Python and ML libraries/frameworks such as PyTorch, TensorFlow, Scikit-learn, plus strong software engineering fundamentals (testing, code review, CI/CD, API design, performance, and reliability) — enough depth to stay hands-on and credible with your team.
• Experience designing and operating real-time model serving at scale, including containerization, scalable inference, feature retrieval, and safe rollout strategies (canaries, shadowing, backward-compatible schema evolution).
• Strong data engineering fluency: building reliable datasets and features using SQL, Spark/Databricks, and warehouse technologies (e.g., Snowflake), with an understanding of event semantics, identity resolution, and data quality controls.
• Proven experience implementing ML monitoring for both technical and business metrics (drift, calibration, segment performance, latency, error budgets) and running models reliably in production.
• Familiarity with generative AI/LLM infrastructure and operational considerations (latency, cost, safety guardrails) is a strong plus.
• Ability to break down ambiguous, high-impact problems, define crisp interfaces and success metrics, and deliver iteratively while managing stakeholder expectations across engineering leadership, product, and data science.
• Strong leadership and mentoring skills and a proven ability to raise the bar on architecture, engineering quality, and operational rigor for production ML/AI systems.
• Advanced degree (Master's or Ph.D.) in Computer Science, Statistics, Data Science, or a related technical field is preferred.
• Sense of humor is optional but appreciated.
Why you’ll love it here
• Make an Impact : Be part of a mission-driven organization making a positive difference in millions of lives every year.
• Innovative Environment : Work with a diverse, passionate, and talented team in a fast-paced, forward-thinking atmosphere.
• Collaborative Team : Join a fun and collaborative team that works hard and celebrates success together.
• Competitive Benefits : Enjoy competitive pay and comprehensive healthcare benefits.
• Holistic Support : Enjoy financial assistance for things like hybrid work, family planning, along with generous parental leave, flexible time-off policies, and mental health and wellness resources to support your overall well-being.
• Growth Opportunities : Participate in learning, development, and recognition programs to help you thrive and grow.
• Commitment to DEI : Contribute to diversity, equity, and inclusion through ongoing initiatives and
Salary insight
The midpoint of this range ($274k) is about 37% above the median disclosed salary for San Francisco roles listed on ForgeApply ($200k across 8,354 jobs).
See full Machine Learning Engineer salary data for San Francisco →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Gofundme role before you apply.
Tailor my resume for this jobSimilar jobs
- Manager, Machine Learning Engineering — Warner Bros. Discovery · WA Bellevue 205 108th Avenue NE, Suite S200 | CA San Francisco 153 Kearny Street | NY New York 30 Hudson Yards
- Senior Manager, Machine Learning Engineering — Metropolis · Seattle, Washington, United States
- Sr Manager, Machine Learning Engineering — Adobe · San Jose
- Senior Engineering Manager, Machine Learning — Signifyd95 · Remote
- Senior Manager, Machine Learning — Upstart · Remote
- Engineering Manager – Machine Learning Engineering — The Aerospace Corporation · Chantilly, VA
- Senior Manager, Machine Learning Platform Engineer — Gilead Sciences · Foster City, California, United States
- Manager, Machine Learning Engineer (US) — Clio · Remote
More like this: Machine Learning & AI Jobs · Machine Learning & AI Jobs in San Francisco · Browse all jobs
Free ATS checker · How to Autofill Greenhouse Job Applications (Without Sending Junk)