ForgeApply · Job listing
Senior Platform AI Engineer
Drata
See all 43 open roles at Drata →
Tailor your resume for this Drata job in about a minute.
ForgeApply rewrites your resume for this exact posting, then autofills the application on Drata's site with it. You review everything before it's sent. Free trial, no card required.
About this role
Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from.
Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for:
- Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go.
- Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them.
- A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level.
- Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience.
- A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here https://drata.com/about/careers/life and follow us on LinkedIn https://www.linkedin.com/company/drata/posts/?feedView=all for company news, employee stories, and career updates.
Job Summary:
Drata's AI Platform team builds the production infrastructure that powers AI features across our compliance platform — from MCP servers that make Drata's data available to AI agents, to LLM workflow orchestration that automates SOC 2, TPRM, and policy analysis. You'll own the systems that sit between our AI models and our customers: tool definitions that agents actually understand, deployment pipelines that handle model upgrades without breaking output quality, and orchestration layers that manage multi-step agent workflows with persistent state.
This is not a traditional infrastructure role. You'll debug prompt templates alongside Terraform modules. You'll design API schemas optimized for LLM token budgets, not just HTTP throughput. When a model upgrade changes behavior across 15 workflows, you'll assess quality impact — not just confirm the containers are healthy.
You'll work closely with our agent developers, product engineers, and an embedded SRE partner, sitting at the intersection of AI development and production reliability.
Our north star is simple: minimize the time it takes to launch a new agent in production. You're someone who asks "are we solving the right problem?" before writing the first line of code, who builds systems that make five other engineers faster, not just yourself, and who's equally proud of what they chose not to build.
What you'll do:
MCP SERVER DEVELOPMENT & AI-OPTIMIZED API DESIGN
- Design and build MCP (Model Context Protocol) servers that expose Drata's platform to AI agents. This means making architectural decisions about tool granularity, naming conventions for agent disambiguation, response compression for LLM context windows, and workspace isolation for multi-tenant access. You'll own the protocol layer that determines whether agents can reliably find and use the right tools — writing semantic parameter descriptions, contextual hints, and tool schemas that optimize for model comprehension, not just developer ergonomics.
AGENT ORCHESTRATION & WORKFLOW INFRASTRUCTURE
- Build and operate the infrastructure for deploying multi-step agent workflows — state management across complex reasoning chains, tool routing and execution runtimes, and long-running agentic processes that persist over time. Own the orchestration layer that coordinates agent planning, tool calls, and human-in-the-loop patterns. Design systems that handle agent failure modes gracefully: retries on ambiguous tool outputs, fallback strategies when models produce unexpected results, and observability into multi-step execution traces.
LLM OPERATIONS & MODEL LIFECYCLE MANAGEMENT
- Own the operational side of our LLM workflows: model upgrades across production pipelines (assessing behavior changes, not just version bumps), prompt versioning and A/B testing, AI workflow deployment with custom container compatibility, and output quality monitoring.
- Manage token capacity planning — understanding model costs, context limits, batching strategies, and rate governance across workflows. When an AI workflow fails, you'll investigate whether it's a prompt template issue, a model behavior change, or an infrastructure problem. Making that distinction requires understanding both systems.
PRODUCTION AI INFRASTRUCTURE & RAG SYSTEMS
- Operate and evolve our production AI stack: vector storage and indexing (designing chunking strategies and metadata schemas for retrieval quality), document parsing pipelines, multi-region deployment, and cost optimization across LLM providers. You'll make RAG architecture decisions — embedding strategies, retrieval filtering, data model coordination — where the engineering challenge is search quality, not just system uptime. Implement caching layers and token-aware request routing to manage spend as AI workloads scale.
PLATFORM ENABLEMENT & DEVELOPER EXPERIENCE
- Build CI/CD patterns specific to AI workflows (reproducible deployments, SDK version compatibility, workflow rollback semantics). Own AI-specific observability — token usage dashboards, response quality metrics, agent execution traces, and cost-per-workflow tracking alongside traditional infrastructure monitoring. Enable product engineering teams to ship AI features faster by providing rel
Salary insight
The midpoint of this range ($226k) is about 13% above the median disclosed salary for San Francisco roles listed on ForgeApply ($200k across 8,773 jobs).
See full Machine Learning Engineer salary data for San Francisco →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Drata role before you apply.
Tailor my resume for this jobSimilar jobs
- Senior Platform Engineer, AI — Trueanomalyinc · Denver, CO or Long Beach, CA
- Senior AI Platform Engineer — Alpaca · Remote
- Senior AI Platform Engineer — Afresh · San Francisco, CA
- Senior AI Platform Engineer — Placerlabs · United States, Remote
- Senior AI Platform Engineer — Adobe · San Jose | San Francisco
- Senior Platform Engineer, AI Systems — Strike · Remote
- Senior AI Platform & Tools Engineer — Crunchyroll · Los Angeles, California, United States
- Senior AI Engineer - AI Platform — Clickup · Remote
More like this: Machine Learning & AI Jobs · Machine Learning & AI Jobs in San Francisco · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)