ForgeApply
Try it free

ForgeApply · Job listing

Senior AI Platform & Agentic Infrastructure Engineer

Okx

San Jose, California, US$178k – $321konsite

Apply in about a minute — without sacrificing quality.

ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.

About this role

Who We Are

At OKX, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom. OKX is a leading crypto exchange, and the developer of OKX Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). OKX is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. Across our multiple offices globally, we are united by our core principles: We Before Me , Do the Right Thing , and Get Things Done . These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er. OKX is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more.

About the Opportunity

OKX’s Internal Audit function has an early but working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data pipelines, and AI-enabled tools that let a very small team punch far above its weight. Your job is not to rebuild it at today’s maturity. Your job is to take it to a regulator-grade production standard and well beyond , and to drive AI Enablement across the department. You own the foundation, the agent runtime, and the harness: a resilient cloud platform, the agentic runtime and evaluation harnesses that make agents trustworthy in a regulated setting, and the governed data and AI infrastructure everything else depends on. We hire on demonstrated building , not on claims or credentials. We expect you to be more capable than the hiring manager in your domain: you will co-own and challenge the tooling strategy, not just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your counterpart’s stack when needed.

Environment

Cloud on AWS or GCP . Claude as the starting point in a deliberately multi-model architecture (Anthropic, OpenAI, Google), using the right model for each job, including the plugins and integrations you build. Google Workspace for reporting and evidence. JIRA for the audit team’s workflow. Lark for team communications, alert bots, and the corporate wiki. You inherit a working Google-native prototype estate (Apps Script web apps, Drive-synced automation, locally scheduled jobs, Claude Code agent tooling) and evolve it without breaking daily use. Infrastructure-as-code, CI/CD, and observability throughout. Treat this stack as the starting point, not a constraint: you build production-grade systems end to end on what exists today, and you are expected to propose, prove, and adopt better components as demand and capabilities evolve. Production-grade here means service levels sized for an internal assurance platform: board-cycle windows are sacred, recovery is measured in hours, and this is not a 24/7 pager culture.

In your first year

• Hive Mind runs in the cloud with HA, DR, SLOs, and audit logging that passes an internal controls review.

• A governed data foundation with provenance and lineage is live across multiple audit domains, integrated with OKX's group data infrastructure where it exists.

• An agentic runtime and harness with evaluation and red-teaming gates what reaches production.

• The hiring manager is out of the operational loop: no scheduled job runs on a personal machine, every system has a runbook and a non-founder owner. Decommissioning the founder’s laptop as infrastructure is a literal milestone. When these goals compete, the priority order is: keep the estate alive, then the cloud migration with observability, then the harness gating production, then the data foundation.

How we assess demonstrated building

We assess demonstrated building in ways that respect your time and your confidentiality obligations to current and former employers: a portfolio deep-dive (walk us through systems you built and kept running, at the level of detail your obligations allow; we want architecture, decisions, and trade-offs, never proprietary code, data, or documents), a short, time-capped build exercise on a synthetic problem unrelated to OKX's business (a small agentic workflow with an evaluation harness, used for assessment only and never put to use by OKX; the work remains yours) that you defend live, walking us through your design decisions and extending it on the spot, a systems-design session on taking a prototype estate to production, and references focused on whether you built and operated systems in production.

Trust and compliance

This role handles highly sensitive audit data at a global crypto exchange. Expect background checks, confidentiality obligations, and personal-trading and material-non-public-information

What You’ll Be Doing

• Inherit, operate, and progressively migrate the working prototype estate (Google Workspace–native automation across Apps Script, Drive, and the Docs/Sheets/Slides APIs; locally scheduled jobs; and Claude Code agent tooling) to the target platform without interrupting daily and board-cycle workflows. Working software wins arguments; migrate by strangling, not rewriting. Reuse and integrate with OKX's prevailing and evolving AI capabilities and data infrastructure, including enterprise-approved models and gateways, Model Context Protocol (MCP) servers and connectors, security tooling, and group data platforms, before building parallel capability.

• Re-architect Hive Mind into a resilient AWS or GCP platform with high availability (HA), disaster recovery (DR), defined service-level objectives (SLOs), and full observability, and own it in production.

• Build the agentic runtime and harness : orchestration, multi-model routing that sends each task to the right model, tool, MCP, and plugin integration across model providers, and the evaluation, red-team, and regression harnesses that grade agents before and after production, with evaluation gates, judge calibration, and cost controls that ho

Salary insight

The midpoint of this range ($250k) is about 23% above the median disclosed salary for San Francisco roles listed on ForgeApply ($203k across 6,339 jobs).

See full DevOps / SRE salary data for San Francisco

Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.

Ready to apply to Okx?

Apply in about a minute

Similar jobs

More like this: DevOps & SRE Jobs · DevOps & SRE Jobs in San Francisco · More jobs at Okx · Browse all jobs