ForgeApply
Try it free

ForgeApply · Job listing

Lead Platform Engineer - NBA

Humana

Remote · US$129k – $178k

See all 255 open roles at Humana

Tailor your resume for this Humana job in about a minute.

ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for Humana's site. Free trial, no card required.

About this role

Become a part of our caring community   We are seeking a seasoned Lead Software Engineer to architect and deliver the foundational services that enable real time recommendations to become dependable, auditable, and scalable outcomes. In this position, you will own the design and implementation of the State Machine (managing authoritative state and legal transitions) and Transactional Outbox (ensuring exactly-once intent emission for downstream consumers). Your solutions must be robust, traceable, and maintain high performance under significant concurrency and latency demands. This role is hands-on, combining technical leadership with active engineering: you will architect systems, set technical standards, mentor peers, and collaborate closely across platform, data, and product teams.

About the Platform The Next Best Action (NBA) Decision Intelligence Platform is Humana's enterprise system for personalized, omnichannel member engagement. It decides the next best action for each Medicare Advantage member — what to communicate, when, and through which channel — and delivers it reliably across the member journey. The platform pairs a real-time decisioning and orchestration stack with reinforcement-learning and agentic-AI models trained on the Databricks Lakehouse, all operating under the compliance and auditability requirements of a regulated healthcare environment.   Role Summary The Lead Full Stack Engineer owns the NBA platform's quality, reliability, scalability, and operational excellence. This role takes responsibility for the platform once capabilities have been delivered by the feature pods, ensuring services are production-ready, observable, resilient, and capable of operating at enterprise scale. The role spans automated testing, platform observability, reliability engineering, and production readiness across a distributed ecosystem built on Java, Python, Node.js, and TypeScript. This is a hands-on technical leadership role: you write code, lead a small team of engineers and contractors, and are accountable for the platform's operational health, performance, and long-term sustainability.   Key Responsibilities •Pod delivery — Own delivery for the platform scale and reliability pod, including planning, execution, quality, and operational readiness across the NBA platform. •Platform reliability — Own uptime, performance, resilience, and operational excellence across platform services; proactively identify bottlenecks, failure points, and scaling constraints before they become production incidents. •Quality engineering — Build and maintain automated testing frameworks spanning unit, integration, end-to-end, regression, and performance testing; establish quality standards that all delivery pods must meet before release. •Hands-on development — Contribute production code alongside the team, building reliability tooling, test automation, observability capabilities, and platform infrastructure improvements. •Observability and monitoring — Own the platform's logging, metrics, distributed tracing, alerting, and monitoring strategy; ensure engineering teams have deep visibility into system behavior across environments. •Error and failure management — Define and enforce patterns for exception handling, fault tolerance, recovery, and operational diagnostics across distributed services. •Scalability engineering — Evaluate platform behavior under load and lead initiatives that improve performance, throughput, reliability, and cost efficiency as membership volume and engagement activity grow. •Production readiness — Define and enforce release readiness criteria, operational quality gates, and handoff standards that services must satisfy before entering production ownership. •Team leadership — Manage and mentor engineers and contractors; conduct code reviews, uphold engineering standards, and foster a culture of quality and operational excellence. •Cross-team coordination — Partner closely with decisioning, orchestration, activation, platform, and data teams to identify reliability risks early and ensure production concerns are addressed throughout delivery.

Use your skills to make an impact   Required Qualifications • Bachelor's degree in computer science or related field • 8+ years of full stack software engineering experience, with at least 1–2 years leading platform reliability, quality engineering, site reliability, or production operations initiatives. • Strong experience building production software across modern backend technology stacks, including Java, Python, Node.js, or TypeScript. • Hands-on experience designing and implementing automated testing frameworks, including integration testing, end-to-end testing, load testing, and quality automation. • Strong understanding of modern observability practices, including structured logging, metrics collection, distributed tracing, monitoring, and alerting. • Experience designing reliability patterns for distributed systems, including fault tolerance, retries, resilience, and failure recovery mechanisms. • Demonstrated ability to lead a small engineering pod and manage contractor resources while remaining an active individual contributor. • Experience operating cloud-native applications on Kubernetes and container-based platforms. • Clear communicator who can articulate operational risks, quality concerns, and engineering tradeoffs to technical leaders and senior executives. • Experience using modern AI-assisted development tools such as Claude, GitHub Copilot, or similar technologies to improve engineering productivity and quality.

Preferred Qualifications • Experience with Kubernetes platform operations, application scaling strategies, and reliability engineering in large-scale distributed environments. • Experience with Kafka and event-driven architectures, including diagnosing and mitigating failures across asynchronous systems. • Familiarity with Site Reliability Engineering (SRE) principles, service-level objectives (SLOs), service-lev

Tailor your resume for this Humana role before you apply.

Tailor my resume for this job

Similar jobs

More like this: DevOps & SRE Jobs · Remote DevOps & SRE Jobs · Browse all jobs

Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)