ForgeApply · Job listing
Staff Site Reliability Engineer
Datavant2
Apply in about a minute — without sacrificing quality.
ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.
About this role
Datavant is the data collaboration platform trusted for healthcare. Guided by our mission to make the world’s health data secure, accessible and actionable, we provide critical data solutions for organizations across the healthcare ecosystem - including providers, health plans, researchers, and life sciences companies. From fulfilling a single patient’s request for their medical records to powering the AI revolution in healthcare, Datavanters are building the future of how data is connected and used to improve health.
By joining Datavant today, you’re stepping onto a driven and highly collaborative team that is passionate about creating transformative change in healthcare.
We are seeking a Staff Site Reliability Engineer to play a key role in designing, optimizing, and securing the underlying cloud infrastructure that powers our organization’s cloud environments. As we continue to transform, consolidate, and evolve our cloud environments, this role will be instrumental in architecting scalable, secure, automated, and resilient infrastructure, ensuring seamless workload migration, and integrating new cloud environments into our ecosystem.
With a focus on core cloud infrastructure, including networking, identity and access management, security, storage, compute, and cross-cloud integrations, you will collaborate closely with peers in security, development, and other areas of the platform engineering team to ensure our cloud environment adheres to best practices, governance frameworks, and automation-first principles.
What You Will Do
Technical Leadership & Architecture
• Own the technical direction of SRE functions for designated teams from architecture definition and test planning through implementation and ongoing operations
• Lead design and development of tooling and processes that improve service delivery at scale: reliability, resiliency, efficiency, visibility, and quality
• Drive adoption of researched, verified, and standardized engineering practices across the teams you support
• Enhance system designs and infrastructure implementations to make them more efficient, standardized, and maintainable
• Take end-to-end ownership of major projects defining architecture, test plans, and implementation and empower others to do the same
• Push for and advocate superior architectural solutions; provide a compelling, data-backed case for change
• Drive transformation and standardization of foundational infrastructure services as the environment evolves and the business grows
M&A Integration & Environment Consolidation
• Architect and lead the integration of newly acquired cloud environments into Datavant's infrastructure, ensuring security, reliability, and consistency across an expanding multi-cloud footprint
• Design and implement hybrid-cloud and cross-cloud connectivity strategies to ensure interoperability as M&A activity adds new environments
• Own and evolve the modular network security edge, ensuring it remains flexible enough to unify additional environments while providing a consistent integration testing surface for development teams
• Drive environment consolidation initiatives bringing order to disparate hybrid and multi-cloud architectures and reducing operational fragmentation across the organization
• Collaborate with security, platform engineering, and development teams to ensure newly integrated environments adhere to Zero Trust principles, governance frameworks, and automation-first practices
• Optimize processes, develop policies, and enable secure, self-service infrastructure for engineering partners operating across consolidated environments
Operational Excellence
• Drive resolution of complex and systemic operational issues without needing guidance
• Lead enhancements and improvements to service delivery practices for all designated teams
• Define and evolve SLO/SLI frameworks and error budget policies across your scope
• Establish and continuously improve on-call practices, incident response processes, and postmortem culture
• Identify opportunities for efficiency gains and actively share that knowledge across the team
Mentorship & Team Enablement
• Serve as a mentor to SREs and developers on designated teams guiding technical growth, reviewing work, and raising the engineering bar
• Responsible for the mentorship and growth of SREs within your scope of influence
• Train and lead others on the team in completing complex tasks; delegate effectively and empower others to own architectures and implementations independently
• Drive best practices and standards for SRE functions within designated teams
Cross-functional Influence
• Partner deeply with development, security, and platform engineering teams to embed reliability, scalability, and operability into the software development lifecycle
• Drive improvements to development processes based on data-driven analysis and team feedback
• Communicate complex engineering challenges and trade-offs clearly to both technical and non-technical stakeholders
• Leverage AI tools and agents strategically to accelerate delivery and improve team workflows
Automation & Platform
• Lead the design and implementation of automation that improves platform reliability, deployment safety, and operational efficiency
• Architect and build scalable observability solutions instrumentation strategy, dashboards, alerting, and on-call tooling
• Own Infrastructure as Code standards and practices (Terraform, Ansible) for designated teams, writing IaC that supports simple, secure, and scalable deployments
• Implement policy-based automation (AWS SCPs, Azure Policy) to maintain governance across cloud environments, including newly integrated M&A environments
• Champion CI/CD pipeline improvements that reduce toil and increase delivery confidence
What You Need to Succeed
• 8+ years of experience in site reliability engineering, platform engineering, or infrastructure architecture
• Strong expertise in cloud infr
Ready to apply to Datavant2?
Apply in about a minuteSimilar jobs
- Staff Site Reliability Engineer — Diligentcorporation · New York, New York, United States
- Staff Site Reliability Engineer — Stackblitz · Remote
- Staff Site Reliability Engineer — Shein · San Diego
- Staff Site Reliability Engineer — Legora · New York City
- Staff Site Reliability Engineer — Earnin · Mountain View, US
- Staff Site Reliability Engineer — Figureai · San Jose, CA
- Staff Site Reliability Engineer — Idme · Mountain View, California, United States
- Staff Site Reliability Engineer — Flosports · Remote
More like this: DevOps & SRE Jobs · Remote DevOps & SRE Jobs · More jobs at Datavant2 · Browse all jobs