ForgeApply
Try it free

ForgeApply · Job listing

Site Reliability Engineer

Datavant2

Remote · US

Apply in about a minute — without sacrificing quality.

ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.

About this role

Datavant is the data collaboration platform trusted for healthcare. Guided by our mission to make the world’s health data secure, accessible and actionable, we provide critical data solutions for organizations across the healthcare ecosystem - including providers, health plans, researchers, and life sciences companies. From fulfilling a single patient’s request for their medical records to powering the AI revolution in healthcare, Datavanters are building the future of how data is connected and used to improve health.

By joining Datavant today, you’re stepping onto a driven and highly collaborative team that is passionate about creating transformative change in healthcare.

We are seeking a Site Reliability Engineer to play a key role in designing, optimizing, and securing the underlying cloud infrastructure that powers our organization’s cloud environments. As we continue to transform, consolidate, and evolve our cloud environments, this role will be instrumental in architecting scalable, secure, automated, and resilient infrastructure, ensuring seamless workload migration, and integrating new cloud environments into our ecosystem.

With a focus on core cloud infrastructure, including networking, identity and access management, security, storage, compute, and cross-cloud integrations, you will collaborate closely with peers in security, development, and other areas of the platform engineering team to ensure our cloud environment adheres to best practices, governance frameworks, and automation-first principles.

What You Will Do

Reliability & Technical Ownership

• Lead the design and implementation of reliability improvements across assigned services, with minimal guidance

• Identify systemic inefficiencies in architecture, implementation, and operational process and drive solutions

• Implement customized solutions to complex operational problems derived from technical requirements

• Review code, systems, and configuration with a focus on efficiency gains, optimization, and best practices and hold peers to those standards

• Own SLO/SLI definitions for assigned services and drive teams toward meeting and improving those targets

• Lead incident response, facilitate postmortems, and ensure action items result in durable reliability improvements

M&A Integration & Environment Consolidation

• Support the integration of newly acquired cloud environments into Datavant's existing infrastructure, ensuring reliability, security, and operational consistency from day one

• Contribute to the implementation of hybrid-cloud and cross-cloud connectivity strategies that ensure interoperability across a growing multi-cloud footprint

• Help maintain and expand the modular network security edge, keeping it flexible enough to absorb additional environments as M&A activity requires

• Drive standardization of infrastructure and operational practices across consolidated environments, reducing fragmentation and toil

• Partner with security and platform engineering teams to ensure newly integrated environments adhere to governance frameworks and automation-first principles

Service Delivery & Process Improvement

• With limited guidance, develop tools and processes to improve team service delivery including scaling, resiliency, efficiency, visibility, quality, and operations management

• Analyze service delivery data and team feedback to drive meaningful improvements to development processes

• Participate in and help evolve on-call practices, runbooks, and alerting strategy for assigned teams

• Address communication gaps and produce clear documentation of process changes and technical standards

Collaboration & Mentorship

• Teach and lead more junior engineers on team processes, technical implementations, and SRE best practices

• Accept and promote sound engineering decisions including those that weren't your own and build alignment around them

• Collaborate cross-functionally with development, security, and platform engineering teams to embed reliability thinking into the software development lifecycle

Automation & Tooling

• Build and maintain Infrastructure as Code (Terraform, Ansible) to support scalable, repeatable, and secure deployments

• Develop and improve CI/CD pipelines, automated testing, and deployment tooling

• Implement policy-based automation (AWS SCPs, Azure Policy) to maintain governance across cloud environments

• Extend observability coverage through instrumentation, dashboards, and alerting using Datadog, CloudWatch, and Azure Log Analytics

• Leverage AI tools and agents to accelerate and improve daily engineering workflows

What You Need to Succeed

• 5+ years of experience in site reliability engineering, DevOps, or platform/infrastructure engineering

• Strong expertise in cloud infrastructure, including networking, security, compute, storage, and IAM

• Experience supporting workload migrations and integrating new cloud environments into existing architectures

• Hands-on experience with Infrastructure as Code and automation (Terraform and Ansible)

• Strong security knowledge, including IAM, encryption, network security, and compliance frameworks (SOC2, HITRUST, NIST)

• Demonstrated ability to solve complex operational problems independently and drive solutions end-to-end

• Strong proficiency in at least one systems language (Python, Go, or similar) and comfort across multiple languages and configuration formats

• Proven ability to conduct meaningful code and system reviews not just for correctness, but for efficiency and architectural soundness

• Strong communication skills, including the ability to document technical decisions, address gaps in understanding, and build team alignment

• Experience leveraging AI agents to accelerate daily workload

Competency with the following technologies:

• Compute: AWS EC2, Azure VMs, Kubernetes / containerized workloads

• Network security: AWS (VPC, TGW, Peering, SG, ALB, NLB); Azure (VNets, NSG, AGW); VPN

• Observabil

Ready to apply to Datavant2?

Apply in about a minute

Similar jobs

More like this: DevOps & SRE Jobs · Remote DevOps & SRE Jobs · More jobs at Datavant2 · Browse all jobs