ForgeApply
Try it free

ForgeApply · Job listing

Data Platform Infrastructure Manager

SailPoint

United States, US$125k – $210konsite

See all 46 open roles at SailPoint

Tailor your resume for this SailPoint job in about a minute.

ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for SailPoint's site. Free trial, no card required.

About this role

About SailPoint

SailPoint is the leader in identity security for the cloud enterprise. Our identity security solutions secure and enable thousands of companies worldwide, giving our customers unmatched visibility into the entirety of their digital workforce and ensuring workers have the right access to do their jobs—no more, no less. Built on a foundation of AI and machine learning, SailPoint Identity Security delivers the right level of access to the right identities and resources at the right time—matching the scale, velocity, and changing needs of today’s cloud-oriented, modern enterprise.

About the role

As the Data Infrastructure Platform Manager, you will lead the team responsible for the runtime and platform layer beneath SailPoint’s batch, streaming, analytics, and machine learning workloads. Your team owns the availability, scalability, security, performance, lifecycle, and cost efficiency of shared services such as Airflow, Flink, Spark on AWS EMR, Kafka, Snowflake, and Iceberg. Data and product engineering teams own the workload-specific pipelines and processing logic that run on those services.

Your day will span people leadership, production operations, technical strategy, and cross-functional execution. You will review service health and incidents, set priorities across operational and roadmap work, coach and unblock engineers, make architectural and investment tradeoffs, and partner with data engineering, Developer Platform, SRE, Observability, Security, and Infrastructure teams. You will make the platform easier to consume through paved roads, CI/CD, configuration as code, observability, automation, and self-service—all while protecting reliability, quality, and cost efficiency in a fast-moving environment.

About the team

The Data Infrastructure Platform team designs, builds, and operates the production-grade data processing infrastructure that powers SailPoint Identity Security. We provide reliable, scalable, and secure data platforms as services so data and product engineers can focus on DAGs, streaming jobs, pipelines, models, and business logic rather than provisioning and operating the underlying infrastructure. The team values engineering and operations excellence, service ownership, practical automation, constructive debate, continuous learning, and a “strong opinions, loosely held” mindset.

Roadmap for success

Success in this role will be measured through the following outcomes:

30 days • Build working relationships with team members and key partners across data and product engineering, Data Operations, Developer Platform, SRE, Observability, Security, and Infrastructure. • Document and align stakeholders on the team charter, service catalog, ownership boundaries, escalation paths, and the distinction between platform ownership and workload-specific pipeline ownership. • Complete an initial assessment of the team’s people, platforms, roadmap, on-call load, incidents, operational risks, capacity, and cloud costs; share a prioritized set of immediate risks and opportunities. • Establish a regular operating cadence for team priorities, service health, incidents, roadmap delivery, and cross-team dependencies.

90 days • Publish an outcome-oriented 12-month platform roadmap, developed with technical leads and partner teams, that balances reliability, scalability, security, developer experience, and cost. • Define the operating model for the team’s highest-criticality services, including named ownership, service-level objectives, observability expectations, on-call practices, capacity planning, lifecycle management, and incident follow-up. • Select and begin delivery of the first high-value paved-road or self-service improvement for deploying and operating Airflow DAGs, Flink or Spark jobs, or another priority data workload. • Set clear performance expectations and development goals for each team member, identify capability or staffing gaps, and establish a hiring and development plan where needed. • Baseline the platform’s key reliability, delivery, toil, utilization, and cost measures so subsequent improvements can be demonstrated with data.

6 months • Have service-level objectives, actionable dashboards, alerts, and recurring service reviews in place for the highest-criticality data platform services, with measurable progress against the 90-day reliability baseline. • Deliver at least one production self-service or standardized delivery capability that reduces the effort and lead time required for developers to deploy data workloads safely across supported environments. • Implement a cost and capacity management program with service-level visibility, accountable owners, prioritized optimization work, and documented efficiency gains. • Strengthen incident response, change management, disaster recovery, vulnerability remediation, and operational runbooks; demonstrate reduced recurring toil or faster recovery for priority failure modes. • Establish a clear platform approach for supporting machine learning workloads, including appropriate use of AWS SageMaker or equivalent capabilities, model lifecycle needs, and operational guardrails.

1 year • Operate the data infrastructure platform as a mature internal product with a clear service catalog, documented support model, paved roads, self-service capabilities, standardized CI/CD, and transparent reliability and cost reporting. • Deliver the highest-priority roadmap outcomes and demonstrate measurable year-over-year improvement in platform availability, incident recovery, deployment lead time, developer effort, operational toil, and cost efficiency. • Build a healthy, high-performing team with clear ownership, strong technical leadership, meaningful career growth, effective succession coverage, and the capability to execute both roadmap and operational work predictably. • Establish a durable multi-year strategy for Airflow, Flink, Spark and AWS EMR, Kafka, Snowflake, Iceberg, and ML infrastructure that anticipates

Tailor your resume for this SailPoint role before you apply.

Tailor my resume for this job

Similar jobs

Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)