ForgeApply · Job listing
Senior Site Reliability Engineer
CDW
See all 74 open roles at CDW →
Tailor your resume for this CDW job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for CDW's site. Free trial, no card required.
About this role
As a Senior Site Reliability Engineer (SRE) at CDW, you will improve the reliability, scalability, performance, security, and operational excellence of applications and platforms supporting Managed Services. This includes customer-connectivity platforms that enable CDW service delivery teams to monitor, troubleshoot, and manage customer environments. You will serve as a senior technical escalation point, resolving complex operational issues, reducing operational demands on development teams, and partnering with software engineering, infrastructure, security, and business stakeholders to improve service resilience. This role combines software engineering, systems engineering, automation, and operational leadership to reduce risk, improve the customer experience, and accelerate delivery. The Senior SRE will also mentor engineers, establish reliability engineering practices, and influence architectural and operational decisions across the organization. What You Will Do: • Drive service reliability, scalability, performance, security, and operational excellence across Managed Services applications and customer-connectivity platforms. • Serve as a senior escalation point for complex operational issues, resolving incidents that exceed the current SRE team’s expertise and reducing operational interruptions for development teams. • Design, develop, and maintain automation using Ansible and Python to reduce operational toil, improve consistency, and increase platform reliability. • Establish and mature reliability and observability practices, including SLIs, SLOs, Error Budgets, metrics, logs, traces, alerting, and dashboards. • Lead major incident response, Problem Management, root cause analysis, post-incident reviews, and corrective actions that address systemic issues and prevent recurrence. • Perform operational readiness and resiliency reviews while identifying reliability risks, operational gaps, technical debt, and opportunities for continuous improvement. • Troubleshoot complex issues across applications, Kubernetes environments, infrastructure, networking, identity services, databases, certificates, cloud services, and third-party integrations. • Mentor engineers and partner with development, infrastructure, security, and business teams to improve technical standards, operational practices, documentation, release quality, and production readiness. • Participate in a scheduled primary and secondary on-call rotation after completing training and demonstrating readiness to independently support the environment.
What We Expect of You: • Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field and 7+ years of experience in Software Engineering, Site Reliability Engineering, DevOps, Platform Engineering, or a related discipline; or 10+ years of equivalent professional experience. • 5+ years of experience administering Linux-based systems in enterprise environments. • 5+ years of experience developing automation solutions using Ansible. • 3+ years of experience developing automation and operational tooling using Python. • 3+ years of experience supporting Kubernetes or other container orchestration platforms in production environments. • Experience supporting business-critical production applications and distributed systems. • Experience leading major incident response, Problem Management, root cause analysis, and corrective action initiatives. • Experience working within ITIL-aligned Incident, Problem, and Change Management processes. • Experience implementing observability solutions using metrics, logs, traces, alerting, and dashboards. • Experience defining and measuring service reliability using SLIs, SLOs, and Error Budgets. • Strong knowledge of networking, CI/CD pipelines, source control, certificates, secrets management, and modern software delivery practices. • Demonstrated ability to troubleshoot complex technical issues, influence technical direction, mentor engineers, and communicate effectively with technical and non-technical stakeholders. • Experience reviewing, troubleshooting, and making minor enhancements to existing Java-based applications. This is not primarily a Java application-development role, is a plus. • Experience with Spring Boot, REST APIs, messaging technologies, and microservices architectures, is a plus. • Experience with OpenTelemetry, Dynatrace, Prometheus, Grafana, or similar observability platforms, is a plus. • Experience supporting Azure, AWS, or GCP cloud services and cloud-native architectures, is a plus. • Experience troubleshooting and supporting PostgresSQL, MySQL, MongoDB, DB2, IBMi (AS/400), or other enterprise database platforms, is a plus. • Experience implementing OAuth 2.0, JWT, RBAC, and secure application practices, is a plus. • Familiarity with Agile, Scrum, or SAFe delivery methodologies, is a plus. • Experience supporting enterprise-scale Managed Services environments, is a plus. • Kubernetes, cloud, security, or automation-related certifications, is a plus.
Pay range: $106,000 - $150,180 depending on experience and skill set Annual bonus target of 10% subject to terms and conditions of plan Benefits overview: https://cdw.benefit-info.com/ Salary ranges may be subject to geographic differentials
CDW is committed to being an AI-fluent organization We’re looking for people who bring curiosity, a learner’s mindset, and a willingness to engage with ever-evolving technology and tools. We value adopting AI as a partner, openness to experimentation, and a shared interest in learning together on AI. Our goal is to create a culture where AI enhances—not replaces—human creativity and decision-making. You don’t need to be an expert today; what matters is your readiness to explore, adapt, and grow with us as we integrate AI responsibly and effectively into our work. Additionally, CDW is committed to fostering an equitable, transparent, and respectful hiring process for all applicants. Du
Tailor your resume for this CDW role before you apply.
Tailor my resume for this jobSimilar jobs
- Senior Site Reliability Engineer — Datavant2 · Remote
- Senior Site Reliability Engineer — Autodesk · San Francisco, CA
- Senior Site Reliability Engineer — Rb · Boston, MA
- Senior Site Reliability Engineer — Juullabs · United States of America
- Senior Site Reliability Engineer — Securityscorecard · Remote
- Senior Site Reliability Engineer — Securityscorecard · Austin, TX (Hybrid)
- Senior Site Reliability Engineer — Cribl · Remote
- Senior Site Reliability Engineer — Vardaspace · El Segundo, California, United States
More like this: DevOps & SRE Jobs · Remote DevOps & SRE Jobs · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)