ForgeApply · Job listing
Data Site Reliability Engineer (SRE)
GDIT
See all 218 open roles at GDIT →
Tailor your resume for this GDIT job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for GDIT's site. Free trial, no card required.
About this role
Type of Requisition: Regular
Clearance Level Must Currently Possess: None
Clearance Level Must Be Able to Obtain: None
Public Trust/Other Required: BI Full 6C (T4)
Job Family: IT Infrastructure and Operations
Job Qualifications: Skills: CI/CD, Containerization, Structured Query Language (SQL) Development Certifications: None Experience: 5 + years of related experience US Citizenship Required: No
Job Description: Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program. The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States.
GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Data Site Reliability Engineer (SRE) will work as part of the CMM Data Modernization and Governance team responsible for delivering integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill the AO's data and analytics objectives in support of the CMM program.
The successful candidate will be responsible for providing technical leadership for the day-to-day operational support, reliability, performance, and continuous improvement of the CMM data platforms, pipelines, applications, and analytics services. This role ensures that data services remain secure, available, reliable, and aligned with established service levels, data governance standards, architecture principles, and operational procedures.
THE DATA SITE RELIABILITY ENGINEER (SRE) WILL EXECUTE THE FOLLOWING RESPONSIBILITIES • Provide comprehensive real-time monitoring, incident and event management, capacity planning, and operational reporting to support application deployments, maintain system health, predict demand, and align cloud operations with evolving business and security objectives.
• Maintain and audit user roles and responsibilities in cloud environments.
• Integrate Single Sign On (SSO), Multi-Factor Authentication (MFA) and group identity management managed through the Judiciary Enterprise Network Information Exchange (JENIE) for enforcing least privilege access.
• Adhere to guidelines prescribed by the Government and continuously assess and improve credential management processes for all user credentials.
• Provide Disaster Recovery (DR) and Continuity of Operations (COOP) options. This must include high-availability options, including fault-tolerant and automated failover designs.
• Integrate DevSecOps tools and processes seamlessly with enterprise systems (Integrated Development Environments (IDEs), ticketing, monitoring, etc.) to avoid fragmentation and ensure unified security posture.
• Provide and manage a centralized secrets management system with automated rotation, access logging, and policy enforcement to securely store, manage, and control access to sensitive information and to prevent unauthorized access and data breaches for any administrative user account.
• Integrate security tools (example: SAST, DAST, SCA, CSPM) into pipelines for continuous assessment and remediation.
• Implement unified, automated, continuous monitoring (24/7/365) systems and tools for security, performance, and compliance across all environments, leveraging dashboards and alerting for real-time visibility. Provide supplemental monitoring of event response activities beyond normal business hours (7a.m – 6p.m Eastern Time). Systems and tools shall capture data without including a required response to alerts.
• Ensure automated generation and management of Software Bill of Materials (SBOM) for all deployed artifacts, supporting transparency and compliance.
• Provide diagnostics, metrics’ gathering, and performance tuning services.
• Provide canary release function for end-user testing to support beta testing.
• Configure an alert mechanism so that the support teams can react in an instance of unusual behavior.
• Implement and operate a comprehensive incident and event management process, including integration with enterprise SIEM solutions, automated alerting, escalation workflows, and root cause analysis for all critical incidents.
• Provide engineering support to ensure prompt detection, logging, diagnosis, escalation, and resolution of incidents to restore normal service operations as quickly as possible and minimize impact.
• Perform systems support in identifying, analyzing, and eliminating the root causes of recurring incidents to minimize continued adverse impacts and potential degradation of services.
• Make recommendations for the improvement of Incident and Problem management consistent with industry’s best practices for the cloud.
• Maintain knowledge base of known issues, resolutions, and best practices for operational continuity.
• Perform automated health checks across the full stack (Operating System, Application, Database and PaaS services) at agreed levels on an agreed frequency.
• Provide a monthly issues management report. The report shall include cloud-related incidents, any stability and performance issues, configurations issues, quantity of tickets received, and time duration to resolve tickets.
• Develop and implement thresholds, rules, and response procedures based on product team’s recommendation.
• Monitor resource utilization (e.g., CPU, Memory, Disk Space) for the cloud hosted Virtual Machines (VMs) and other cloud services.
• Manage the resolution procedures for any threshold breaches for cloud resources.
• Improves system reliability, observability, automation, scalability, and operational resilience through engineering practices.
• Monitors, maintains, and optimizes cloud infrastructure, databases, and platform services for reliability and performance.
• Act as FinOps Analyst and perform cost optimization.
QUALIFICA
Tailor your resume for this GDIT role before you apply.
Tailor my resume for this jobSimilar jobs
- Senior Site Reliability Engineer (SRE) — Prizepicks · Remote
- Software Engineer, Site Reliability (SRE) — Sierra · San Francisco, CA
- Sr Mgr, Site Reliability Engineer (SRE) — Disney · Orlando, FL
- Infrastructure Site Reliability Engineer — Nebius · United States
- Director Site Reliability Engineering — Websteronline · CT Southington
- Engineer - Site Reliability Engineering — LSEG · USA-St. Louis-795 Office Pkwy
- Sr Site Reliability Engineer — Rdccareers · Austin, Texas, United States
- Site Reliability Engineer — Cisco · RTP, North Carolina, United States
More like this: DevOps & SRE Jobs · Remote DevOps & SRE Jobs · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)