ForgeApply
Try it free

ForgeApply · Job listing

SME Platform Engineer

GDIT

USA VA Arlington, US$191k – $259konsite

See all 236 open roles at GDIT

Tailor your resume for this GDIT job in about a minute.

ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for GDIT's site. Free trial, no card required.

About this role

Type of Requisition: Regular

Clearance Level Must Currently Possess: Top Secret/SCI

Clearance Level Must Be Able to Obtain: Top Secret SCI + Polygraph

Public Trust/Other Required: None

Job Family: IT Infrastructure and Operations

Job Qualifications: Skills: CI/CD, Cloud Infrastructure, Cluster Administration, Kubernetes, Linux Server Administration Certifications: None Experience: 15 + years of related experience US Citizenship Required: Yes

Job Description: YOUR IMPACT  

Own your opportunity to work with the largest government agency in the nation. Make an impact by advancing the Department of   War ’s mission to keep our country safe and secure.  

OUR COMPANY  

Iron   EagleX   (IEX), a wholly owned subsidiary of General Dynamics Information Technology   (GDIT) , delivers agile IT and Intelligence solutions. Combining small-team flexibility with global scale, IEX leverages emerging technologies to provide innovative, user-focused solutions that empower organizations and end users to   operate   smarter, faster, and more securely in dynamic environments.   

JOB DESCRIPTION  

Iron   EagleX   is   seeking   a   SME   Platform   Engineer   to   support our   Engineering   team in Crystal City, VA. This role will   lead the design, implementation, and management of our secure, on-premises cloud infrastructure. In this role, you will be the driving force behind our advanced computing environments, ensuring the seamless orchestration of containerized applications and large-scale data science platforms. You will work at the intersection of infrastructure, security, and machine learning, managing robust compute clusters and providing foundational support for AI model training and deployment. The ideal candidate has deep   expertise   in Kubernetes ecosystem tools,   GitOps   methodologies, and strict compliance standards.  

M EANINGFUL WORK AND PERSONAL IMPACT  

As a   SME   Platform   Engineer ,   your work   will directly empower our data science and engineering teams to push the boundaries of machine learning and data analytics. By building a nd   maintain ing   resilient GPU and Ray clusters, you will accelerate the fine-tuning and deployment of advanced models. Your commitment to security and compliance will ensure our critical syste ms   rem ain   protected against vulnerabilities, providing a safe, compliant, and highly performant foundation for the organization's most impactful technical initiatives. You will not just be managing infrastructure; you will be enabling innovation.  

JOB DUTIES   (INCLUDE BUT ARE NOT LIMITED TO)  

• Infrastructure & Orchestration: Architect, deploy, and manage on-premises cloud infrastructure using RKE2 and   maintain   storage solutions like Longhorn and Object   storage .  

• Platform Enablement: Host and   maintain   robust data science environments, including software such as POSIT Workbench/Connect and Hive   Metastore .  

• AI/ML Infrastructure: Manage and scale robust GPU clusters, Ray Clusters for fine-tuning machine learning models, and VLLM Routers for efficient model inference.  

• CI/CD & Automation: Build,   maintain , and   optimize   CI/CD pipelines using Git, Helm charts, and   ArgoCD   for reliable software delivery.  

• Security & Compliance: Ensure continuous FIPS compliance across the environment. Actively manage and mitigate critical and high-level vulnerabilities.  

• Identity & Access: Implement and   maintain   robust authentication and authorization mechanisms using   Keycloak   and Open Policy Agent (OPA).  

• System Administration: Pull and manage container images from secure registries such as Harbor,   Docker   Hub , or   Containeryard . Manage all core capabilities and troubleshoot issues effectively via the command-line console.  

REQUIRED SKILLS  

• Demonstrated experience designing, deploying, administering, and troubleshooting production Kubernetes environments; hands-on experience with RKE2 or similar.  

• Strong Linux systems administration skills, including the ability to manage, diagnose, and troubleshoot infrastructure and platform services through the command line (CLI).  

• Experience implementing authentication and authorization solutions using technologies such as   Keycloak , Open Policy Agent (OPA), OIDC, RBAC, or comparable identity and access management frameworks.  

• Hands-on experience with Git-based development and deployment workflows, including Helm charts, CI/CD pipelines, and   GitOps   practices.  

• Experience with Argo CD or similar tools for declarative,   GitOps -based continuous delivery.  

• Experience managing Kubernetes storage solutions, including distributed block storage and object storage; experience with Longhorn or comparable technologies preferred.  

• Experience hosting and administering data science or analytics platforms; experience with Posit Workbench, Posit Connect, Hive   Metastore , or similar technologies.  

• Experience configuring and operating systems   in accordance with   FIPS or comparable security and compliance requirements.  

• Demonstrated experience   identifying , prioritizing, and remediating critical and high-severity system and application vulnerabilities.  

• Experience administering GPU-enabled compute environments supporting AI/ML, high-performance computing, or other compute-intensive workloads.  

• Experience supporting distributed AI/ML workloads using Ray or comparable distributed computing frameworks, including model training and fine-tuning use cases.  

• Experience deploying or supporting large language model inference and serving technologies; experience with   vLLM   and related routing capabilities preferred.  

• Experience pulling, managing, securing, and troubleshooting container images using private or public registries such as Harbor, Docker Hub, Container Yard, or equivalent container registry platforms.  

DESIRED SKILLS  

• Hands-on experience with Infrastructure as Code ( IaC ) and configura

Tailor your resume for this GDIT role before you apply.

Tailor my resume for this job

Similar jobs

More like this: DevOps & SRE Jobs · Browse all jobs

Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)