SME Platform Engineer

General Dynamics Information Technology

Arlington (VA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Iron EagleX is seeking a SME Platform Engineer in Crystal City, VA to lead on‑premises cloud infrastructure design, deployment, and management. You will orchestrate containerized apps, manage robust GPU and Ray compute environments, and enable AI/ML platforms while maintaining strict security and compliance.

The role emphasizes Kubernetes, RKE2, GitOps, Keycloak/OPA identity, and secure registries. You will collaborate with engineering and data science teams to enable scalable, secure data

Qualifications

  • Demonstrated experience designing, deploying, administering, and troubleshooting production Kubernetes environments; hands-on experience with RKE2 or similar.
  • Strong Linux systems administration skills, including diagnosing and troubleshooting infrastructure via CLI.
  • Experience implementing authentication and authorization using Keycloak, OPA, OIDC, RBAC.
  • Hands-on experience with Git-based workflows, Helm charts, CI/CD, and GitOps practices.
  • Experience with Argo CD or similar GitOps tools for declarative delivery.
  • Experience managing Kubernetes storage including Longhorn or similar.
  • Experience hosting data science/analytics platforms such as Posit Workbench/Connect, Hive Metastore.
  • Experience configuring systems to meet FIPS or other security/compliance requirements.
  • Demonstrated experience identifying and remediating critical vulnerabilities.
  • Experience administering GPU-enabled compute for AI/ML, HPC, or similar workloads.
  • Experience supporting distributed AI/ML workloads with Ray or similar frameworks.
  • Experience deploying large language model inference/serving technologies; vLLM routing preferred.
  • Experience securing container images from Harbor, Docker Hub, or Container Yard.

Responsibilities

  • Infrastructure & Orchestration: Architect, deploy, and manage on-premises cloud infrastructure using RKE2 and Longhorn/Object storage.
  • Platform Enablement: Host and maintain robust data science environments (Posit Workbench/Connect, Hive Metastore).
  • AI/ML Infrastructure: Manage GPU clusters, Ray clusters, and VLLM Routers for model training and inference.
  • CI/CD & Automation: Build and optimize pipelines with Git, Helm, and ArgoCD for reliable delivery.
  • Security & Compliance: Maintain FIPS compliance and mitigate vulnerabilities.
  • Identity & Access: Implement authentication/authorization with Keycloak and OPA.
  • System Administration: Manage container images from Harbor, Docker Hub, or Containeryard; troubleshoot via CLI.

Skills

Kubernetes
Linux
GitOps
Keycloak
OPA
ArgoCD
Terraform/Ansible
Python/Go/Bash
GPU/HPC
Security & Compliance

Tools

RKE2
Longhorn
Posit Workbench/Connect
Hive Metastore
Harbor/Docker Hub/Containeryard
Harbor
cGPU/Ray Clusters

Job description

Type of Requisition: Regular Clearance Level Must Currently Possess: Top Secret/SCI Clearance Level Must Be Able to Obtain: Top Secret SCI + Polygraph Public Trust/Other Required: None Job Family: IT Infrastructure and Operations Job Qualifications: Skills: CI/CD, Cloud Infrastructure, Cluster Administration, Kubernetes, Linux Server Administration Certifications: None Experience: 15 + years of related experience US Citizenship Required: Yes Job Description:

YOUR IMPACT

Own your opportunity to work with the largest government agency in the nation. Make an impact by advancing the Department of War’s mission to keep our country safe and secure.

OUR COMPANY Iron EagleX (IEX), a wholly owned subsidiary of General Dynamics Information Technology (GDIT), delivers agile IT and Intelligence solutions. Combining small-team flexibility with global scale, IEX leverages emerging technologies to provide innovative, user-focused solutions that empower organizations and end users to operate smarter, faster, and more securely in dynamic environments.

JOB DESCRIPTION

Iron EagleX is seeking a SME Platform Engineer to support our Engineering team in Crystal City, VA. This role will lead the design, implementation, and management of our secure, on-premises cloud infrastructure. In this role, you will be the driving force behind our advanced computing environments, ensuring the seamless orchestration of containerized applications and large-scale data science platforms. You will work at the intersection of infrastructure, security, and machine learning, managing robust compute clusters and providing foundational support for AI model training and deployment. The ideal candidate has deep expertise in Kubernetes ecosystem tools, GitOps methodologies, and strict compliance standards.

MEANINGFUL WORK AND PERSONAL IMPACT

As a SME Platform Engineer, your work will directly empower our data science and engineering teams to push the boundaries of machine learning and data analytics. By building and maintaining resilient GPU and Ray clusters, you will accelerate the fine-tuning and deployment of advanced models. Your commitment to security and compliance will ensure our critical systems remain protected against vulnerabilities, providing a safe, compliant, and highly performant foundation for the organization’s most impactful technical initiatives. You will not just be managing infrastructure; you will be enabling innovation.

JOB DUTIES (INCLUDE BUT ARE NOT LIMITED TO)
  • Infrastructure & Orchestration: Architect, deploy, and manage on-premises cloud infrastructure using RKE2 and maintain storage solutions like Longhorn and Object storage.
  • Platform Enablement: Host and maintain robust data science environments, including software such as POSIT Workbench/Connect and Hive Metastore.
  • AI/ML Infrastructure: Manage and scale robust GPU clusters, Ray Clusters for fine-tuning machine learning models, and VLLM Routers for efficient model inference.
  • CI/CD & Automation: Build, maintain, and optimize CI/CD pipelines using Git, Helm charts, and ArgoCD for reliable software delivery.
  • Security & Compliance: Ensure continuous FIPS compliance across the environment. Actively manage and mitigate critical and high-level vulnerabilities.
  • Identity & Access: Implement and maintain robust authentication and authorization mechanisms using Keycloak and Open Policy Agent (OPA).
  • System Administration: Pull and manage container images from secure registries such as Harbor, Docker Hub, or Containeryard. Manage all core capabilities and troubleshoot issues effectively via the command-line console.
REQUIRED SKILLS
  • Demonstrated experience designing, deploying, administering, and troubleshooting production Kubernetes environments; hands-on experience with RKE2 or similar.
  • Strong Linux systems administration skills, including the ability to manage, diagnose, and troubleshoot infrastructure and platform services through the command line (CLI).
  • Experience implementing authentication and authorization solutions using technologies such as Keycloak, Open Policy Agent (OPA), OIDC, RBAC, or comparable identity and access management frameworks.
  • Hands-on experience with Git-based development and deployment workflows, including Helm charts, CI/CD pipelines, and GitOps practices.
  • Experience with Argo CD or similar tools for declarative, GitOps-based continuous delivery.
  • Experience managing Kubernetes storage solutions, including distributed block storage and object storage; experience with Longhorn or comparable technologies preferred.
  • Experience hosting and administering data science or analytics platforms; experience with Posit Workbench, Posit Connect, Hive Metastore, or similar technologies.
  • Experience configuring and operating systems in accordance with FIPS or comparable security and compliance requirements.
  • Demonstrated experience identifying, prioritizing, and remediating critical and high-severity system and application vulnerabilities.
  • Experience administering GPU-enabled compute environments supporting AI/ML, high-performance computing, or other compute-intensive workloads.
  • Experience supporting distributed AI/ML workloads using Ray or comparable distributed computing frameworks, including model training and fine-tuning use cases.
  • Experience deploying or supporting large language model inference and serving technologies; experience with vLLM and related routing capabilities preferred.
  • Experience pulling, managing, securing, and troubleshooting container images using private or public registries such as Harbor, Docker Hub, Container Yard, or equivalent container registry platforms.
DESIRED SKILLS
  • Hands-on experience with Infrastructure as Code (IaC) and configuration management tools such as Terraform, Ansible, or comparable technologies.
  • Proficiency in scripting or programming languages such as Python, Go, or Bash to support infrastructure automation, platform operations, and troubleshooting.
  • Advanced knowledge of Linux system administration, networking concepts, and protocols within complex or highly available infrastructure environments.
  • Experience implementing and maintaining monitoring, logging, and observability solutions using tools such as Prometheus, Grafana, or comparable platforms.
  • Familiarity with MLOps practices and the machine learning lifecycle, including model development, deployment, monitoring, versioning, and operational support.
  • Experience automating infrastructure provisioning, configuration, deployment, and operational workflows in secure or regulated environments.
  • Familiarity with performance tuning, capacity planning, and resource optimization for Kubernetes, GPU, or other compute-intensive environments.
WHAT YOU’LL NEED
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer Senior Principal
Platform Engineer Senior Principal

General Dynamics - IT • Arlington (VA)

On-site
USD 180,000 - 260,000
Platform Engineer Senior Principal
Platform Engineer Senior Principal

General Dynamics Information Technology • Arlington (VA)

Hybrid
USD 185,000 - 251,000
401(k) with company match
Paid time off
Life insurance
Platform Engineer Senior Principal
Platform Engineer Senior Principal

Iron EagleX, Inc. • Crystal City (TX)

Hybrid
USD 185,000 - 251,000
Platform Operations Engineer - TS/SCI with Security Clearance
Platform Operations Engineer - TS/SCI with Security Clearance

SUNAYU • Chevy Chase (MD)

Hybrid
USD 120,000 - 170,000
Software DevOps Engineer - JobID-0263
Software DevOps Engineer - JobID-0263

Innovative Defense Technologies • Mount Laurel Township (NJ)

Hybrid
USD 114,000 - 231,000
Hybrid work model
On-site Mount Laurel office
DevOps Engineer II
DevOps Engineer II

Torch.AI • McLean (VA), Leawood (KS)

On-site
USD 140,000 - 190,000
Competitive base salary
Quarterly bonuses
Equity participation
Sr Software Development Engineer, SRE (US Federal)
Sr Software Development Engineer, SRE (US Federal)

Workday, Inc. • Reston (VA)

On-site
USD 140,000 - 170,000
Principal Agentic AI Full Stack Software Engineer
Principal Agentic AI Full Stack Software Engineer

General Dynamics Information Technology • Arlington (VA)

On-site
USD 153,000 - 207,000
Dev Ops System Administrator ,Scottsdale, Arizona, US, Full Time/Per (MK)
Dev Ops System Administrator ,Scottsdale, Arizona, US, Full Time/Per (MK)

Central Business Solutions, Inc • Scottsdale (AZ)

On-site
USD 126,000 - 138,000
9/80 option
Generous PTO
401(k) with company match
+1
Senior Principal Platform Engineer
Senior Principal Platform Engineer

Clarity Innovations • Jessup (MD)

On-site
USD 130,000 - 150,000