Senior GPU Systems Engineer for Secure AI Clusters

RPMGlobal

Bethesda, Northern (MD, KY)

Hybrid

USD 180,000 - 280,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

RPMGlobal seeks an experienced engineer to design, implement, and optimize GPU clusters for enterprise AI mission systems. You will work with AI/ML teams to integrate NVIDIA GPUs into Linux environments and drive performance across hardware and software layers.

Responsibilities include debugging tooling, automating with Bash and Python, and ensuring DoD security compliance. A TS/SCI clearance and DoD 8570.11 IAT Level II certs are required.

Qualifications

  • Active TS/SCI clearance with ability to obtain CI Polygraph.
  • Bachelor's degree with at least 10 years in the field.
  • Experience managing NVIDIA GPU data center platforms (DGX, HGX, H200, H100, L4s).
  • Knowledge of enterprise server components, storage/ controllers, HBAs, SSDs.
  • Strong Linux expertise (RHEL, Ubuntu, Oracle, Rocky).
  • Proven problem-solving skills; team collaboration.
  • IAT Level II certification per DoD 8570.11 (Level III acceptable).
  • U.S. citizenship required for government contracts.

Responsibilities

  • Design, configure, and maintain GPU clusters.
  • Define architectures for performance and power efficiency.
  • Integrate GPUs with Linux-based systems with AI/ML engineers.
  • Optimize GPU drivers for compatibility and performance.
  • Analyze GPU performance and improve efficiency across layers.
  • Build and maintain debugging tools and profiling utilities for Linux.
  • Develop tooling with Bash, Python, Ansible, Puppet, Salt.
  • Maintain documentation and Linux best practices.
  • Support ATO activities and ensure compliance with security standards.

Skills

Active TS/SCI
NVIDIA GPU data center platforms
Linux proficiency
Problem solving
Team collaboration
IAT Level II certification

Education

Bachelor's degree with 10 years of experience
Master's degree with 10 years of experience
PhD with 10 years of experience

Tools

Kubernetes
Argo
Airflow
Kubeflow
Prometheus
Grafana
Slurm
LSF
RHEL
Ubuntu
Oracle Linux
Rocky Linux

Job description

RPMGlobal seeks an experienced engineer to design, implement, and optimize GPU clusters for enterprise AI mission systems. You will work with AI/ML teams to integrate NVIDIA GPUs into Linux environments and drive performance across hardware and software layers.

Responsibilities include debugging tooling, automating with Bash and Python, and ensuring DoD security compliance. A TS/SCI clearance and DoD 8570.11 IAT Level II certs are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Systems Engineer: Enterprise AI Clusters (TS/SCI)
Senior GPU Systems Engineer: Enterprise AI Clusters (TS/SCI)

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 180,000
GPU Systems Engineer 4
GPU Systems Engineer 4

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 180,000 - 280,000
GPU Systems Engineer 3
GPU Systems Engineer 3

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 180,000
Senior GPU Cluster Engineer – Linux Systems, On-Site
Senior GPU Cluster Engineer – Linux Systems, On-Site

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
HPC & GPU Systems Engineer - Onsite (TS/SCI)
HPC & GPU Systems Engineer - Onsite (TS/SCI)

Xcelerate Solutions • Bethesda (MD)

On-site
USD 120,000 - 170,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Graphics Processing Unit (GPU) Engineer - TS/SCI
Graphics Processing Unit (GPU) Engineer - TS/SCI

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
HPC & GPU Systems Engineer — Linux/Kubernetes (TS/SCI)
HPC & GPU Systems Engineer — Linux/Kubernetes (TS/SCI)

VMD Corp • Bethesda (MD)

On-site
USD 110,000 - 150,000