Senior GPU Systems Engineer: Enterprise AI Clusters (TS/SCI)

RPMGlobal

Bethesda (MD)

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

RPMGlobal in Bethesda, MD is seeking a systems engineer to design, deploy, and optimize enterprise GPU clusters for DoD-grade AI mission systems. You will work with Linux, NVIDIA GPUs, high-speed networking, and secure configurations to support mission-critical workloads.

The role requires TS/SCI with CI poly eligibility, strong problem solving, and the ability to collaborate across cross-disciplinary teams to deliver reliable, scalable infrastructure.

Qualifications

  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with 6+ years of experience, or equivalent.
  • Experience managing NVIDIA GPU data center platforms (DGX/HGX/H200/H100/L4s).
  • Knowledge of enterprise server components (storage/network controllers, HBAs, SSDs).
  • Strong Linux exp (RHEL, Ubuntu, Oracle, Rocky).
  • DoD 8570.11 IAT Level II certification or higher.
  • U.S. citizenship required.

Responsibilities

  • Design, configure, and maintain GPU clusters.
  • Collaborate to optimize architectures for performance and power efficiency.
  • Work with AI/ML engineers to integrate GPUs with Linux systems.
  • Optimize GPU drivers for reliability and performance.
  • Analyze performance, identify bottlenecks, and improve efficiency.
  • Build debugging tools, profiling utilities, and performance analyses for Linux.
  • Use Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain documentation and Linux best practices.
  • Support ATO activities and ensure DoD security compliance.

Skills

Team collaboration
Problem solving
Linux expertise

Education

Bachelor's Degree
High School Diploma/GED
Associates Degree
Master's Degree
PhD

Tools

NVIDIA DGX/HGX/H200/H100/L4s
Kubernetes
Prometheus
Grafana
Slurm
LSF

Job description

RPMGlobal in Bethesda, MD is seeking a systems engineer to design, deploy, and optimize enterprise GPU clusters for DoD-grade AI mission systems. You will work with Linux, NVIDIA GPUs, high-speed networking, and secure configurations to support mission-critical workloads.

The role requires TS/SCI with CI poly eligibility, strong problem solving, and the ability to collaborate across cross-disciplinary teams to deliver reliable, scalable infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Systems Engineer for Secure AI Clusters
Senior GPU Systems Engineer for Secure AI Clusters

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 180,000 - 280,000
GPU Systems Engineer 4
GPU Systems Engineer 4

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior GPU Cluster Engineer – Linux Systems, On-Site
Senior GPU Cluster Engineer – Linux Systems, On-Site

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
GPU Systems Engineer 3
GPU Systems Engineer 3

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 180,000
HPC & GPU Systems Engineer — Linux/Kubernetes (TS/SCI)
HPC & GPU Systems Engineer — Linux/Kubernetes (TS/SCI)

VMD Corp • Bethesda (MD)

On-site
USD 110,000 - 150,000
Graphics Processing Unit (GPU) Engineer - TS/SCI
Graphics Processing Unit (GPU) Engineer - TS/SCI

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
Senior HPC & GPU Systems Engineer
Senior HPC & GPU Systems Engineer

Xcelerate-Solutions-5 • Bethesda (MD)

On-site
USD 120,000 - 160,000
HPC Engineer (TS/SCI) – GPU, Clusters & Slurm
HPC Engineer (TS/SCI) – GPU, Clusters & Slurm

VT Group (VTG) • McLean (VA)

On-site
USD 130,000 - 190,000
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Systems Engineer - TS/SCI
Systems Engineer - TS/SCI

Xcelerate-Solutions-5 • Bethesda (MD)

On-site
USD 120,000 - 160,000