GPU Systems Engineer 3

RPMGlobal

Bethesda (MD)

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

RPMGlobal in Bethesda, MD is seeking a systems engineer to design, deploy, and optimize enterprise GPU clusters for DoD-grade AI mission systems. You will work with Linux, NVIDIA GPUs, high-speed networking, and secure configurations to support mission-critical workloads.

The role requires TS/SCI with CI poly eligibility, strong problem solving, and the ability to collaborate across cross-disciplinary teams to deliver reliable, scalable infrastructure.

Qualifications

  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with 6+ years of experience, or equivalent.
  • Experience managing NVIDIA GPU data center platforms (DGX/HGX/H200/H100/L4s).
  • Knowledge of enterprise server components (storage/network controllers, HBAs, SSDs).
  • Strong Linux exp (RHEL, Ubuntu, Oracle, Rocky).
  • DoD 8570.11 IAT Level II certification or higher.
  • U.S. citizenship required.

Responsibilities

  • Design, configure, and maintain GPU clusters.
  • Collaborate to optimize architectures for performance and power efficiency.
  • Work with AI/ML engineers to integrate GPUs with Linux systems.
  • Optimize GPU drivers for reliability and performance.
  • Analyze performance, identify bottlenecks, and improve efficiency.
  • Build debugging tools, profiling utilities, and performance analyses for Linux.
  • Use Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain documentation and Linux best practices.
  • Support ATO activities and ensure DoD security compliance.

Skills

Team collaboration
Problem solving
Linux expertise

Education

Bachelor's Degree
High School Diploma/GED
Associates Degree
Master's Degree
PhD

Tools

NVIDIA DGX/HGX/H200/H100/L4s
Kubernetes
Prometheus
Grafana
Slurm
LSF

Job description

Position Summary

Support enterprise AI mission systems by designing, developing, and optimizing GPU clusters, with deep focus on operating systems, hardware, GPU platforms, and high-speed networking in a secure customer environment.

Essential Duties and Responsibilities
  • Design, configure, and maintain GPU clusters.
  • Collaborate with a multidisciplinary team to define and optimize architectures for performance, power efficiency, and required features.
  • Work closely with AI/ML engineers to integrate GPUs with Linux-based systems.
  • Optimize GPU drivers for compatibility, reliability, and performance.
  • Analyze GPU performance, identify bottlenecks, and develop strategies to improve efficiency across hardware and software layers.
  • Build and maintain debugging tools, profiling utilities, and performance analysis software for Linux environments.
  • Leverage Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain technical documentation, architectural specifications, and Linux best practices.
  • Support ATO activities and ensure compliance with federal security standards.
Required Qualifications
  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with a minimum of six years of experience in the category field. Three additional years of experience may be substituted for the bachelor's degree.
  • Experience managing NVIDIA GPU data center platforms, including DGX, HGX, H200, H100, and L4s.
  • Knowledge of enterprise server components, including storage/network controllers, HBAs, and SSDs.
  • Strong expertise with Linux distributions, including RHEL, Ubuntu, Oracle, and Rocky.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Meet DoD 8570.11 IAT Level II certification requirements at a minimum; IAT Level III is also acceptable.
  • U.S. citizenship is required due to the nature of the government contracts supported.
Preferred Qualifications
  • Experience with Kubernetes cluster management and AI/ML workflow orchestration, including Argo, Airflow, and Kubeflow.
  • Familiarity with GPU virtualization and cloud computing.
  • Experience with Prometheus and Grafana for monitoring.
  • Knowledge of distributed resource scheduling systems such as Slurm, LSF, or similar tools.
Required Education and Experience Equivalency
  • High School Diploma/GED with 9 years of experience.
  • Associates Degree with 9 years of experience.
  • Bachelors' Degree with 6 years of experience.
  • Masters' Degree with 6 years of experience.
  • PhD with 6 years of experience.
RequiredCertifications
  • DoD 8570.11 IAT Level II certification: Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP.
Required Security Clearance
  • Active TS/SCI with ability to obtain a CI Polygraph.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer 4
GPU Systems Engineer 4

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 180,000 - 280,000
Graphics Processing Unit (GPU) Engineer - TS/SCI
Graphics Processing Unit (GPU) Engineer - TS/SCI

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
Senior GPU Systems Engineer: Enterprise AI Clusters (TS/SCI)
Senior GPU Systems Engineer: Enterprise AI Clusters (TS/SCI)

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 180,000
Senior GPU Systems Engineer for Secure AI Clusters
Senior GPU Systems Engineer for Secure AI Clusters

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 180,000 - 280,000
Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos Inc • Bethesda (MD)

On-site
USD 70,000 - 90,000
GPU Systems Infrastructure Engineer
GPU Systems Infrastructure Engineer

Blue Signal Search • Fremont (CA)

On-site
USD 120,000 - 170,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior System Engineer – GPU Platforms
Senior System Engineer – GPU Platforms

Jobtailor • San Jose (CA)

On-site
USD 150,000 - 210,000
Data Center Technician L2
Data Center Technician L2

Covestic Inc • Reno (NV)

On-site
USD 90,000 - 130,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Calance • Costa Mesa (CA)

Hybrid
USD 180,000 - 240,000