GPU Systems Engineer 4

RPMGlobal

Bethesda (MD)

On-site

USD 140,000 - 190,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RPMGlobal seeks a senior systems engineer to design, deploy, and optimize GPU clusters for enterprise AI mission systems in a secure government environment.

The role emphasizes operating systems, hardware, NVIDIA platforms (DGX/HGX/H200/H100/L4s), and high-speed networking, with Linux configuration, automation, and performance analysis at the core. Active TS/SCI clearance is required.

Qualifications

  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with 10 years of experience in the field.
  • Experience managing NVIDIA GPU data center platforms (DGX, HGX, H200, H100, L4s).
  • Knowledge of enterprise server components, including storage/network controllers, HBAs, and SSDs.
  • Strong Linux distributions expertise (RHEL, Ubuntu, Oracle, Rocky).
  • Excellent problem-solving and teamwork capabilities.
  • Meets DoD 8570.11 IAT Level II certification; Level III acceptable.
  • U.S. citizenship is required.

Responsibilities

  • Design, configure, and maintain GPU clusters.
  • Collaborate with multidisciplinary teams to define architectures for performance and power efficiency.
  • Integrate GPUs with Linux-based systems for AI/ML workflows.
  • Optimize GPU drivers for reliability and performance.
  • Analyze GPU performance to identify bottlenecks and improve cross-layer efficiency.
  • Develop debugging tools, profiling utilities, and performance analysis software for Linux.
  • Leverage Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain architectural specifications and Linux best practices.
  • Support ATO activities and ensure compliance with federal security standards.

Skills

NVIDIA GPUs
Linux admin
Python automation
Shell scripting
Security clearance
Team collaboration
Performance optimization

Education

Bachelor's degree
Master's degree
PhD

Tools

Kubernetes
Argo
Airflow
Kubeflow
Prometheus
Grafana
Slurm
LSF

Job description

Position Summary

Support enterprise AI mission systems by designing, developing, and optimizing GPU clusters, with deep focus on operating systems, hardware, GPU platforms, and high-speed networking in a secure customer environment.

Essential Duties and Responsibilities
  • Design, configure, and maintain GPU clusters.
  • Collaborate with a multidisciplinary team to define and optimize architectures for performance, power efficiency, and required features.
  • Work closely with AI/ML engineers to integrate GPUs with Linux-based systems.
  • Optimize GPU drivers for compatibility, reliability, and performance.
  • Analyze GPU performance, identify bottlenecks, and develop strategies to improve efficiency across hardware and software layers.
  • Build and maintain debugging tools, profiling utilities, and performance analysis software for Linux environments.
  • Leverage Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain technical documentation, architectural specifications, and Linux best practices.
  • Support ATO activities and ensure compliance with federal security standards.
Required Qualifications
  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with a minimum of ten years of experience in the category field.
  • Experience managing NVIDIA GPU data center platforms, including DGX, HGX, H200, H100, and L4s.
  • Knowledge of enterprise server components, including storage/network controllers, HBAs, and SSDs.
  • Strong expertise with Linux distributions, including RHEL, Ubuntu, Oracle, and Rocky.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Meet DoD 8570.11 IAT Level II certification requirements at a minimum; IAT Level III is also acceptable.
  • U.S. citizenship is required due to the nature of the government contracts supported.
Preferred Qualifications
  • Experience with Kubernetes cluster management and AI/ML workflow orchestration, including Argo, Airflow, and Kubeflow.
  • Familiarity with GPU virtualization and cloud computing.
  • Experience with Prometheus and Grafana for monitoring.
  • Knowledge of distributed resource scheduling systems such as Slurm, LSF, or similar tools.
Required Education and Experience Equivalency
  • Bachelors' Degree with 10 years of experience.
  • Masters' Degree with 10 years of experience.
  • PhD with 10 years of experience.
RequiredCertifications
  • DoD 8570.11 IAT Level II certification: Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP.
Required Security Clearance
  • Active TS/SCI with ability to obtain a CI Polygraph.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer 3
GPU Systems Engineer 3

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 260,000
GPU Systems Engineer 4
GPU Systems Engineer 4

Base-2 Solutions • Bethesda (MD)

On-site
USD 160,000 - 220,000
401(k) with company match
Company-paid health premiums
PTO & holidays
Graphics Processing Unit (GPU) Engineer - TS/SCI
Graphics Processing Unit (GPU) Engineer - TS/SCI

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
GPU Systems Engineer 3
GPU Systems Engineer 3

Base-2 Solutions • Bethesda (MD)

On-site
USD 130,000 - 170,000
Medical premiums covered
Dental premiums covered
Vision premiums covered
+2
Senior GPU Systems Architect for AI Clusters
Senior GPU Systems Architect for AI Clusters

RPMGlobal • Bethesda (MD)

On-site
USD 140,000 - 190,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior System Engineer – GPU Platforms
Senior System Engineer – GPU Platforms

Jobtailor • San Jose (CA)

On-site
USD 150,000 - 210,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Operations & Infrastructure Engineer
AI Operations & Infrastructure Engineer

Invictus International Consulting, LLC • Fort Meade (MD)

On-site
USD 100,000 - 130,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000