Senior GPU Systems Architect for AI Clusters

RPMGlobal

Bethesda (MD)

On-site

USD 140,000 - 190,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RPMGlobal seeks a senior systems engineer to design, deploy, and optimize GPU clusters for enterprise AI mission systems in a secure government environment.

The role emphasizes operating systems, hardware, NVIDIA platforms (DGX/HGX/H200/H100/L4s), and high-speed networking, with Linux configuration, automation, and performance analysis at the core. Active TS/SCI clearance is required.

Qualifications

  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with 10 years of experience in the field.
  • Experience managing NVIDIA GPU data center platforms (DGX, HGX, H200, H100, L4s).
  • Knowledge of enterprise server components, including storage/network controllers, HBAs, and SSDs.
  • Strong Linux distributions expertise (RHEL, Ubuntu, Oracle, Rocky).
  • Excellent problem-solving and teamwork capabilities.
  • Meets DoD 8570.11 IAT Level II certification; Level III acceptable.
  • U.S. citizenship is required.

Responsibilities

  • Design, configure, and maintain GPU clusters.
  • Collaborate with multidisciplinary teams to define architectures for performance and power efficiency.
  • Integrate GPUs with Linux-based systems for AI/ML workflows.
  • Optimize GPU drivers for reliability and performance.
  • Analyze GPU performance to identify bottlenecks and improve cross-layer efficiency.
  • Develop debugging tools, profiling utilities, and performance analysis software for Linux.
  • Leverage Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain architectural specifications and Linux best practices.
  • Support ATO activities and ensure compliance with federal security standards.

Skills

NVIDIA GPUs
Linux admin
Python automation
Shell scripting
Security clearance
Team collaboration
Performance optimization

Education

Bachelor's degree
Master's degree
PhD

Tools

Kubernetes
Argo
Airflow
Kubeflow
Prometheus
Grafana
Slurm
LSF

Job description

RPMGlobal seeks a senior systems engineer to design, deploy, and optimize GPU clusters for enterprise AI mission systems in a secure government environment.

The role emphasizes operating systems, hardware, NVIDIA platforms (DGX/HGX/H200/H100/L4s), and high-speed networking, with Linux configuration, automation, and performance analysis at the core. Active TS/SCI clearance is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer III — AI Clusters & Linux
GPU Systems Engineer III — AI Clusters & Linux

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 260,000
GPU Systems Engineer 4
GPU Systems Engineer 4

RPMGlobal • Bethesda (MD)

On-site
USD 140,000 - 190,000
GPU Systems Engineer 3
GPU Systems Engineer 3

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 260,000
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
HPC AI Systems Architect (On-Prem GPU Cluster)
HPC AI Systems Architect (On-Prem GPU Cluster)

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Senior GPU Cluster Engineer – Linux Systems, On-Site
Senior GPU Cluster Engineer – Linux Systems, On-Site

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
Senior GPU Systems Engineer - TS/SCI, On-Site Bethesda
Senior GPU Systems Engineer - TS/SCI, On-Site Bethesda

Base-2 Solutions • Bethesda (MD)

On-site
USD 160,000 - 220,000
401(k) with company match
Company-paid health premiums
PTO & holidays
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior GPU HPC Systems Engineer
Senior GPU HPC Systems Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits