GPU Systems Engineer 4

Base-2 Solutions

Bethesda (MD)

On-site

USD 170,000 - 230,000

Full time

2 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

100% company-paid medical premiums
100% company-paid dental premiums
100% company-paid vision premiums
401(k) with 4% match
Life insurance up to $200,000

Job summary

Base-2 Solutions is seeking a Senior GPU Systems Engineer in Bethesda, MD to design, optimize, and operate secure GPU clusters in a Linux environment. You will work with AI/ML engineers and DoD standards to ensure peak performance, security, and reliability of enterprise-grade GPU systems.

The candidate will lead efforts on driver optimization, high-speed networking, and automation tooling, while maintaining thorough documentation and compliance with required certifications.

Qualifications

  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with a minimum of ten years of experience in the category field.
  • Experience managing NVIDIA GPU data center platforms, including DGX, HGX, H200, H100, and L4s.
  • Knowledge of enterprise server components, including storage/network controllers, HBAs, and SSDs.
  • Strong expertise with Linux distributions, including RHEL, Ubuntu, Oracle, and Rocky.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Meet DoD 8570.11 IAT Level II certification requirements at a minimum; IAT Level III is also acceptable.
  • U.S. citizenship is required due to the nature of the government contracts supported.

Responsibilities

  • Design, configure, and maintain GPU clusters.
  • Collaborate with a multidisciplinary team to define architectures for performance and efficiency.
  • Integrate GPUs with Linux-based systems with AI/ML engineers.
  • Optimize GPU drivers for reliability and performance.
  • Analyze GPU performance and identify bottlenecks across hardware and software.
  • Build and maintain debugging tools and profiling utilities for Linux.
  • Utilize Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain technical documentation and Linux best practices.
  • Support security compliance activities and federal standards.

Skills

Linux distributions
Python
Shell scripting
Team collaboration

Education

Bachelor's degree + 10 years experience
Master's degree + 10 years experience
PhD + 10 years experience

Tools

DGX platforms
HGX platforms
Kubernetes
Slurm
Prometheus
Grafana

Job description

  • Required Security Clearance: Top Secret/SCI
  • Location: Bethesda, MD
  • Work Type: On-Site
  • Shift: First
  • Requisition ID: 3690
  • Standard Title: Senior GPU Systems Engineer
  • Required Security Clearance: Top Secret/SCI
  • Location: Bethesda, MD
  • Work Type: On-Site
  • Shift: First
  • Referral Eligibility: Eligible
  • U.S. Citizenship Required? Yes
Position Summary

Support enterprise AI mission systems by designing, developing, and optimizing GPU clusters, with deep focus on operating systems, hardware, GPU platforms, and high-speed networking in a secure customer environment.

Essential Duties And Responsibilities
  • Design, configure, and maintain GPU clusters.
  • Collaborate with a multidisciplinary team to define and optimize architectures for performance, power efficiency, and required features.
  • Work closely with AI/ML engineers to integrate GPUs with Linux-based systems.
  • Optimize GPU drivers for compatibility, reliability, and performance.
  • Analyze GPU performance, identify bottlenecks, and develop strategies to improve efficiency across hardware and software layers.
  • Build and maintain debugging tools, profiling utilities, and performance analysis software for Linux environments.
  • Leverage Bash, Python, Ansible, Puppet, and Salt for tooling and automation.
  • Maintain technical documentation, architectural specifications, and Linux best practices.
  • Support ATO activities and ensure compliance with federal security standards.
Required Qualifications
  • Active TS/SCI with ability to obtain a CI Polygraph.
  • Bachelor's degree with a minimum of ten years of experience in the category field.
  • Experience managing NVIDIA GPU data center platforms, including DGX, HGX, H200, H100, and L4s.
  • Knowledge of enterprise server components, including storage/network controllers, HBAs, and SSDs.
  • Strong expertise with Linux distributions, including RHEL, Ubuntu, Oracle, and Rocky.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Meet DoD 8570.11 IAT Level II certification requirements at a minimum; IAT Level III is also acceptable.
  • U.S. citizenship is required due to the nature of the government contracts supported.
Preferred Qualifications
  • Experience with Kubernetes cluster management and AI/ML workflow orchestration, including Argo, Airflow, and Kubeflow.
  • Familiarity with GPU virtualization and cloud computing.
  • Experience with Prometheus and Grafana for monitoring.
  • Knowledge of distributed resource scheduling systems such as Slurm, LSF, or similar tools.
Required Education and Experience Equivalency
  • Bachelors' Degree with 10 years of experience.
  • Masters' Degree with 10 years of experience.
  • PhD with 10 years of experience.
Required Certifications
  • DoD 8570.11 IAT Level II certification: Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP.
Required Security Clearance
  • Active TS/SCI with ability to obtain a CI Polygraph.
Pay & Benefit Highlights
Compensation
  • Competitive fixed salary or hourly pay (based on experience, skills, location, and internal equity).
  • Employee referral bonuses up to $10,000 per hired referral.
  • Additional bonus opportunities for exceptional performance and contributions to business development and company growth (role-dependent).
Health
  • 100% company-paid medical premiums for employees and eligible dependents.
  • Choose from multiple plan options with CareFirst, Kaiser, and UnitedHealthcare, including PPO, POS, HMO, and HSA-compatible plans.
  • 100% company-paid dental premiums for employees and eligible dependents.
  • 100% company-paid vision premiums for employees and eligible dependents.
Income Protection
  • 100% company-paid premiums for short-term disability.
  • 100% company-paid premiums for long-term disability.
  • 100% company-paid premiums for accidental death & dismemberment (AD&D).
  • 100% company-paid premiums for life insurance up to $200,000.
Retirement
  • 401(k) with immediate vesting: 4% company match plus a 4% non-elective company contribution (8% total).
  • 401(k) pre-tax and Roth options.
Leave
  • Up to 20 days of flexible paid time off (PTO).
  • 11 paid floating holidays.
Work-Life Balance
  • Flexible work schedules, including flex time and compressed work periods (contract and project-dependent).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer 3 with Security Clearance
GPU Systems Engineer 3 with Security Clearance

Base-2 Solutions • Chevy Chase (MD)

On-site
USD 150,000 - 190,000
Company-paid health premiums
401(k) with match
Paid time off and holidays
GPU Systems Engineer 3
GPU Systems Engineer 3

Base-2 Solutions • Bethesda (MD)

On-site
USD 120,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+2
Graphics Processing Unit (GPU) Engineer - TS/SCI
Graphics Processing Unit (GPU) Engineer - TS/SCI

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
Senior Systems Engineer (Virtualization / GPU Infrastructure)
Senior Systems Engineer (Virtualization / GPU Infrastructure)

Endepth Solutions • Laurel (MD)

On-site
USD 195,000 - 255,000
Affordable healthcare options
401(k) with company match
Paid Time Off (PTO)
+3
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior Systems Engineer (Virtualization / GPU Infrastructure)
Senior Systems Engineer (Virtualization / GPU Infrastructure)

EnDepth Solutions, LLC • Laurel (MD)

On-site
USD 195,000 - 255,000
Affordable healthcare options
401(k) with company match
Annual training reimbursement
+1
Senior Full Stack Software Engineer - DGX Cloud
Senior Full Stack Software Engineer - DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Graphic Processing Unit (GPU) Engineer - TS/SCI
Graphic Processing Unit (GPU) Engineer - TS/SCI

VMD Corp • Bethesda (MD)

On-site
USD 160,000 - 210,000
Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos • Bethesda (MD)

On-site
USD 87,000 - 158,000
Health and Wellness programs
Income Protection
Paid Leave
+1
Senior Full Stack Software Engineer - DGX Cloud
Senior Full Stack Software Engineer - DGX Cloud

NVIDIA Gruppe • North Carolina

On-site
USD 224,000 - 357,000