Senior Systems Engineer - HPC & GPU Infrastructure

RPMGlobal

Bethesda, Northern (MD, KY)

Hybrid

USD 140,000 - 200,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Base-2 Solutions seeks a Senior Systems Engineer for HPC and GPU infrastructure to design, deploy, and optimize Linux-based clusters supporting Intelligence Community customers. This on-site Bethesda, MD role emphasizes performance, security, and scalable architectures.

You will install, configure, and manage HPC components, Slurm/PBS schedulers, container tech, and cloud integrations, while ensuring DoD compliance and robust documentation.

Qualifications

  • Bachelor's degree in CS/EE or related field.
  • 12+ years of relevant systems engineering experience.
  • DoD 8570.11 IAT Level II (Level III acceptable) or equivalent.

Responsibilities

  • Install and maintain GPU/HPC hardware on-premises and in the cloud.
  • Analyze cluster performance and optimize across Linux environments.
  • Install and configure Slurm, PBS, and workload-management tools.
  • Develop power-management and efficiency strategies for GPUs.

Skills

Linux
HPC
GPU hardware
Python
Bash
Teamwork
Communication

Education

Bachelor's or higher in CS/EE/related

Tools

Docker
Kubernetes
Slurm
PBS
Ansible
Puppet
Salt
Terraform

Job description

Position Summary

Base-2 Solutions is seeking a Senior Systems Engineer - HPC & GPU Infrastructure to design, develop, and optimize high-performance computing and GPU clusters supporting Intelligence Community customers. This is a 100% on-site position at the customer site in Bethesda, Maryland.

Essential Duties and Responsibilities
  • Contribute to the installation and maintenance of GPU and HPC hardware on-premises and in the cloud, providing insight into hardware performance and interaction with software components.
  • Analyze HPC and GPU cluster performance, identify bottlenecks, and develop strategies to improve performance across applications in Linux, addressing hardware and software considerations.
  • Install and configure HPC and GPU job-scheduling and workload-management platforms, including Slurm, PBS, Apache Airflow, and Kubernetes.
  • Apply power-management techniques to optimize GPU power consumption on mobile and desktop Linux platforms and continuously assess power-efficiency strategies.
  • Design and execute Linux-based tests to validate GPU performance and functionality, including stress testing, benchmarking, and debugging; maintain and expand the testing suite.
  • Maintain comprehensive technical documentation, including architectural specifications, code documentation, and Linux-specific best practices for GPU development.
  • Monitor trends, innovations, and competitive developments in the GPU industry; contribute to research and propose Linux-specific approaches to GPU design and optimization.
Required Qualifications
  • Relevant systems engineering experience meeting an applicable education and experience pathway in the equivalency section.
  • Expertise in operating system integration for Linux.
  • Strong understanding of computer hardware architecture, particularly as it relates to Linux systems.
  • Knowledge of parallel computing, graphics algorithms, and real-time rendering in Linux environments.
  • Excellent problem-solving skills and ability to collaborate within a team.
  • Strong communication skills for conveying technical information in a Linux context.
  • Proficiency with scripting languages such as Python or BASH.
  • Proficiency with automation tools such as Ansible, Puppet, Salt, Terraform, and similar tools.
  • Candidate must, at a minimum, meet DoD 8570.11 IAT Level II certification requirements; an IAT Level III certification is also acceptable.
  • US Citizenship is required due to the nature of the government contracts supported.
Preferred Qualifications
  • Knowledge of GPU virtualization, cloud computing, and emerging Linux-based technologies.
  • Experience with container technologies, including Docker and Kubernetes.
  • Experience with Prometheus/Grafana for monitoring.
  • Knowledge of distributed resource scheduling systems.
  • Understanding of data center networking hardware and cabling concepts.
  • Understanding of networking technologies such as DHCP, DNS, TCP/IP, VLANs, HSRP, and SNMP.
  • Knowledge of data center networking security principles, including firewall ACLs, IPS/IDS, and policy-based routing.
Required Education and Experience Equivalency
  • Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field with 12+ years of relevant systems engineering experience.
  • Additional years of relevant systems engineering experience in lieu of a degree.
Required Certifications
  • Candidate must, at a minimum, meet DoD 8570.11 IAT Level II certification requirements, currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification. An IAT Level III certification would also be acceptable, including CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, or CCSP.
Required Security Clearance
  • Active Top Secret/SCI
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Systems Engineer (Mid-Career) - HPC & GPU Infrastructure
Systems Engineer (Mid-Career) - HPC & GPU Infrastructure

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 110,000 - 160,000
Senior HPC & GPU Systems Engineer - On‑Site (Top Secret)
Senior HPC & GPU Systems Engineer - On‑Site (Top Secret)

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 110,000 - 160,000
Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos Inc • Bethesda (MD)

On-site
USD 70,000 - 90,000
Senior Systems Engineer
Senior Systems Engineer

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior HPC & GPU Systems Engineer — On-Site Bethesda
Senior HPC & GPU Systems Engineer — On-Site Bethesda

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 140,000 - 200,000
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

FiveInsights • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems Engineer HPC&GPU Infrastructure - TS/SCI
Systems Engineer HPC&GPU Infrastructure - TS/SCI

Sunayu Llc • Bethesda (MD), Northern (KY)

Hybrid
USD 140,000 - 180,000
Medical Plan Options
Dental and Vision
Life/AD&D Insurance
+5
Systems Engineer - HPC & GPU Infrastructure - TS/SCI
Systems Engineer - HPC & GPU Infrastructure - TS/SCI

Xcelerate Solutions • Bethesda (MD)

On-site
USD 120,000 - 170,000
Platform Engineer
Platform Engineer

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 140,000 - 190,000