Systems Engineer (Mid-Career) - HPC & GPU Infrastructure

RPMGlobal

Bethesda, Northern (MD, KY)

Hybrid

USD 110,000 - 160,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Base-2 Solutions seeks a mid-career Systems Engineer specialized in HPC and GPU infrastructure to design, develop, and optimize GPU clusters for Intelligence Community customers. This is a 100% on-site role at the customer site in Bethesda, Maryland, requiring DoD 8570 IAT Level II/III or equivalent certification and US Citizenship.

The engineer will install and maintain on-premises and cloud-based GPU and HPC hardware, optimize performance, and implement workload management with Slurm, PBS,

Qualifications

  • Bachelor's degree or higher in Computer Science, Electrical Engineering or a related field.
  • 4+ years of systems engineering experience with Linux environments.
  • US Citizenship and security clearance eligibility.

Responsibilities

  • Contribute to installation and maintenance of on-premises and cloud GPU/HPC hardware.
  • Analyze HPC/GPU performance and identify bottlenecks for Linux applications.
  • Install and configure workload-management platforms (Slurm, PBS, Airflow, Kubernetes).
  • Apply power-management techniques to optimize GPU consumption.
  • Design and execute Linux-based GPU tests including benchmarking and debugging.
  • Maintain technical documentation and architectural specifications.
  • Monitor GPU industry trends and propose Linux-specific optimization approaches.
  • Share regular technical updates with the team.

Skills

Problem solving
Team collaboration
Strong communication
US Citizenship

Education

Bachelor's or higher in CS/EE
4+ years of systems engineering

Tools

Slurm
PBS
Apache Airflow
Kubernetes
Docker
Terraform
Ansible
Puppet
Salt

Job description

Position Summary

Base-2 Solutions is seeking a mid‑career Systems Engineer specializing in HPC and GPU infrastructure to design, develop, and optimize GPU clusters for Intelligence Community customers. This is a 100% on‑site position at the customer site in Bethesda, Maryland.


Essential Duties and Responsibilities


  • Contribute to the installation and maintenance of on‑premises and cloud‑based GPU and HPC hardware, assessing hardware performance and its interaction with software components.

  • Analyze HPC and GPU cluster performance, identify bottlenecks, and develop strategies to improve performance across Linux applications while addressing hardware and software considerations.

  • Install and configure HPC and GPU job‑scheduling and workload‑management platforms, including Slurm, PBS, Apache Airflow, and Kubernetes.

  • Apply power‑management techniques to optimize GPU power consumption across mobile and desktop Linux platforms.

  • Design and execute Linux‑based GPU performance and functionality tests, including stress testing, benchmarking, and debugging; maintain and expand the testing suite.

  • Maintain comprehensive technical documentation, including architectural specifications, code documentation, and Linux‑specific best practices for GPU development.

  • Monitor trends, innovations, and competitive developments in the GPU industry; contribute to research and propose Linux‑specific approaches to GPU design and optimization.

  • Share regular technical updates and insights with the team.


Required Qualifications


  • Systems engineering experience meeting an applicable education/experience pathway in the equivalency section.

  • Expertise in operating system integration for Linux.

  • Strong understanding of computer hardware architecture, particularly as it relates to Linux systems.

  • Knowledge of parallel computing, graphics algorithms, and real‑time rendering in Linux environments.

  • Excellent problem‑solving skills and ability to collaborate within a team.

  • Strong communication skills for conveying technical information in a Linux context.

  • Proficiency with scripting languages such as Python or BASH.

  • Proficiency with automation tools such as Ansible, Puppet, Salt, Terraform, and related tools.

  • US Citizenship.

  • DoD 8570.11 IAT Level II certification requirements or an acceptable IAT Level III certification.


Preferred Qualifications


  • Knowledge of GPU virtualization, cloud computing, and emerging Linux‑based technologies.

  • Experience with container technologies, including Docker and Kubernetes.

  • Experience with Prometheus and Grafana for monitoring.

  • Knowledge of distributed resource‑scheduling systems.

  • Understanding of data centre networking hardware and cabling concepts.

  • Understanding of networking technologies such as DHCP, DNS, TCP/IP, VLANs, HSRP, and SNMP.

  • Knowledge of data centre networking security principles, including firewall ACLs, IPS/IDS, and policy‑based routing.


Required Education and Experience Equivalency


  • Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field with 4+ years of relevant systems engineering experience.

  • Additional years of relevant experience in lieu of a degree.


Required Certifications


  • Meet DoD 8570.11 IAT Level II certification requirements, currently one of Security+ CE, CCNA‑Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification, or an acceptable IAT Level III certification, including one of CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, or CCSP.


Required Security Clearance


  • Active Top Secret/SCI

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Engineer - HPC & GPU Infrastructure
Senior Systems Engineer - HPC & GPU Infrastructure

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 140,000 - 200,000
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

FiveInsights • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos Inc • Bethesda (MD)

On-site
USD 70,000 - 90,000
Senior HPC & GPU Systems Engineer - On‑Site (Top Secret)
Senior HPC & GPU Systems Engineer - On‑Site (Top Secret)

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 110,000 - 160,000
Systems Engineer - HPC & GPU Infrastructure - TS/SCI
Systems Engineer - HPC & GPU Infrastructure - TS/SCI

Xcelerate Solutions • Bethesda (MD)

On-site
USD 120,000 - 170,000
Senior Systems Engineer
Senior Systems Engineer

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 120,000 - 180,000
Platform Engineer
Platform Engineer

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 140,000 - 190,000
Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos • Bethesda (MD)

On-site
USD 87,100 - 157,450
Health and Wellness programs
Income Protection
Paid Leave
+1
Senior Systems Engineer
Senior Systems Engineer

Talentify • Bethesda (MD)

On-site
USD 120,000 - 160,000