Graphics Processing Unit (GPU) Engineer - TS/SCI

Sunayu

Bethesda (MD)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
Short-Term & Long-Term Disability

Job summary

Sunayu, LLC is seeking a skilled Systems Engineer in Bethesda, MD. The role involves designing and optimizing GPU clusters to support enterprise AI, collaborating closely with AI/ML engineers.

Candidates should have extensive experience in systems engineering, particularly with NVIDIA GPUs and Linux systems. Strong problem-solving skills and an active TS/SCI clearance are required for this fully on-site position.

Benefits include multiple medical plans, PTO, and a 401(k) match.

Qualifications

  • At least 12 years of related technical experience, or additional experience in lieu of a degree.
  • 10+ years of relevant systems engineering experience.
  • Experience managing NVIDIA GPU data center platforms.

Responsibilities

  • Design, configure, and maintain GPU clusters.
  • Ensure smooth GPU integration with Linux-based systems.
  • Analyze GPU performance and develop optimization strategies.
  • Maintain technical documentation and ensure compliance.

Skills

GPU Cluster Engineering
Operating System Integration
Performance Optimization
Tooling and Automation
Compliance & Documentation
Excellent problem-solving skills

Education

Bachelor's or higher in Computer Science or Engineering

Tools

Bash
Python
Ansible
Puppet
Salt
Kubernetes
Prometheus
Grafana

Job description

Location: Bethesda, MD

Category: Systems Engineering

Travel Required: No

Remote Type: No

Clearance: TS/SCI

Sunayu, LLC is looking for a highly skilled Systems Engineer with deep expertise in operating systems, hardware, GPU, and high-speed networking. In this role, you will design, develop, and optimize GPU clusters that power enterprise AI for the mission customers.

This is a 100% on-site position.

Responsibilities
  • GPU Cluster Engineering: Design, configure, and maintain GPU clusters. Collaborate with a multidisciplinary team to define and optimize architectures, ensuring they meet performance, power efficiency, and feature requirements.
  • Operating System Integration: Work closely with AI/ML engineers to ensure smooth GPU integration with Linux-based systems. Optimize GPU drivers for compatibility, reliability, and performance. Provide regular maintenance and updates.
  • Performance Optimization: Analyze GPU performance, identify bottlenecks, and develop strategies to improve efficiency across hardware and software layers.
  • Tooling and Automation: Build and maintain debugging tools, profiling utilities, and performance analysis software for Linux environments. Leverage scripting and configuration tools such as Bash, Python, Ansible, Puppet, and Salt.
  • Compliance & Documentation: Maintain technical documentation, architectural specifications, and Linux best practices. Support ATO (Authority to Operate) and ensure compliance with federal security standards.
Qualifications (You Bring)
  • Bachelor's or higher degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field with at least 12 years of related technical experience. Additional years of experience may be considered in lieu of a degree.
  • 10+ years of relevant systems engineering experience.
  • Experience in managing NVIDIA GPU data center platforms (DGX, HGX, H200, H100, L4s).
  • Knowledge of enterprise server components (storage/network controllers, HBA, SSDs).
  • Strong expertise with Linux distributions (RHEL, Ubuntu, Oracle, and Rocky).
  • Excellent problem-solving skills and ability to collaborate within a team.
  • Candidate must meet DoD 8570.11-IAT Level II certification requirements (Security+ CE, CCNA-Security, GICSP, GSEC, SSCP with appropriate computing environment CE) or have an IAT Level III certification (CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, CCSP).
Security Clearance

Active TS/SCI clearance with Polygraph required OR active TS/SCI and willingness to obtain and maintain a Poly.

US Citizenship is required due to the nature of the government contracts we support.

Preferred Qualifications
  • Experience with Kubernetes cluster management and AI/ML workflow orchestration (Argo, Airflow, and Kubeflow).
  • Familiarity with GPU virtualization and cloud computing.
  • Experience with Prometheus/Grafana for monitoring.
  • Knowledge of distributed resource scheduling systems (Slurm (preferred), LSF, etc.).
Pay Rate

Salary range considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, education/training, key skills, as well as market and business considerations when extending an offer.

Benefits
  • 3 Medical Plan Options
  • Dental and Vision
  • FSA, DCFSA, HSA
  • Life/AD&D Insurance
  • Short-Term & Long-Term Disability
  • Employee Assistance Program (EAP)
  • Training and Educational AssistancePaid Time Off (PTO)
  • 11 Federal holidays
  • 401k plan with up to a 6% match (100% immediate vesting)
Equal Opportunity Employer

Sunayu, LLC is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, gender expression, national origin, age, protected veteran status, disability status, marital status, genetic information, medical condition, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer 4
GPU Systems Engineer 4

Base-2 Solutions • Bethesda (MD)

On-site
USD 170,000 - 230,000
100% company-paid medical premiums
100% company-paid dental premiums
100% company-paid vision premiums
+2
GPU Systems Engineer 3 with Security Clearance
GPU Systems Engineer 3 with Security Clearance

Base-2 Solutions • Chevy Chase (MD)

On-site
USD 150,000 - 190,000
Company-paid health premiums
401(k) with match
Paid time off and holidays
Graphic Processing Unit (GPU) Engineer - TS/SCI
Graphic Processing Unit (GPU) Engineer - TS/SCI

VMD Corp • Bethesda (MD)

On-site
USD 160,000 - 210,000
GPU Systems Engineer 3
GPU Systems Engineer 3

Base-2 Solutions • Bethesda (MD)

On-site
USD 120,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+2
Senior GPU Cluster Engineer – Linux Systems, On-Site
Senior GPU Cluster Engineer – Linux Systems, On-Site

Sunayu • Bethesda (MD)

On-site
USD 120,000 - 150,000
3 Medical Plan Options
Dental and Vision
401k plan with up to a 6% match
+1
Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos • Bethesda (MD)

On-site
USD 87,000 - 158,000
Health and Wellness programs
Income Protection
Paid Leave
+1
Senior TS/SCI GPU Systems Engineer - On-Site Bethesda
Senior TS/SCI GPU Systems Engineer - On-Site Bethesda

Base-2 Solutions • Bethesda (MD)

On-site
USD 170,000 - 230,000
100% company-paid medical premiums
100% company-paid dental premiums
100% company-paid vision premiums
+2
Senior DSP Engineer | Technical SIGINT Analyst
Senior DSP Engineer | Technical SIGINT Analyst

Grey Matters Defense Solutions, LLC • Arlington (VA)

On-site
USD 140,000 - 190,000
25% employer contribution to SEP IRA
Individual Benefit Account to cover medical insurance and funded time off
Employee assistance program
+2
GPU Systems Engineer (TS/SCI) – Secure AI Clusters
GPU Systems Engineer (TS/SCI) – Secure AI Clusters

Base-2 Solutions • Bethesda (MD)

On-site
USD 120,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+2
Senior Systems Engineer (Virtualization / GPU Infrastructure)
Senior Systems Engineer (Virtualization / GPU Infrastructure)

Endepth Solutions • Laurel (MD)

On-site
USD 195,000 - 255,000
Affordable healthcare options
401(k) with company match
Paid Time Off (PTO)
+3