Systems Engineer - HPC & GPU Infrastructure

Leidos

Bethesda (MD)

On-site

USD 87,100 - 157,450

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health and Wellness programs
Income Protection
Paid Leave
Retirement

Job summary

Leidos, located in Bethesda, MD, seeks a Mid-Career Systems Engineer specializing in HPC & GPU Infrastructure. This on-site position involves designing and optimizing GPU clusters for the Intelligence Community.

The ideal candidate will have a Bachelor’s degree in Computer Science or a related field, at least 4 years of relevant experience, and expertise in Linux systems, hardware architecture, and performance optimization. Strong communication skills and relevant certifications are essential.

Leidos provides a competitive pay range of $87,100.00 - $157,450.00, alongside Health and Wellness programs, Paid Leave, and Retirement benefits.

Qualifications

  • 4+ years of relevant systems engineering experience.
  • Candidate must meet DoD 8570.11- IAT Level II certification requirements.
  • Strong understanding of computer hardware as it relates to Linux.

Responsibilities

  • Contribute to the installation and maintenance of GPU and HPC hardware.
  • Analyze HPC/GPU cluster performance and develop enhancements.
  • Design and execute tests to validate GPU performance on Linux.

Skills

Operating system integration for Linux
Computer hardware architecture
Scripting languages (Python, BASH)
Automation tools (Ansible, Puppet, etc.)
Problem-solving skills
Communication skills
Understanding of parallel computing

Education

Bachelor's or higher degree in Computer Science or related field

Tools

Kubernetes
Docker
Prometheus/Grafana

Job description

Description

Leidos is looking for a mid-career Systems Engineer (Mid-Career) – HPC & GPU Infrastructure with a deep understanding of operating systems, hardware, Kubernetes, and NVIDIA GPU products. As a Systems Engineer – HPC & GPU Infrastructure, you will play a pivotal role in designing, developing, and optimizing GPU clusters for the IC community customers.

This is a 100% on-site position. All work must be performed at the customer site in Bethesda at the Intelligence Community Campus.

Responsibilities
  • HPC and GPU environment engineering: Contribute to the installation and maintenance of GPU and HPC hardware on-prem and in the cloud, providing insights into hardware performance to ensure efficient interaction with software components.
  • Performance Optimization: Analyze HPC/GPU cluster performance, identify bottlenecks, and develop strategies to enhance performance across various applications in Linux, addressing both hardware and software considerations. Regularly monitor and improve performance.
  • HPC/GPU tooling: Install and configure HPC/GPU job scheduling and workload management platforms such as Slurm , PBS , Apache Airflow , Kubernetes
  • Power Efficiency: Work on power management techniques to optimize GPU power consumption, ensuring efficient operation on both mobile and desktop Linux platforms. Continuously assess and enhance power efficiency strategies.
  • Testing and Validation: Design and execute tests to validate GPU performance and functionality on Linux, including stress testing, benchmarking, and debugging to ensure robust operation. Maintain and expand the testing suite.
  • Documentation: Maintain comprehensive technical documentation, including architectural specifications, code documentation, and Linux-specific best practices for GPU development. Keep documentation up to date with changes and improvements.
  • Industry Insight: Stay updated on the latest trends, innovations, and competitive landscapes within the GPU industry, contributing to research efforts and proposing Linux-specific approaches to GPU design and optimization. Share regular updates and insights with the team.
You Bring
  • Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field. Additional years of experience may be considered in lieu of a degree.
  • 4+ years of relevant systems engineering experience
  • Expertise in operating system integration for Linux.
  • Strong understanding of computer hardware architecture, particularly as it relates to Linux systems.
  • Knowledge of parallel computing, graphics algorithms, and real-time rendering in Linux environments.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Strong communication skills for conveying technical information in a Linux context.
  • Proficiency with scripting languages such as Python or BASH.
  • Proficiency with automation tools such Ansible, Puppet, Salt, Terraform, etc.
  • Candidate must, at a minimum, meet DoD 8570.11- IAT Level II certification requirements (currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification). An IAT Level III certification would also be acceptable (CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, CCSP).
Clearance
  • Active TS/SCI clearance with Polygraph required OR active TS/SCI and willingness to obtain and maintain a Poly.
  • US Citizenship is required due to the nature of the government contracts we support.
Preferred Qualifications
  • Knowledge of GPU virtualization, cloud computing, and emerging Linux-based technologies in the field.
  • Experience with container technologies (Docker, Kubernetes)
  • Experience with Prometheus/Grafana for monitoring
  • Knowledge of distributed resource scheduling systems
  • Understanding data center networking hardware and cabling concepts.
  • Understanding of networking technologies such as DHCP, DNS, TCP/IP, VLANs, HSRP, and SNMP.
  • Knowledge of data center networking security principles Firewall ACLs, IPS/IDS, and Policy Based Routing.
Pay Range

Pay Range: $87,100.00 - $157,450.00

The Leidos pay range for this job level is a general guideline only and not a guarantee of compensation or salary. Additional factors considered in extending an offer include responsibilities of the job, education, experience, knowledge, skills, and abilities, as well as internal equity, alignment with market data, applicable bargaining agreement (if any), or other law.

Pay and Benefits

Pay and benefits are fundamental to any career decision. Employment benefits include competitive compensation, Health and Wellness programs, Income Protection, Paid Leave and Retirement.

Commitment to Non-Discrimination

All qualified applicants will receive consideration for employment without regard to sex, race, ethnicity, age, national origin, citizenship, religion, physical or mental disability, medical condition, genetic information, pregnancy, family structure, marital status, ancestry, domestic partner status, sexual orientation, gender identity or expression, veteran or military status, or any other basis prohibited by law. Leidos will also consider for employment qualified applicants with criminal histories consistent with relevant laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos Inc • Bethesda (MD)

On-site
USD 70,000 - 90,000
Systems and Platform Engineer
Systems and Platform Engineer

Leidos Inc • Bethesda (MD)

On-site
USD 107,000 - 196,000
Systems and Platform Engineer
Systems and Platform Engineer

Quest International • Bethesda (MD)

Hybrid
USD 107,900 - 195,050
On-site at ICC Bethesda
Competitive pay and benefits
Health and wellness programs
Systems and Platform Engineer
Systems and Platform Engineer

Leidos • Bethesda (MD)

On-site
USD 107,000 - 196,000
Health & Wellness
Income Protection
Paid Leave
+1
GPU Software Engineer
GPU Software Engineer

Leidos Inc • Arlington (VA)

On-site
USD 107,000 - 196,000
4 weeks Paid Time Off
11 paid Holidays
401K with a 6% company match
+3
Junior Software Engineer
Junior Software Engineer

Leidos • Bowie (MD)

On-site
USD 70,000 - 126,000
Health and Wellness programs
Income Protection
Paid Leave and Retirement
Computer Systems Analyst II
Computer Systems Analyst II

Leidos Inc • Huntsville (AL)

On-site
USD 70,000 - 126,000
GPU HPC Systems Engineer – Linux, Kubernetes & On-Prem
GPU HPC Systems Engineer – Linux, Kubernetes & On-Prem

Leidos • Bethesda (MD)

On-site
USD 87,000 - 158,000
Health and Wellness programs
Income Protection
Paid Leave
+1
Linux Server Systems Engineer
Linux Server Systems Engineer

Quest Oracle Community • Bethesda (MD)

Hybrid
USD 87,000 - 157,000
Linux Server Systems Engineer
Linux Server Systems Engineer

Koitecc Solutions • Bethesda (MD)

Hybrid
USD 94,000 - 157,000