Systems Engineer - HPC & GPU Infrastructure

Leidos Inc

Bethesda (MD)

On-site

USD 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Leidos Inc in Bethesda, Maryland is seeking a Junior Systems Engineer for HPC and GPU Infrastructure. This role involves designing and optimizing GPU clusters for the IC community customers, requiring a strong background in Linux and hardware integration.

The ideal candidate will have at least 2 years of experience in systems engineering, expertise in operating systems, and a bachelor's degree in a related field. A TS/SCI clearance is required for this role.

Qualifications

  • 2+ years of relevant systems engineering experience.
  • Strong understanding of graphics algorithms and real-time rendering.
  • Excellent communication skills for technical information.
  • Proficiency with automation tools like Ansible or Terraform.

Responsibilities

  • Contribute to the installation and maintenance of GPU and HPC hardware.
  • Analyze performance and develop strategies for optimization.
  • Install and configure workload management platforms like Slurm.

Skills

Linux operating system integration
Computer hardware architecture
Scripting languages (Python, BASH)
Automation tools (Ansible, Puppet)
Collaboration in a team
Problem-solving skills

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Docker
Prometheus/Grafana

Job description

Leidos is looking for a junior-careerSystems Engineer- HPC & GPU Infrastructurewith a deep understanding of operating systems, hardware, Kubernetes, and NVIDIA GPU products. As a Systems Engineer (Mid-Career) - HPC & GPU Infrastructure, you will play a pivotal role in designing, developing, and optimizing GPU clusters for the IC community customers.

Responsibilities
  • HPC and GPU environment engineering: Contribute to the installation and maintenance of GPU and HPC hardware on-prem and in the cloud, providing insights into hardware performance to ensure efficient interaction with software components.
  • Performance Optimization: Analyze HPC/GPU cluster performance, identify bottlenecks, and develop strategies to enhance performance across various applications in Linux, addressing both hardware and software considerations. Regularly monitor and improve performance.
  • HPC/GPU tooling: Install and configure HPC/GPU job scheduling and workload management platforms such as Slurm, PBS, Apache Airflow, Kubernetes.
  • Power Efficiency: Work on power management techniques to optimize GPU power consumption, ensuring efficient operation on both mobile and desktop Linux platforms. Continuously assess and enhance power efficiency strategies.
  • Testing and Validation: Design and execute tests to validate GPU performance and functionality on Linux, including stress testing, benchmarking, and debugging to ensure robust operation. Maintain and expand the testing suite.
  • Documentation: Maintain comprehensive technical documentation, including architectural specifications, code documentation, and Linux-specific best practices for GPU development. Keep documentation up to date with changes and improvements.
  • Industry Insight: Stay updated on the latest trends, innovations, and competitive landscapes within the GPU industry, contributing to research efforts and proposing Linux-specific approaches to GPU design and optimization. Share regular updates and insights with the team.
Qualifications
  • Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field. Additional years of experience may be considered in lieu of a degree.
  • 2+ years of relevant systems engineering experience
  • Expertise in operating system integration for Linux.
  • Strong understanding of computer hardware architecture, particularly as it relates to Linux systems.
  • Knowledge of parallel computing, graphics algorithms, and real-time rendering in Linux environments.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Strong communication skills for conveying technical information in a Linux context.
  • Proficiency with scripting languages such as Python or BASH.
  • Proficiency with automation tools such Ansible, Puppet, Salt, Terraform, etc.
  • Candidate must, at a minimum, meet DoD 8570.11- IAT Level II certification requirements (currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification). An IAT Level III certification would also be acceptable (CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, CCSP).
Clearance
  • Active TS/SCI clearance with Polygraph required OR active TS/SCI and willingness to obtain and maintain a Poly.
  • US Citizenship is required due to the nature of the government contracts we support.
Preferred Qualifications
  • Knowledge of GPU virtualization, cloud computing, and emerging Linux-based technologies in the field.
  • Experience with container technologies (Docker, Kubernetes)
  • Experience with Prometheus/Grafana for monitoring
  • Knowledge of distributed resource scheduling systems
  • Understanding data center networking hardware and cabling concepts.
  • Understanding of networking technologies such as DHCP, DNS, TCP/IP, VLANs, HSRP, and SNMP.
  • Knowledge of data center networking security principles Firewall ACLs, IPS/IDS, and Policy Based Routing.
Commitment to Non-Discrimination

All qualified applicants will receive consideration for employment without regard to sex, race, ethnicity, age, national origin, citizenship, religion, physical or mental disability, medical condition, genetic information, pregnancy, family structure, marital status, ancestry, domestic partner status, sexual orientation, gender identity or expression, veteran or military status, or any other basis prohibited by law. Leidos will also consider for employment qualified applicants with criminal histories consistent with relevant laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Engineer - HPC & GPU Infrastructure
Systems Engineer - HPC & GPU Infrastructure

Leidos • Bethesda (MD)

On-site
USD 87,000 - 158,000
Health and Wellness programs
Income Protection
Paid Leave
+1
Systems and Platform Engineer
Systems and Platform Engineer

Leidos • Bethesda (MD)

On-site
USD 107,000 - 196,000
Health & Wellness
Income Protection
Paid Leave
+1
Systems and Platform Engineer
Systems and Platform Engineer

Stryker Corporation • Bethesda (MD)

On-site
Confidential
On-site at ICC Bethesda
Competitive pay and benefits
Health and wellness programs
Systems and Platform Engineer
Systems and Platform Engineer

Leidos Inc • Bethesda (MD)

On-site
USD 107,000 - 196,000
HPC & GPU Systems Engineer — Linux, Kubernetes
HPC & GPU Systems Engineer — Linux, Kubernetes

Leidos Inc • Bethesda (MD)

On-site
USD 70,000 - 90,000
Systems Engineer II
Systems Engineer II

Via Logic LLC • Huntsville (AL)

On-site
USD 100,000 - 140,000
Junior Software Engineer
Junior Software Engineer

Leidos Inc • Bowie (MD)

On-site
USD 70,000 - 126,000
GPU HPC Systems Engineer – Linux, Kubernetes & On-Prem
GPU HPC Systems Engineer – Linux, Kubernetes & On-Prem

Leidos • Bethesda (MD)

On-site
USD 87,000 - 158,000
Health and Wellness programs
Income Protection
Paid Leave
+1
Systems and Platform Engineer
Systems and Platform Engineer

Via Logic LLC • Bethesda (MD)

On-site
USD 150,000 - 190,000
Junior Software Engineer
Junior Software Engineer

Leidos • Bowie (MD)

On-site
USD 70,000 - 126,000
Health and Wellness programs
Income Protection
Paid Leave and Retirement