Staff HPC Infrastructure Engineer

gh

Deutschland

Hybrid

EUR 134.000 - 184.000

Vollzeit

Vor 10 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Guardant Health is seeking a Staff-level HPC engineer to manage and advance its high-performance computing backbone. You will work across Linux systems, storage, and networking, mentoring teammates and coordinating with vendors to maintain robust, scalable infrastructure.

The role supports a hybrid work model with on-site collaboration and remote work options, including global locations, and requires participation in a 24/7 on-call rotation to keep systems healthy and performant.

Qualifikationen

  • Bachelors degree in Computer Science or a related field with 8 years of relevant experience; Masters degree with 6 years; or PhD with 3 years.
  • Strong experience in systems and/or infrastructure engineering, Linux/Unix administration, and TCP/IP networking.
  • Hands-on with automation tools (Ansible or equivalent).
  • Experience with high-performance networking and large-scale data storage in HPC environments.
  • Experience with on-premise and cloud-based infrastructure (AWS, GCP, Azure).

Aufgaben

  • Maintain and troubleshoot HPC clusters and file systems in production.
  • Develop next-generation HPC solutions and research integration approaches.
  • Mentor junior engineers on best HPC practices.
  • Work with offsite consultants and vendors for upgrades and support.
  • Participate in 24/7 on-call rotation and ensure system reliability.

Kenntnisse

HPC skills
On-call rotation
Troubleshooting
Python scripting
Shell scripting
Linux administration
TCP/IP networking
Cloud experience
Ansible
Docker
Kubernetes

Ausbildung

Bachelor9s degree in Computer Science or related
Master/PhD in a relevant field (plus)

Tools

Linux/Unix
TCP/IP networking
AWS/GCP/Azure
GPFS
Slurm
InfiniBand
Docker
Kubernetes
Ansible

Jobbeschreibung

Company Description Guardant Health is a leading precision oncology company focused on guarding wellness and giving every person more time free from cancer. Founded in 2012, Guardant™ is transforming patient care and accelerating new cancer therapies by providing critical insights into what drives disease through its advanced blood and tissue tests, real‑world data and AI analytics. Guardant tests help improve outcomes across all stages of care, including screening to find cancer early, monitoring for recurrence in early‑stage cancer, and treatment selection for patients with advanced cancer. For more information, visit guardanthealth.com and follow the company on LinkedIn , X (Twitter) and Facebook . Guardant's HPC team builds and operates the computational technology backbone of the company. This includes scalable data storage that holds petabytes of genomics data, high‑performance compute clusters running a custom bioinformatics pipeline in production and R&D environments, and the software infrastructure that hosts an ecosystem of services for internal data processing and external data integration. The HPC engineering team is looking for a Staff-level engineer with broad, all-round HPC competency and specialized depth in one or more of the following:

Guardant Health is a leading precision oncology company focused on guarding wellness and giving every person more time free from cancer. Founded in 2012, Guardant™ is transforming patient care and accelerating new cancer therapies by providing critical insights into what drives disease through its advanced blood and tissue tests, real‑world data and AI analytics. Guardant tests help improve outcomes across all stages of care, including screening to find cancer early, monitoring for recurrence in early‑stage cancer, and treatment selection for patients with advanced cancer. For more information, visit guardanthealth.com and follow the company on LinkedIn , X (Twitter) and Facebook . Guardant's HPC team builds and operates the computational technology backbone of the company. This includes scalable data storage that holds petabytes of genomics data, high‑performance compute clusters running a custom bioinformatics pipeline in production and R&D environments, and the software infrastructure that hosts an ecosystem of services for internal data processing and external data integration. The HPC engineering team is looking for a Staff-level engineer with broad, all-round HPC competency and specialized depth in one or more of the following:

Guardant Health is a leading precision oncology company focused on guarding wellness and giving every person more time free from cancer. Founded in 2012, Guardant™ is transforming patient care and accelerating new cancer therapies by providing critical insights into what drives disease through its advanced blood and tissue tests, real‑world data and AI analytics. Guardant tests help improve outcomes across all stages of care, including screening to find cancer early, monitoring for recurrence in early‑stage cancer, and treatment selection for patients with advanced cancer. For more information, visit guardanthealth.com and follow the company on LinkedIn , X (Twitter) and Facebook . Guardant's HPC team builds and operates the computational technology backbone of the company. This includes scalable data storage that holds petabytes of genomics data, high‑performance compute clusters running a custom bioinformatics pipeline in production and R&D environments, and the software infrastructure that hosts an ecosystem of services for internal data processing and external data integration. The HPC engineering team is looking for a Staff-level engineer with broad, all-round HPC competency and specialized depth in one or more of the following:

Red Hat-family OS management, networking, storage, Kubernetes, or Slurm.

Essential Duties and Responsibilities
  • Broad HPC Skills (all-round)
  • Manage multiple HPC clusters and cluster file systems
  • Integrate cloud bursting as part of the HPC abstraction work
  • Research, develop, and implement the next generation HPC solutions
  • Troubleshoot the production system stack down to source code level, e.g shell scripts, Python, and others
  • Maintain, monitor, and support the infrastructure environment and/or facilities
  • Use and maintain enhanced production monitoring and addition capability
  • Support improvements for increased system reliability and performance
  • Support multiple systems or applications of medium to high complexity complexity defined by size, technology used, and system feeds and interfaces) with multiple concurrent users, ensuring control, integrity, and accessibility
  • Support systems at remote locations, including internationally
  • Mentor junior engineers on HPC best practices
  • Work with offsite consultants to maintain the infrastructure
  • Work with vendors to troubleshoot, upgrade, and repair systems as needed
  • Represent HPC infrastructure networking and storage-integration topics in cross-functional planning with networking, SQA, DevOps/SRE, and the MSP
  • Set up and ownership supporting XDMoD instances for HPC metric and monitoring
  • Participate in a 24/7 on-call rotation
Networking
  • Act as the technical peer for HPC networking and interconnect initiatives with the dedicated networking engineer
  • Collaborate on design, performance tuning, and troubleshooting of HPC Ethernet
  • Work with enterprise networking on integration of HPC systems with the bandwidth-on-demand system that connects our sites and cloud infrastructure
  • Work with the networking infrastructure team to manage and optimize connectivity to and from HPC systems and global locations
Storage
  • Act as the technical peer for the architecture and integration strategy for. HPC storage in partnership with the dedicated storage engineer and MSP
  • Serve as a technical point of contact for the MSP storage relationship and help define and evolve SLAs, validate delivery, and escalation technical issues
  • Support the transition of day-to-day storage operations to the MSP without loss of performance or reliability
Required Qualifications
  • Bachelor’s degree in Computer Science or a related field with 8–12 years of relevant experience; Master’s degree with 6–8 years of relevant experience; or PhD with 3–5 years of relevant experience Strong experience in systems and/or infrastructure engineering, including Linux/Unix administration and TCP/IP networking.
  • Hands‑on experience with automation tools, such as Ansible or equivalent technologies.
  • Experience with high‑performance networking technologies, such as InfiniBand, RoCE, RDMA, or equivalent, including troubleshooting in production environments.
  • Experience supporting large‑scale data storage and high‑performance computing (HPC)/compute environments.
  • Experience working with both on‑premise and cloud‑based infrastructure, such as AWS, Google Cloud Platform (GCP), Azure, or similar environments.
  • Experience developing and supporting software release, operations, and infrastructure automation processes and toolsets.
  • Strong experience creating and maintaining system administration and technical documentation.
Preferred Qualifications
  • Cisco Certified Network Professional (CCNP) certification Experience with Arista and compatible networking, up to and including 400 Gb/s links Experience administering IBM's General Parallel File System (GPFS) Experience administering the Slurm scheduler Experience using Warewulf Linux support and OS management.
  • RedHat family is a must, but Debian or Suse is nice to have.
  • Experience with cloud bursting technologies
  • Experience with wide area file systems
  • Experience with Docker and Apptainer container technologies
  • Experience with Kubernetes Operating infrastructure compliant with HIPAA and SOX standards
AI & Digital Fluency

Demonstrate curiosity, sound judgment, and the ability to critically evaluate and responsibly leverage AI-enabled tools in accordance with company policies, ethical standards, and regulatory requirements to improve the efficiency, effectiveness, and quality of work.

Hybrid Work Model

This section is applicable to onsite employees who are eligible for hybrid work location as specified by management and related policies. Guardant has defined days for in‑person/onsite collaboration and work‑from‑home days for individual‑focused time. All U.S. employees who live within 50 miles of a Guardant facility will be required to be onsite on Mondays, Tuesdays, and Thursdays. We have found aligning our scheduled in‑office days allows our teams to do the best work and creates the focused thinking time our innovative work requires. At Guardant, our work model has created flexibility for better work‑life balance while keeping teams connected to advance our science for our patients.

Primary Location: Remote-USA-CA

Primary Location Base Pay Range: $155,700 - $214,150

Other US Location(s) Base Pay Range: $147,100 - $202,300

If the role is performed in Colorado, the pay range for this job is: $155,700 - $214,150 Employee may be required to lift routine office supplies and use office equipment. Majority of the work is performed in a desk/office environment; however, there may be exposure to high noise levels, fumes, and biohazard material in the laboratory environment. Ability to sit for extended periods of time.

Guardant Health is committed to providing reasonable accommodations in our hiring processes for candidates with disabilities, long‑term conditions, mental health conditions, or sincerely held religious beliefs. If you need support, please reach out to Peopleteam@guardanthealth.com

A background screening including criminal history is required for this role. GH will consider qualified applicants with criminal arrest or conviction histories in a manner consistent with applicable law including but not limited to the LA County Fair Chance Policies and the Fair Chance Act (Gov. Code Section 12952).

Guardant Health is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability.

All your information will be kept confidential according to EEO guidelines.

To learn more about the information collected when you apply for a position at Guardant Health, Inc. and how it is used, please review our Privacy Notice for Job Applicants.

Please visit our career page at: http://www.guardanthealth.com/jobs/

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Mainframe Storage Administrator
Mainframe Storage Administrator

Highmark Health • Deutschland

Hybrid
EUR 68.000 - 109.000
HPC Systems Administrator
HPC Systems Administrator

Helsing • München

Vor Ort
EUR 55.000 - 75.000
Competitive salary
Relocation support
Health & wellness support
+2
HPC Engineer Services
HPC Engineer Services

Embedded Shishya • Deutschland

Remote
EUR 60.000 - 80.000
HPC Systems Administrator Systems Architecture Munich
HPC Systems Administrator Systems Architecture Munich

helsing.ai • München

Vor Ort
EUR 85.000 - 120.000
Competitive salary
Relocation support
Learning allowance
+4
Federal Platform Engineer
Federal Platform Engineer

Horizon3 • Deutschland

Hybrid
EUR 104.000 - 164.000
Competitive compensation
Hybrid & remote work
Growth opportunities
+1
Staff / Principal Forward-Deployed Architect, Data Modernization
Staff / Principal Forward-Deployed Architect, Data Modernization

Transformcap • Deutschland

Hybrid
EUR 164.000 - 217.000
Equity packages
Robust medical/dental/vision insurance
Flexible working hours
+1
Senior Software Engineer
Senior Software Engineer

paretocaptiveservicesllc • Deutschland

Hybrid
EUR 104.000 - 147.000
Fully paid medical, dental, and vision
Flexible PTO
401k company contribution
+4
Information Security Engineer (m/f/d)
Information Security Engineer (m/f/d)

Northern Data AG • Frankfurt

Hybrid
EUR 85.000 - 120.000
Home office flexibility
Wellbeing initiatives
Sustainability focus
HPC Systems Administrator
HPC Systems Administrator

EngineersOfAI • München

Vor Ort
EUR 60.000 - 80.000
Senior Software Engineer (AI CICD)
Senior Software Engineer (AI CICD)

Socket.dev • Deutschland

Remote
EUR 136.356 - 159.805
Flexible & Remote-First Culture
Stock options
Health insurance
+1