Senior DevOps Engineer – HPC / EDA

Whiz Global LLC

Rancho Cordova (CA)

Hybride

USD 140 000 - 170 000

Plein temps

Il y a 4 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Une candidature conçue pour ce poste — un CV et une lettre de motivation personnalisés qui correspondent à l’offre.

Passez les filtres ATS

Résumé du poste

Whiz Global LLC is seeking a Senior DevOps Engineer to support and maintain HPC and EDA infrastructure. The role requires strong Linux administration, SLURM workload management, and IaC tooling.

You will implement automation with Terraform and Ansible, manage Azure-based environments, and collaborate with IDAM and storage teams to ensure reliable, scalable platform operations. This is a hybrid/remote California position.

Qualifications

  • 5 years of experience in DevOps, Platform Engineering, Linux Systems Engineering, or related infrastructure role.
  • Hands-on experience administering HPC clusters using SLURM or an equivalent workload manager.
  • Experience supporting EDA, scientific computing, or HPC environments.
  • Strong hands-on expertise in Ansible, including playbook and role development, automation, and configuration management.
  • Experience with Terraform and Infrastructure as Code (IaC).
  • Strong Linux administration skills (SUSE Linux Enterprise Server (SLES 15) and Ubuntu).
  • Experience with Microsoft Azure cloud compute environments.
  • Working knowledge of SSSD, LDAP, Active Directory, and Okta in enterprise Linux environments.
  • Experience with NetApp or comparable enterprise storage platforms, NFS, AutoFS, and filesystem administration.
  • Familiarity with Git, GitHub, pull requests, code reviews, and repository management.
  • Experience with monitoring and logging tools such as Splunk and Datadog.
  • Scripting skills in Python, Bash, and/or Perl.
  • Experience using ServiceNow for production change management.
  • Ability to create MOPs, runbooks, architecture diagrams, and technical implementation documentation.
  • Strong troubleshooting, analytical, communication, and cross-functional collaboration skills.

Responsabilités

  • Administer and support SLURM-based HPC compute environments, including workload management, partition configuration, and infrastructure migration planning.
  • Plan and execute infrastructure changes and migrations using Terraform and Ansible.
  • Support Azure-based EDA user environments, including ThinLinc/VNC access and related services.
  • Coordinate with EDA, TD NAND, storage, and IDAM teams to implement infrastructure changes and platform upgrades.
  • Develop and maintain formal Methods of Procedure (MOPs), operational runbooks, and service cutover documentation.

Description du poste

Job Description: Senior DevOps Engineer -- HPC / EDA

Job Title: Senior DevOps Engineer -- HPC / EDA

Contract Duration: 12 Months

Department: IT Datacenter (ITDC)

Work Location: Rancho Cardova, CA (Hybrid/Remote)

Position Summary

We are seeking an experienced Senior DevOps Engineer to support and maintain High-Performance Computing (HPC) and Electronic Design Automation (EDA) infrastructure. The ideal candidate will have strong hands‑on experience in Linux systems administration, SLURM workload management, infrastructure automation, Terraform, Ansible, Azure cloud environments, enterprise identity management, and storage administration.

The candidate will work closely with infrastructure, EDA, storage, and Identity and Access Management (IDAM) teams to deliver reliable, secure, and scalable HPC platform operations. This role requires a senior‑level engineer who can independently manage production infrastructure changes, troubleshoot complex technical issues, develop automation, and maintain comprehensive technical documentation.

Key Responsibilities

1. HPC / EDA Platform Operations
  • Administer and support SLURM‑based HPC compute environments, including workload management, partition configuration, and infrastructure migration planning.

  • Plan and execute infrastructure changes and migrations using Terraform and Ansible.

  • Support Azure‑based EDA user environments, including ThinLinc/VNC access and related services.

  • Coordinate with EDA, TD NAND, storage, and IDAM teams to implement infrastructure changes and platform upgrades.

  • Develop and maintain formal Methods of Procedure (MOPs), operational runbooks, and service cutover documentation.

2. Automation & Infrastructure as Code
  • Develop, maintain, and enhance Ansible playbooks and roles for Linux provisioning, authentication, system configuration, and platform administration.

  • Ensure compatibility of Ansible playbooks across multiple versions and SLES 15 environments.

  • Use Terraform to automate infrastructure provisioning and support configuration management and migration activities.

  • Manage infrastructure code through Git and GitHub, including pull requests, code reviews, and internal repository contributions.

  • Support artifact and binary management using Artifactory.

  • Implement production changes through established change management processes using ServiceNow.

3. Identity & Access Management
  • Configure and troubleshoot enterprise authentication and identity integration for Linux and HPC environments.

  • Work with SSSD, LDAP, Active Directory, and Okta to support centralized authentication and access management.

  • Audit and reconcile Linux user and group identity information, including UID/GID consistency across multiple directory and authentication domains.

  • Validate authentication, authorization, and access behavior across compute and storage environments.

  • Extend SSSD‑based corporate authentication to new compute environments and automate configurations using Ansible.

4. Monitoring, Logging & Operational Readiness
  • Evaluate and implement log management solutions for HPC systems, including potential Splunk integration.

  • Monitor, troubleshoot, and resolve production Linux service issues involving ThinLinc/VNC, AutoFS, Datadog, and related infrastructure services.

  • Develop and maintain operational scripts using Python, Bash, and Perl as required.

  • Support production readiness assessments, infrastructure validation, and operational improvement initiatives.

  • Create and maintain technical documentation, architecture diagrams, implementation guides, and end‑user instructions in Confluence.

5. Enterprise Storage & Filesystem Administration
  • Support enterprise storage platforms, including NetApp Storage Virtual Machines (SVMs) and comparable storage solutions.

  • Work with NFS, AutoFS, RootSquash, and Linux filesystem configurations.

  • Assist with storage tier design, capacity planning, and IOPS performance considerations.

  • Coordinate storage‑related changes with infrastructure and HPC teams to ensure platform reliability and performance.

Required Skills & Qualifications

  • 5 years of experience in DevOps, Platform Engineering, Linux Systems Engineering, or a related infrastructure role.

  • Hands‑on experience administering HPC clusters using SLURM or an equivalent workload manager.

  • Experience supporting EDA, scientific computing, or high‑performance computing environments.

  • Strong hands‑on expertise in Ansible, including playbook and role development, automation, and configuration management.

  • Experience with Terraform and Infrastructure as Code (IaC).

  • Strong Linux administration skills, preferably with SUSE Linux Enterprise Server (SLES 15) and Ubuntu.

  • Experience with Microsoft Azure cloud compute environments.

  • Working knowledge of SSSD, LDAP, Active Directory, and Okta in enterprise Linux environments.

  • Experience with NetApp or comparable enterprise storage platforms, NFS, AutoFS, and filesystem administration.

  • Familiarity with Git, GitHub, pull requests, code reviews, and repository management.

  • Experience with monitoring and logging tools such as Splunk and Datadog.

  • Scripting skills in Python, Bash, and/or Perl.

  • Experience using ServiceNow for production change management.

  • Ability to create MOPs, runbooks, architecture diagrams, and technical implementation documentation.

  • Strong troubleshooting, analytical, communication, and cross‑functional collaboration skills.

  • Ability to work independently and manage complex infrastructure tasks with minimal supervision.

Preferred Qualifications

  • Experience with SLES 12 and/or SLES 15 in enterprise environments.

  • Experience migrating configuration artifacts and binaries to Artifactory.

  • Background in semiconductor, NAND, storage, or high‑tech manufacturing IT environments.

  • Experience with HPC datacenter migrations and large‑scale infrastructure transitions.

  • Familiarity with enterprise storage performance optimization and capacity planning.

Ideal Candidate Profile

The ideal candidate is a hands‑on senior infrastructure engineer with a strong combination of Linux administration, HPC/EDA operations, SLURM, Ansible, Terraform, Azure, enterprise authentication, and storage expertise. The candidate should be comfortable troubleshooting complex production environments, automating repetitive tasks, documenting infrastructure changes, and collaborating with multiple technical teams.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior DevOps Engineer for HPC/EDA Platforms (Hybrid/Remote)
Senior DevOps Engineer for HPC/EDA Platforms (Hybrid/Remote)

Whiz Global LLC • Rancho Cordova (CA)

Hybride
USD 140 000 - 170 000
Senior DevOps Engineer - 26-03060
Senior DevOps Engineer - 26-03060

Akraya, Inc. • San Jose (CA)

Sur place
USD 90 000 - 96 000
EDA Infrastructure / CAD Engineer
EDA Infrastructure / CAD Engineer

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

Sur place
USD 120 000 - 180 000
CAE Engineer
CAE Engineer

Saigepartners • San Jose (CA)

Sur place
USD 90 000 - 130 000
Senior DevOps Engineer for HPC & EDA Infrastructure
Senior DevOps Engineer for HPC & EDA Infrastructure

E-Space • Santa Clara (CA)

Sur place
USD 140 000 - 200 000
Competitive salaries
Paid holidays
Paid time off
+4
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine • Berkeley (CA)

À distance
USD 140 000 - 190 000
Paid time off
401k match
Health care benefits
Sr. Linux Engineer (Contract)
Sr. Linux Engineer (Contract)

OVT group • Santa Clara (CA), Northern (KY)

Sur place
USD 83 000 - 103 000
HPC Systems Engineer
HPC Systems Engineer

EITR Technologies LLC • Annapolis (MD)

Sur place
USD 120 000 - 180 000
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine Performance Solutions • Berkeley (CA)

À distance
USD 140 000 - 190 000
Paid time off
401k match
Health care benefits
ZR_2754_JOB
ZR_2754_JOB

Zohorecruit • Jacksonville (FL)

À distance
USD 180 000 - 240 000