HPC & AI Cluster Engineer: Scale & Automate

NVIDIA Corporation

Zürich

Vor Ort

CHF 120.000 - 190.000

Vollzeit

Vor 13 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

NVIDIA is seeking a senior engineer to design, implement, and maintain large-scale HPC/AI clusters with robust monitoring, logging, and alerting. You will manage Linux job/workload schedules, orchestration tools, and develop CI/CD pipelines to streamline deployment across compute resources.

You will build tooling for automated deployment and operational monitoring, enable self-service consumption of resources, and deploy monitoring for servers, network and storage.

Qualifikationen

  • A degree in Computer Science, Engineering, or a related field and 8+ years of experience.
  • Knowledge of HPC and AI solution technologies from CPU’s and GPU’s to high speed interconnects and supporting software.
  • Experience with job scheduling workloads and orchestration tools such as Slurm, K8s.
  • Excellent knowledge of Windows and Linux networking (sockets, firewalld, iptables, wireshark, etc.) and internals, ACLs and OS level security protection and common protocols e.g. TCP, DHCP, DNS, etc.
  • Experience with multiple storage solutions such as Lustre, GPFS, Weka.io. Familiarity with newer and emerging storage technologies.
  • Python programming and bash scripting experience.
  • Comfortable with automation and configuration management tools such as Jenkins, Ansible, Puppet/chef
  • Deep knowledge of Networking Protocols like InfiniBand, Ethernet
  • Deep understanding and experience with virtual systems (for example VMware, Hyper-V, KVM, or Citrix)
  • Familiarity with cloud computing platforms (e.g. AWS, Azure, Google Cloud)

Aufgaben

  • Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting
  • Manage Linux job/workload schedules and orchestration tools
  • Develop and maintain continuous integration and delivery pipelines
  • Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources
  • Deploy monitoring solutions for the servers, network and storage
  • Perform troubleshooting bottom up from bare metal, operating system, software stack and application level
  • Being a technical resource, develop, re-define and document standard methodologies to share with internal teams
  • Support Research & Development activities and engage in POCs/POVs for future improvements

Tools

Jenkins
Ansible
Puppet/Chef

Jobbeschreibung

NVIDIA is seeking a senior engineer to design, implement, and maintain large-scale HPC/AI clusters with robust monitoring, logging, and alerting. You will manage Linux job/workload schedules, orchestration tools, and develop CI/CD pipelines to streamline deployment across compute resources.

You will build tooling for automated deployment and operational monitoring, enable self-service consumption of resources, and deploy monitoring for servers, network and storage.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • Zürich

Vor Ort
CHF 170.000 - 210.000
Senior HPC AI Cluster Architect
Senior HPC AI Cluster Architect

CH01 NVIDIA Switzerland AG • Schweiz

Vor Ort
CHF 120.000 - 180.000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

CH01 NVIDIA Switzerland AG • Schweiz

Vor Ort
CHF 120.000 - 180.000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA AI • Zürich

Vor Ort
CHF 180.000 - 260.000
HPC-AI Cluster Architect for Next-Gen Systems
HPC-AI Cluster Architect for Next-Gen Systems

NVIDIA • Zürich

Vor Ort
CHF 170.000 - 210.000
Senior HPC‑AI Systems Architect
Senior HPC‑AI Systems Architect

NVIDIA AI • Zürich

Vor Ort
CHF 180.000 - 260.000
Remote Senior Performance Engineer - AI & HPC Systems
Remote Senior Performance Engineer - AI & HPC Systems

NVIDIA Corporation • Zürich

Vor Ort
CHF 140.000 - 210.000
HPC Performance Engineer — Large-Scale GPU Clusters
HPC Performance Engineer — Large-Scale GPU Clusters

NVIDIA AI • Val-de-Travers

Vor Ort
CHF 120.000 - 180.000
Hybrid HPC Platform Engineer: Linux, Automation & AI Infra
Hybrid HPC Platform Engineer: Linux, Automation & AI Infra

Master in Integrated Building Systems ETH Zürich • Zürich

Hybrid
CHF 110.000 - 150.000
Flexible work
Public transport
Car sharing
+3
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Corporation • Zürich

Vor Ort
CHF 170.000 - 210.000