HPC-AI Cluster Architect for Next-Gen Systems

NVIDIA

Zürich

Vor Ort

CHF 170.000 - 210.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

NVIDIA in Zürich is seeking an experienced HPC‑AI Engineer to design, deploy and operate large‑scale computing clusters for research and production workloads, spanning CPU/GPU compute, high‑speed interconnects and storage.

You will own deployment pipelines, monitoring and automation, work with researchers and customers to optimize workflows, and collaborate with HPC, OS and networking teams to bring up scalable performance platforms.

Qualifikationen

  • A degree in Computer Science, Engineering, or a related field and 8+ years of experience
  • Knowledge of HPC and AI solution technologies from CPU’s and GPU’s to high speed interconnects and supporting software
  • Experience with job scheduling workloads and orchestration tools such as Slurm, K8s
  • Excellent knowledge of Windows and Linux networking (sockets, firewalld, iptables, wireshark, etc.) and internals, ACLs and OS level security protection and common protocols e.g. TCP, DHCP, DNS, etc.
  • Experience with multiple storage solutions such as Lustre, GPFS, Weka.io. Familiarity with newer and emerging storage technologies.
  • Python programming and bash scripting experience
  • Comfortable with automation and configuration management tools such as Jenkins, Ansible, Puppet/chef
  • Deep knowledge of Networking Protocols like InfiniBand, Ethernet
  • Deep understanding and experience with virtual systems (for example VMware, Hyper-V, KVM, or Citrix)
  • Familiarity with cloud computing platforms (e.g. AWS, Azure, Google Cloud)

Aufgaben

  • Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting
  • Manage Linux job/workload schedules and orchestration tools
  • Develop and maintain continuous integration and delivery pipelines
  • Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources
  • Deploy monitoring solutions for the servers, network and storage
  • Perform troubleshooting bottom up from bare metal, operating system, software stack and application level
  • Being a technical resource, develop, re-define and document standard methodologies to share with internal teams
  • Support Research & Development activities and engage in POCs/POVs for future improvements

Kenntnisse

HPC expertise
AI solutions
Linux troubleshooting
Scripting
System design

Ausbildung

Degree in Computer Science or Engineering

Tools

Slurm
Kubernetes
Docker
Jenkins
Ansible
Puppet
VMware
Hyper-V
Git
Weka.io
Lustre

Jobbeschreibung

NVIDIA in Zürich is seeking an experienced HPC‑AI Engineer to design, deploy and operate large‑scale computing clusters for research and production workloads, spanning CPU/GPU compute, high‑speed interconnects and storage.

You will own deployment pipelines, monitoring and automation, work with researchers and customers to optimize workflows, and collaborate with HPC, OS and networking teams to bring up scalable performance platforms.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC AI Cluster Architect
Senior HPC AI Cluster Architect

CH01 NVIDIA Switzerland AG • Schweiz

Vor Ort
CHF 120.000 - 180.000
Senior AI Network Architect for Scalable Distributed HPC
Senior AI Network Architect for Scalable Distributed HPC

NVIDIA Gruppe • Rüti (ZH)

Vor Ort
CHF 221.000 - 507.000
Comprehensive benefits package
HPC & AI Cluster Engineer: Scale & Automate
HPC & AI Cluster Engineer: Scale & Automate

NVIDIA Corporation • Zürich

Vor Ort
CHF 120.000 - 190.000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • Zürich

Vor Ort
CHF 170.000 - 210.000
HPC DevOps Engineer for AI/ML Platforms on Kubernetes
HPC DevOps Engineer for AI/ML Platforms on Kubernetes

ETH Zürich • Lugano

Hybrid
CHF 80.000 - 100.000
Public transport season tickets
Car sharing
Wide range of sports offered by ASVZ
+2
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

CH01 NVIDIA Switzerland AG • Schweiz

Vor Ort
CHF 120.000 - 180.000
HPC & AI Platform Automation Engineer
HPC & AI Platform Automation Engineer

ETH Zürich • Zürich

Vor Ort
CHF 80.000 - 120.000
Kubernetes‑Driven DevOps Engineer for AI/ML HPC
Kubernetes‑Driven DevOps Engineer for AI/ML HPC

ETH Zürich • Lugano

Vor Ort
CHF 90.000 - 120.000
Public transport season tickets
Childcare benefits
Attractive pension benefits
Senior AI Platform & GPU HPC Engineer
Senior AI Platform & GPU HPC Engineer

Finders SA • Basel

Vor Ort
CHF 140.000 - 190.000
Senior HPC and AI Network Software Architect
Senior HPC and AI Network Software Architect

NVIDIA Gruppe • Rüti (ZH)

Vor Ort
CHF 221.000 - 507.000
Comprehensive benefits package