AI Infra Engineer — GPU Clusters & HPC

Veeda AI

Zürich

On-site

CHF 120,000 - 180,000

Full time

17 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Veeda AI seeks a Senior HPC Infrastructure Engineer to design, deploy, and operate GPU clusters on Slurm and Kubernetes, ensuring compute health, interconnects, and storage. You will implement security, observability, and end-to-end ownership in a fast-moving AI research environment.

You will collaborate with model and data teams, manage PoC rounds with vendors, and help forecast resource needs for ongoing projects in Zürich.

Qualifications

  • Bachelor's degree or equivalent hands-on HPC/infrastructure experience.

Responsibilities

  • GPU Cluster Operations: Design, deploy, and operate Slurm-on-Kubernetes GPU clusters.
  • Access Control and Security: Design, deploy, and operate identity provisioning and access control.
  • Network and Data: Design and maintain data transmission and caching services for timely delivery.

Skills

Linux networking
Kernel tuning
PCIe/NUMA
Hardware diagnostics
Distributed storage

Education

Bachelor's degree in CS/CE or equivalent

Tools

Ansible
Terraform
Helm
Python scripting

Job description

Veeda AI seeks a Senior HPC Infrastructure Engineer to design, deploy, and operate GPU clusters on Slurm and Kubernetes, ensuring compute health, interconnects, and storage. You will implement security, observability, and end-to-end ownership in a fast-moving AI research environment.

You will collaborate with model and data teams, manage PoC rounds with vendors, and help forecast resource needs for ongoing projects in Zürich.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC & AI Network Architect for Scalable AI Infra
Senior HPC & AI Network Architect for Scalable AI Infra

NVIDIA • Zürich

On-site
CHF 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Senior Software Engineer: AI Inference & HPC
Senior Software Engineer: AI Inference & HPC

NVIDIA Switzerland AG • Zürich

On-site
CHF 140,000 - 210,000
Senior HPC AI Network Architect: Scalable Infra
Senior HPC AI Network Architect: Scalable Infra

NVIDIA Corporation • Zürich

On-site
CHF 180,000 - 240,000
Senior ML Engineer, GPU Compute & Infrastructure - On-site
Senior ML Engineer, GPU Compute & Infrastructure - On-site

Flexion • Zürich

On-site
CHF 150,000 - 210,000
Enhanced pension plan
Relocation & permit sponsorship
Enhanced holiday & paid leave perks
+1
HPC DevOps Engineer for AI/ML Platforms on Kubernetes
HPC DevOps Engineer for AI/ML Platforms on Kubernetes

ETH Zürich • Lugano

Hybrid
CHF 80,000 - 100,000
Public transport season tickets
Car sharing
Wide range of sports offered by ASVZ
+2
Kubernetes‑Driven DevOps Engineer for AI/ML HPC
Kubernetes‑Driven DevOps Engineer for AI/ML HPC

ETH Zürich • Lugano

On-site
CHF 90,000 - 120,000
Public transport season tickets
Childcare benefits
Attractive pension benefits
Senior GPU Networking Architect for AI Data Centers
Senior GPU Networking Architect for AI Data Centers

NVIDIA Corporation • Zürich

On-site
CHF 180,000 - 240,000
SLURM & HPC Orchestration Engineer for AI Workloads
SLURM & HPC Orchestration Engineer for AI Workloads

Roche • Kaiseraugst

On-site
CHF 150,000 - 190,000
Remote Senior Performance Engineer - AI & HPC Systems
Remote Senior Performance Engineer - AI & HPC Systems

NVIDIA Corporation • Zürich

On-site
CHF 140,000 - 210,000