Lead HPC-AI Cluster Engineer

NVIDIA AI

Layer-de-la-Haye

On-site

GBP 90,000 - 130,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA is seeking an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team in the United Kingdom. You will design and maintain large-scale HPC/AI clusters, manage Linux job workloads, and develop CI/CD pipelines for automated deployment and monitoring of infrastructure.

You will work with researchers, developers, and customers to craft improved workflows and solutions. Requirements include 8+ years in CS/Engineering, expertise in HPC/AI tech, Slurm/K8s, Linux

Qualifications

  • Degree in Computer Science, Engineering, or related field and 8+ years of experience.
  • Experience with HPC and AI solution technologies from CPUs/GPUs to high-speed interconnects.
  • Experience with job scheduling workloads and orchestration tools such as Slurm, K8s.
  • Excellent knowledge of Windows and Linux networking and OS security.
  • Experience with Lustre/GPFS/GPUs storage tech. Familiarity with modern storage tech.
  • Python and Bash scripting experience.
  • Familiarity with automation/configuration tools: Jenkins, Ansible, Puppet/Chef.
  • Deep knowledge of InfiniBand/Ethernet networking and virtualization platforms.

Responsibilities

  • Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting.
  • Manage Linux job/workload schedules and orchestration tools.
  • Develop and maintain CI/CD pipelines.
  • Develop tooling to automate deployment and management of large-scale infra.
  • Deploy monitoring for servers, network and storage.
  • Troubleshoot from bare metal to application level.
  • Document standard methodologies for internal teams.
  • Support R&D activities and engage in PoCs for future improvements.

Skills

HPC/AI system design
Linux system administration
Scripting (bash, Python)
Technical troubleshooting

Education

Bachelor's degree in CS/Engineering

Tools

Slurm
Kubernetes
Jenkins
Ansible
Puppet/Chef
VMware/Hyper-V/KVM
Azure/AWS/Google Cloud

Job description

NVIDIA is seeking an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team in the United Kingdom. You will design and maintain large-scale HPC/AI clusters, manage Linux job workloads, and develop CI/CD pipelines for automated deployment and monitoring of infrastructure.

You will work with researchers, developers, and customers to craft improved workflows and solutions. Requirements include 8+ years in CS/Engineering, expertise in HPC/AI tech, Slurm/K8s, Linux

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI GPU Compute Cluster Architect
Lead AI GPU Compute Cluster Architect

Radiant • Greater London

On-site
GBP 120,000 - 190,000
25 days leave
Medical insurance
Cycle to Work
+3
Platform Engineer - HPC & AI Infra Architect
Platform Engineer - HPC & AI Infra Architect

Carbon3.ai • United Kingdom

On-site
GBP 85,000 - 120,000
Senior Solution Engineer – GPU and AI Infrastructure (F673F3F)
Senior Solution Engineer – GPU and AI Infrastructure (F673F3F)

Referment • Greater London

On-site
GBP 90,000 - 120,000
HPC IT Lead Engineer for GPU/CPU Simulations
HPC IT Lead Engineer for GPU/CPU Simulations

Linuxconfig • Greater London

Hybrid
GBP 90,000 - 130,000
Equity options
10% employer pension contribution
Free office lunches
+10
AI Data Center Engineer: GPU HPC & Networking
AI Data Center Engineer: GPU HPC & Networking

Hamilton Barnes Associates Limited • United Kingdom

On-site
GBP 45,000 - 75,000
HPC & AI Platform Engineer (GPU/Networking)
HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Senior Field Engineer - AI/ML HPC & Kubernetes Solutions
Senior Field Engineer - AI/ML HPC & Kubernetes Solutions

CoreWeave • Greater London

On-site
GBP 98,000 - 130,000
Medical Insurance
Dental Insurance
Pension Plan
+5
Senior HPC Engineer
Senior HPC Engineer

Gazelle Global Consulting Limited • Stevenage

Hybrid
GBP 60,000 - 90,000
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment
MLOps & Infrastructure Engineer for AI HPC
MLOps & Infrastructure Engineer for AI HPC

EngineersOfAI • United Kingdom

On-site
GBP 40,000 - 70,000