AI Kernel & GPU HPC Cluster Engineer

Blue Signal Search

Santa Clara (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Blue Signal Search seeks an experienced AI Kernel / Cluster Engineer to optimize production GPU and HPC clusters, focusing on Linux kernel level operations and driver management for reliable AI workloads.

You will collaborate with AI and infrastructure teams to ensure high availability, performance, and scalability using SLURM/Kubernetes, NVIDIA tools, and robust automation with Bash and Python.

Qualifications

  • Minimum 3 years managing production GPU or HPC clusters.
  • Advanced Ubuntu Linux administration experience with kernel-level expertise (driver management, tuning, boot config).
  • Production LXC container administration experience.
  • GPU networking expertise including InfiniBand and RoCEv2.
  • NVIDIA tools (DCGM, UFM, CUDA, NCCL, SHARP) proficiency.
  • Experience with SLURM or Kubernetes for workload orchestration.
  • Bash/Python automation for diagnostics and operations.

Responsibilities

  • Manage and optimize production GPU/HPC clusters for AI workloads.
  • Administer Ubuntu Linux at the kernel level, including drivers and tuning.
  • Monitor GPU infrastructure health using NVIDIA tools and resolve issues.
  • Configure and maintain LXC containers for reliable isolation and networking.
  • Support high-performance cluster networking (InfiniBand, RoCEv2).
  • Manage workloads with SLURM or Kubernetes to maximize efficiency.
  • Develop Bash/Python automation to streamline operations.
  • Troubleshoot Linux, GPU, containers, storage, and networking; document solutions.

Skills

Linux troubleshooting
Analytical thinking
Strong communication
Automation scripting (Bash/Python)

Tools

NVIDIA DCGM
NVIDIA UFM
CUDA
NCCL
SHARP
SLURM
Kubernetes
InfiniBand
RoCEv2
LXC
Ubuntu kernel tuning

Job description

Blue Signal Search seeks an experienced AI Kernel / Cluster Engineer to optimize production GPU and HPC clusters, focusing on Linux kernel level operations and driver management for reliable AI workloads.

You will collaborate with AI and infrastructure teams to ensure high availability, performance, and scalability using SLURM/Kubernetes, NVIDIA tools, and robust automation with Bash and Python.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD 180,000 - 240,000
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI HPC Systems Engineer: GPU Clusters & ML Platforms
AI HPC Systems Engineer: GPU Clusters & ML Platforms

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
Senior MLOps Engineer: GPU AI Infra & Production
Senior MLOps Engineer: GPU AI Infra & Production

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Senior GPU Kernel Engineer for AI Infrastructure
Senior GPU Kernel Engineer for AI Infrastructure

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 170,000 - 250,000
Hybrid work model
Medical, dental, vision insurance
Senior GPU Systems Engineer for AI Clusters & Linux
Senior GPU Systems Engineer for AI Clusters & Linux

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1