AI Kernel & GPU HPC Cluster Engineer

Blue Signal Search

Santa Clara (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Blue Signal Search seeks an experienced AI Kernel / Cluster Engineer to optimize production GPU and HPC clusters, focusing on Linux kernel level operations and driver management for reliable AI workloads.

You will collaborate with AI and infrastructure teams to ensure high availability, performance, and scalability using SLURM/Kubernetes, NVIDIA tools, and robust automation with Bash and Python.

Qualifications

  • Minimum 3 years managing production GPU or HPC clusters.
  • Advanced Ubuntu Linux administration experience with kernel-level expertise (driver management, tuning, boot config).
  • Production LXC container administration experience.
  • GPU networking expertise including InfiniBand and RoCEv2.
  • NVIDIA tools (DCGM, UFM, CUDA, NCCL, SHARP) proficiency.
  • Experience with SLURM or Kubernetes for workload orchestration.
  • Bash/Python automation for diagnostics and operations.

Responsibilities

  • Manage and optimize production GPU/HPC clusters for AI workloads.
  • Administer Ubuntu Linux at the kernel level, including drivers and tuning.
  • Monitor GPU infrastructure health using NVIDIA tools and resolve issues.
  • Configure and maintain LXC containers for reliable isolation and networking.
  • Support high-performance cluster networking (InfiniBand, RoCEv2).
  • Manage workloads with SLURM or Kubernetes to maximize efficiency.
  • Develop Bash/Python automation to streamline operations.
  • Troubleshoot Linux, GPU, containers, storage, and networking; document solutions.

Skills

Linux troubleshooting
Analytical thinking
Strong communication
Automation scripting (Bash/Python)

Tools

NVIDIA DCGM
NVIDIA UFM
CUDA
NCCL
SHARP
SLURM
Kubernetes
InfiniBand
RoCEv2
LXC
Ubuntu kernel tuning

Job description

Blue Signal Search seeks an experienced AI Kernel / Cluster Engineer to optimize production GPU and HPC clusters, focusing on Linux kernel level operations and driver management for reliable AI workloads.

You will collaborate with AI and infrastructure teams to ensure high availability, performance, and scalability using SLURM/Kubernetes, NVIDIA tools, and robust automation with Bash and Python.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior MLOps Engineer: GPU AI Infra & Production
Senior MLOps Engineer: GPU AI Infra & Production

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000