AI Kernel / Cluster Engineer

Blue Signal Search

Santa Clara (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Blue Signal Search seeks an experienced AI Kernel / Cluster Engineer to optimize production GPU and HPC clusters, focusing on Linux kernel level operations and driver management for reliable AI workloads.

You will collaborate with AI and infrastructure teams to ensure high availability, performance, and scalability using SLURM/Kubernetes, NVIDIA tools, and robust automation with Bash and Python.

Qualifications

  • Minimum 3 years managing production GPU or HPC clusters.
  • Advanced Ubuntu Linux administration experience with kernel-level expertise (driver management, tuning, boot config).
  • Production LXC container administration experience.
  • GPU networking expertise including InfiniBand and RoCEv2.
  • NVIDIA tools (DCGM, UFM, CUDA, NCCL, SHARP) proficiency.
  • Experience with SLURM or Kubernetes for workload orchestration.
  • Bash/Python automation for diagnostics and operations.

Responsibilities

  • Manage and optimize production GPU/HPC clusters for AI workloads.
  • Administer Ubuntu Linux at the kernel level, including drivers and tuning.
  • Monitor GPU infrastructure health using NVIDIA tools and resolve issues.
  • Configure and maintain LXC containers for reliable isolation and networking.
  • Support high-performance cluster networking (InfiniBand, RoCEv2).
  • Manage workloads with SLURM or Kubernetes to maximize efficiency.
  • Develop Bash/Python automation to streamline operations.
  • Troubleshoot Linux, GPU, containers, storage, and networking; document solutions.

Skills

Linux troubleshooting
Analytical thinking
Strong communication
Automation scripting (Bash/Python)

Tools

NVIDIA DCGM
NVIDIA UFM
CUDA
NCCL
SHARP
SLURM
Kubernetes
InfiniBand
RoCEv2
LXC
Ubuntu kernel tuning

Job description

Our client is expanding a world-class AI infrastructure environment designed to power next generation machine learning and high performance computing workloads. They are seeking an experienced AI Kernel / Cluster Engineer who thrives in highly technical Linux environments and enjoys solving complex infrastructure challenges at scale. This position offers the opportunity to work directly with cutting edge GPU computing platforms, optimize mission critical AI infrastructure, and help ensure the reliability and performance of production systems supporting advanced AI initiatives.

This is an excellent opportunity for an engineer who enjoys working at the intersection of Linux systems engineering, GPU infrastructure, networking, and automation while partnering with highly skilled technical teams to support innovative AI platforms.

What You’ll Do
  • Manage and optimize production GPU and HPC clusters supporting large scale AI training and inference workloads.
  • Administer Ubuntu Linux systems at the kernel level, including driver management, kernel tuning, boot configuration, and operating system optimization.
  • Monitor GPU infrastructure health using NVIDIA tools and proactively resolve performance, hardware, and system issues.
  • Configure and maintain LXC container environments, ensuring reliable resource allocation, networking, and workload isolation.
  • Support high performance cluster networking, including InfiniBand and RoCEv2, while diagnosing connectivity and fabric performance issues.
  • Manage workload scheduling and compute resources using SLURM or Kubernetes to maximize cluster efficiency.
  • Develop Bash and Python automation to streamline operations, improve monitoring, and reduce manual administrative tasks.
  • Troubleshoot complex issues spanning Linux systems, GPU hardware, containers, storage, and networking while documenting solutions and operational best practices.
Required Qualifications
  • Minimum of 3 years managing production GPU or HPC cluster environments.
  • Advanced Ubuntu Linux administration experience with strong kernel level expertise, including driver management, kernel tuning, boot configuration, and operating system internals.
  • Experience administering LXC container environments in production.
  • Strong understanding of GPU networking technologies including InfiniBand and RoCEv2.
  • Hands‑on experience with NVIDIA technologies including DCGM, UFM, CUDA, NCCL, and SHARP.
  • Experience supporting workload orchestration through SLURM or Kubernetes.
  • Proficiency with Bash and Python scripting for automation, diagnostics, and operational efficiency.
  • Demonstrated ability to troubleshoot issues across Linux systems, GPU infrastructure, container platforms, and networking simultaneously.
  • Strong analytical and problem solving abilities with excellent communication skills.
Preferred Qualifications
  • Experience supporting AI and machine learning frameworks including PyTorch, TensorFlow, or JAX.
  • Familiarity with infrastructure automation using Ansible or Terraform.
  • Exposure to TensorRT, ONNX, or related AI optimization technologies.
  • Industry certifications in Linux, networking, AI, cloud, or GPU technologies, including NVIDIA DLI, AWS Machine Learning, CCNP, or similar credentials.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA AI • Indiana (PA)

On-site
USD 150,000 - 190,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
AI Operations & Infrastructure Engineer
AI Operations & Infrastructure Engineer

Invictus International • Geraghty Village (MD)

On-site
USD 100,000 - 130,000
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
AI Operations & Infrastructure Engineer
AI Operations & Infrastructure Engineer

Invictus International Consulting, LLC • Fort Meade (MD)

On-site
USD 100,000 - 130,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000