Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques in New York seeks an experienced infrastructure engineer to design, deploy, and scale large-scale GPU clusters for AI research. You will work across compute, storage, OS, and automation to support hundreds of petabytes and thousands of nodes.

You will profile GPU workloads, remove bottlenecks, and collaborate with researchers to translate findings into speedups. Expect to own end-to-end infrastructure projects from design through long-term support and vendor engagement.

Qualifications

  • 5+ years engineering large-scale Linux systems in HPC/AI or distributed infra.
  • Strong Linux fundamentals, including kernel-level troubleshooting.
  • Hands-on debugging of distributed GPU workloads with a strong mental model of GPU performance.
  • Experience with GPUDirect RDMA and data movement between GPUs and the network.
  • Python for automation; CUDA or C/C++ for reading, profiling, and debugging GPU code.
  • Familiarity with configuration tools such as Salt, Ansible, Puppet, or Chef.
  • Ability to diagnose cross-stack issues across hardware, OS, and network.
  • Clear communication with researchers, engineers, and vendors.

Responsibilities

  • Design, deploy, and scale distributed GPU clusters from hardware selection to production operation.
  • Profile GPU workloads and identify performance bottlenecks across compute, storage, and network.
  • Collaborate with researchers to benchmark workloads and implement speedups.
  • Build automation to operate thousands of nodes, including provisioning and self-healing.
  • Own infrastructure projects end-to-end from scope to long-term support.
  • Engage with vendors to root-cause complex hardware and software issues.

Skills

Linux systems
GPU workloads
Python
CUDA/C++
GPUDirect RDMA
Diagnostics
Communication

Tools

Salt
Ansible
Puppet
Chef

Job description

Career Techniques in New York seeks an experienced infrastructure engineer to design, deploy, and scale large-scale GPU clusters for AI research. You will work across compute, storage, OS, and automation to support hundreds of petabytes and thousands of nodes.

You will profile GPU workloads, remove bottlenecks, and collaborate with researchers to translate findings into speedups. Expect to own end-to-end infrastructure projects from design through long-term support and vendor engagement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Systems Engineer – Large-Scale AI & HPC
Senior GPU Systems Engineer – Large-Scale AI & HPC

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Lead GPU Systems Engineer - HPC & AI Infrastructure
Lead GPU Systems Engineer - HPC & AI Infrastructure

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Wellness programs
+2
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior HPC Architect — Lead GPU Compute & Scale
Senior HPC Architect — Lead GPU Compute & Scale

NVIDIA AI • Illinois

On-site
USD 184,000 - 357,000
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior GPU Systems Engineer: Scale AI Performance
Senior GPU Systems Engineer: Scale AI Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000