Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Career Techniques in New York seeks an experienced infrastructure engineer to design, deploy, and scale large-scale GPU clusters for AI research. You will work across compute, storage, OS, and automation to support hundreds of petabytes and thousands of nodes.

You will profile GPU workloads, remove bottlenecks, and collaborate with researchers to translate findings into speedups. Expect to own end-to-end infrastructure projects from design through long-term support and vendor engagement.

Qualifications

  • 5+ years engineering large-scale Linux systems in HPC/AI or distributed infra.
  • Strong Linux fundamentals, including kernel-level troubleshooting.
  • Hands-on debugging of distributed GPU workloads with a strong mental model of GPU performance.
  • Experience with GPUDirect RDMA and data movement between GPUs and the network.
  • Python for automation; CUDA or C/C++ for reading, profiling, and debugging GPU code.
  • Familiarity with configuration tools such as Salt, Ansible, Puppet, or Chef.
  • Ability to diagnose cross-stack issues across hardware, OS, and network.
  • Clear communication with researchers, engineers, and vendors.

Responsibilities

  • Design, deploy, and scale distributed GPU clusters from hardware selection to production operation.
  • Profile GPU workloads and identify performance bottlenecks across compute, storage, and network.
  • Collaborate with researchers to benchmark workloads and implement speedups.
  • Build automation to operate thousands of nodes, including provisioning and self-healing.
  • Own infrastructure projects end-to-end from scope to long-term support.
  • Engage with vendors to root-cause complex hardware and software issues.

Skills

Linux systems
GPU workloads
Python
CUDA/C++
GPUDirect RDMA
Diagnostics
Communication

Tools

Salt
Ansible
Puppet
Chef

Job description

Career Techniques in New York seeks an experienced infrastructure engineer to design, deploy, and scale large-scale GPU clusters for AI research. You will work across compute, storage, OS, and automation to support hundreds of petabytes and thousands of nodes.

You will profile GPU workloads, remove bottlenecks, and collaborate with researchers to translate findings into speedups. Expect to own end-to-end infrastructure projects from design through long-term support and vendor engagement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Architect for AI Infra at Scale
Senior GPU Cluster Architect for AI Infra at Scale

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • New York (NY)

On-site
USD 184,000 - 356,500
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior HPC Architect: At-Scale GPU Deployment & Automation
Senior HPC Architect: At-Scale GPU Deployment & Automation

NVIDIA • New Mexico

On-site
USD 184,000 - 287,500
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU System Architect: Multi-GPU AI/HPC Scale-Out
Senior GPU System Architect: Multi-GPU AI/HPC Scale-Out

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity compensation
Benefits package
Hybrid work model
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior GPU Data Center Architect for AI Infra
Senior GPU Data Center Architect for AI Infra

Applied Methods Ltd • Northern (KY), New York (NY)

Hybrid
USD 160,000 - 290,000
Bonus program
Equity
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000