Senior GPU Systems Engineer – Large-Scale AI & HPC

Iceberg

New York (NY)

On-site

USD 200,000 - 300,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Iceberg is seeking a GPU Systems Engineer in New York to design and scale GPU infrastructure across thousands of nodes and petabytes of storage. You will profile workloads, identify bottlenecks, and build automation to keep a massive fleet running with minimal human intervention.

You'll work across Linux, GPU infra, GPUDirect RDMA, CUDA, C/C++, Python, networking and distributed systems, with exposure to hardware selection and architecture from design to production.

Qualifications

  • 5+ years of experience with large-scale Linux systems in HPC/AI or distributed infrastructure environments.
  • Hands-on experience with GPU infrastructure design and performance profiling.
  • Strong fundamentals in CUDA, C/C++, and Python.

Responsibilities

  • Design and scale GPU clusters and storage infrastructure.
  • Profile workloads and optimize performance across the stack.
  • Develop automation to minimize manual intervention and outages.
  • Collaborate across hardware, software, and networking teams.

Skills

Linux systems
GPU infrastructure
CUDA
C/C++
Python
Profiling & optimization
Distributed systems
Networking
Performance tuning
Automation

Tools

GPUDirect RDMA
NVIDIA CUDA
NCCL
NVLink
Profiling tools

Job description

Iceberg is seeking a GPU Systems Engineer in New York to design and scale GPU infrastructure across thousands of nodes and petabytes of storage. You will profile workloads, identify bottlenecks, and build automation to keep a massive fleet running with minimal human intervention.

You'll work across Linux, GPU infra, GPUDirect RDMA, CUDA, C/C++, Python, networking and distributed systems, with exposure to hardware selection and architecture from design to production.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Lead GPU Systems Engineer - HPC & AI Infrastructure
Lead GPU Systems Engineer - HPC & AI Infrastructure

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Wellness programs
+2
GPU Fleet Engineer: Build Observability & Automation
GPU Fleet Engineer: Build Observability & Automation

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
GPU Systems Engineer
GPU Systems Engineer

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Senior GPU Systems Engineer: Scale AI Performance
Senior GPU Systems Engineer: Scale AI Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior HPC Architect — Lead GPU Compute & Scale
Senior HPC Architect — Lead GPU Compute & Scale

NVIDIA AI • Illinois

On-site
USD 184,000 - 357,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000