GPU Systems Engineer

Iceberg

New York (NY)

On-site

USD 200,000 - 300,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Iceberg is seeking a GPU Systems Engineer in New York to design and scale GPU infrastructure across thousands of nodes and petabytes of storage. You will profile workloads, identify bottlenecks, and build automation to keep a massive fleet running with minimal human intervention.

You'll work across Linux, GPU infra, GPUDirect RDMA, CUDA, C/C++, Python, networking and distributed systems, with exposure to hardware selection and architecture from design to production.

Qualifications

  • 5+ years of experience with large-scale Linux systems in HPC/AI or distributed infrastructure environments.
  • Hands-on experience with GPU infrastructure design and performance profiling.
  • Strong fundamentals in CUDA, C/C++, and Python.

Responsibilities

  • Design and scale GPU clusters and storage infrastructure.
  • Profile workloads and optimize performance across the stack.
  • Develop automation to minimize manual intervention and outages.
  • Collaborate across hardware, software, and networking teams.

Skills

Linux systems
GPU infrastructure
CUDA
C/C++
Python
Profiling & optimization
Distributed systems
Networking
Performance tuning
Automation

Tools

GPUDirect RDMA
NVIDIA CUDA
NCCL
NVLink
Profiling tools

Job description

New York | $200,000-$300,000 Base + Bonus

I'm working with one of the world's leading trading firms on an interesting GPU Systems Engineer position.

This is a hands-on infrastructure role working at scale, with large CPU and GPU clusters spanning thousands of nodes and hundreds of petabytes of storage.

You'll be designing and scaling GPU infrastructure, profiling workloads, finding performance bottlenecks and building the automation needed to keep a huge fleet running with minimal human intervention.

You'll be working across Linux, GPU infrastructure, GPUDirect RDMA, CUDA, C/C++, Python, networking and distributed systems, with the opportunity to get involved from hardware selection and architecture all the way through to production.

I'm looking for someone with 5+ years working with large-scale Linux systems in HPC, AI or distributed infrastructure environments who genuinely enjoys getting into the weeds when something isn't performing as it should.

Experience with NVIDIA, NCCL, NVLink or large-scale AI infrastructure would be a big plus.

$200k-$300k base + bonus.

If you're a systems engineer who likes working at scale and solving problems across hardware, software and networking, I'd love to speak.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Quant Hedge Fund - Systems Engineer - GPU expert
Quant Hedge Fund - Systems Engineer - GPU expert

Saragossa • New York (NY)

On-site
USD 180,000 - 300,000
Software Engineer - GPU Fleet
Software Engineer - GPU Fleet

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior HPC Architect, Automation and At-Scale Deployment
Senior HPC Architect, Automation and At-Scale Deployment

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Inclusive work environment
Comprehensive benefits
Senior Linux Kernel Systems Software Engineer – CSP Engagements
Senior Linux Kernel Systems Software Engineer – CSP Engagements

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Linux Kernel Systems Software Engineer – CSP Engagements
Senior Linux Kernel Systems Software Engineer – CSP Engagements

NVIDIA • Austin (TX)

On-site
USD 184,000 - 288,000
Equity options
Benefits package
Systems Engineer (GPU Virtualisation) - Systems Integrator
Systems Engineer (GPU Virtualisation) - Systems Integrator

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 170,000 - 300,000
Base salary $170K–$250K
Early-stage equity
On-site in San Francisco