GPU Fleet Engineer: Build Observability & Automation

Iceberg

New York (NY)

On-site

USD 120,000 - 170,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Iceberg is seeking a Software Engineer to join its high-performance GPU infra team in New York. You will build tooling for fleet management, monitoring and network configuration, and work across the stack from applications to kernel drivers to diagnose and resolve issues.

You will own GPU fleet automation, workload tuning, and data-driven optimization to improve utilization, with a focus on observability and scalable infrastructure in a rapidly growing environment.

Qualifications

  • 2+ years of relevant software engineering experience, including strong Python development.
  • Hands-on experience managing and troubleshooting GPU infrastructure.
  • Strong CS fundamentals and sound software design instincts.
  • Solid Linux/UNIX experience and comfort with open-source software.
  • Strong debugging skills and a methodical approach to problem solving.
  • Experience with configuration management and monitoring technologies.
  • BS/MS in Computer Science or a related field.

Responsibilities

  • Build and maintain tooling for GPU fleet management, monitoring, metrics, maintenance and network configuration.
  • Troubleshoot issues across applications, networking, Linux, drivers and the kernel.
  • Work with engineering teams to tune workloads and improve GPU utilization.
  • Analyse GPU job data to identify trends, inefficiencies and opportunities for improvement.
  • Automate manual, slow or error-prone processes.

Skills

Python development
GPU infrastructure troubleshooting
CS fundamentals
Linux/UNIX experience
Debugging skills
Configuration management & monitoring
Software engineering

Education

BS/MS in Computer Science or related field

Job description

Iceberg is seeking a Software Engineer to join its high-performance GPU infra team in New York. You will build tooling for fleet management, monitoring and network configuration, and work across the stack from applications to kernel drivers to diagnose and resolve issues.

You will own GPU fleet automation, workload tuning, and data-driven optimization to improve utilization, with a focus on observability and scalable infrastructure in a rapidly growing environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Fleet Automation Engineer (Hybrid)
GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
Senior GPU Systems Engineer – Large-Scale AI & HPC
Senior GPU Systems Engineer – Large-Scale AI & HPC

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Software Engineer - GPU Fleet
Software Engineer - GPU Fleet

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Senior Observability Product Manager, GPU Fleet
Senior Observability Product Manager, GPU Fleet

Nscale • United States

On-site
USD 200,000 - 280,000
Competitive benefits package
Flexible paid time off
Parental leave
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
GPU Infrastructure Engineer - Scale & Automation
GPU Infrastructure Engineer - Scale & Automation

United States Digital Space LLC • United States

Remote
USD 150,000 - 210,000
Senior Software Engineer: Fleet Intelligence & GPU Telemetry
Senior Software Engineer: Fleet Intelligence & GPU Telemetry

NVIDIA • New York (NY)

On-site
USD 152,000 - 242,000
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
GPU Infrastructure Software Engineer I
GPU Infrastructure Software Engineer I

crusoe • San Francisco (CA)

On-site
USD 117,000 - 135,000
Health insurance
401(k) with match
Stock options/RSUs
+3