Software Engineer - GPU Fleet

Iceberg

New York (NY)

On-site

USD 120,000 - 170,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Iceberg is seeking a Software Engineer to join its high-performance GPU infra team in New York. You will build tooling for fleet management, monitoring and network configuration, and work across the stack from applications to kernel drivers to diagnose and resolve issues.

You will own GPU fleet automation, workload tuning, and data-driven optimization to improve utilization, with a focus on observability and scalable infrastructure in a rapidly growing environment.

Qualifications

  • 2+ years of relevant software engineering experience, including strong Python development.
  • Hands-on experience managing and troubleshooting GPU infrastructure.
  • Strong CS fundamentals and sound software design instincts.
  • Solid Linux/UNIX experience and comfort with open-source software.
  • Strong debugging skills and a methodical approach to problem solving.
  • Experience with configuration management and monitoring technologies.
  • BS/MS in Computer Science or a related field.

Responsibilities

  • Build and maintain tooling for GPU fleet management, monitoring, metrics, maintenance and network configuration.
  • Troubleshoot issues across applications, networking, Linux, drivers and the kernel.
  • Work with engineering teams to tune workloads and improve GPU utilization.
  • Analyse GPU job data to identify trends, inefficiencies and opportunities for improvement.
  • Automate manual, slow or error-prone processes.

Skills

Python development
GPU infrastructure troubleshooting
CS fundamentals
Linux/UNIX experience
Debugging skills
Configuration management & monitoring
Software engineering

Education

BS/MS in Computer Science or related field

Job description

I'm working with a leading quantitative trading firm whose technology infrastructure is a major part of its competitive edge. The engineering environment is highly technical, with teams working across low-latency systems, hardware acceleration, machine learning and large-scale compute.

They're now looking for a Software Engineer to work on their rapidly growing GPU fleet - building the tooling that keeps the infrastructure healthy, observable and efficient.

You’ll have ownership across GPU management, automation, monitoring, metrics, maintenance and network configuration. The role sits close to the infrastructure, so when something goes wrong anywhere from an application down to the OS, drivers or kernel, you’ll be involved in figuring out why.

The work is varied by design. One day you could be building fleet automation, the next investigating a stubborn driver issue, and the next analyzing GPU workloads to find opportunities to improve utilization.

What you’ll be doing
  • Build and maintain tooling for GPU fleet management, monitoring, metrics, maintenance and network configuration.
  • Troubleshoot issues across applications, networking, Linux, drivers and the kernel.
  • Work with engineering teams to tune workloads and improve GPU utilization.
  • Analyse GPU job data to identify trends, inefficiencies and opportunities for improvement.
  • Automate manual, slow or error-prone processes.
What we’re looking for
  • 2+ years of relevant software engineering experience, including strong Python development.
  • Hands‑on experience managing and troubleshooting GPU infrastructure.
  • Strong CS fundamentals and sound software design instincts.
  • Solid Linux/UNIX experience and comfort with open-source software.
  • Strong debugging skills and a methodical approach to problem solving.
  • Experience with configuration management and monitoring technologies.
  • BS/MS in Computer Science or a related field.

CI/CD experience would be a plus.

The opportunity:

a technical infrastructure role where you’ll be working directly with a large and rapidly expanding GPU environment, solving problems across the stack and having a tangible impact on how efficiently the firm’s compute infrastructure operates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Fleet Software Engineer — Automation & Monitoring
GPU Fleet Software Engineer — Automation & Monitoring

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous paid time off
Wellness reimbursements
+2
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Learning and growth opportunities
Visa sponsorship and relocation assistance
+1
GPU Systems Engineer
GPU Systems Engineer

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Distinguished Engineer, GPU Fleet Operations Automation, Distinguished Engineer, GPU Fleet Oper[...]
Distinguished Engineer, GPU Fleet Operations Automation, Distinguished Engineer, GPU Fleet Oper[...]

NVIDIA • New York (NY)

On-site
USD 308,000 - 472,000
Equity
Benefits
Software Engineer, GPU Fleet
Software Engineer, GPU Fleet

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous paid time off
Wellness reimbursements
+2
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Software Engineer (DCIE)
Senior Software Engineer (DCIE)

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 300,000
Health benefits
Paid time off
401(k) match
+1
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Distinguished Engineer, GPU Fleet Operations Automation, Distinguished Engineer, GPU Fleet Oper[...]
Distinguished Engineer, GPU Fleet Operations Automation, Distinguished Engineer, GPU Fleet Oper[...]

NVIDIA • Town of Texas (WI)

On-site
USD 320,000 - 489,000
Equity
Benefits