GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous PTO
Hybrid work
Free meals

Job summary

Tower Research Capital in New York seeks a software engineer to own the tooling behind day-to-day GPU operations—management, monitoring, metrics collection, maintenance, and network configuration.

You will troubleshoot issues from application to kernel, tune workloads for efficiency, and surface trends from GPU job statistics. The role offers a hybrid work setup with a globally distributed team and challenging, impact-driven projects.

Qualifications

  • BS or MS in computer science or a related field.
  • 2+ years of relevant experience, including Python development and hands-on GPU management.
  • Automation-first mindset. When a workflow is manual, slow, or error-prone, you reach for code.
  • Experience deploying, troubleshooting, and tuning a range of GPU hardware.
  • Strong computer science fundamentals and sound software design instincts.
  • Solid working knowledge of Linux/UNIX and comfort with open-source software.
  • Sharp debugging. You get to the bottom of problems quickly and methodically.
  • Familiarity with configuration management and monitoring technologies.
  • Organized, adaptable, and collaborative. You can juggle several tasks with careful attention to detail, work independently or with the team, and pick up new skills fast.

Responsibilities

  • Build and maintain the tooling that automates GPU fleet operations: management, monitoring, metrics collection, maintenance, and network configuration.
  • Troubleshoot software and hardware issues across the fleet, from application and network problems down to the operating system and kernel.
  • Work with engineering teams across the firm to tune workloads and processes so they use GPUs more efficiently.
  • Analyze GPU job statistics to surface trends, inefficiencies, and opportunities for improvement.

Skills

Python development
GPU management
Linux/UNIX
Debugging
CI/CD
Configuration management
Problem solving

Education

BS or MS in computer science

Tools

CUDA
Linux tooling

Job description

Tower Research Capital in New York seeks a software engineer to own the tooling behind day-to-day GPU operations—management, monitoring, metrics collection, maintenance, and network configuration.

You will troubleshoot issues from application to kernel, tune workloads for efficiency, and surface trends from GPU job statistics. The role offers a hybrid work setup with a globally distributed team and challenging, impact-driven projects.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Fleet Engineer: Build Observability & Automation
GPU Fleet Engineer: Build Observability & Automation

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
Software Engineer, GPU Fleet
Software Engineer, GPU Fleet

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
Software Engineer - GPU Fleet
Software Engineer - GPU Fleet

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
Senior Engineer, GPU Fleet Automation & Ops
Senior Engineer, GPU Fleet Automation & Ops

NVIDIA • New York (NY)

On-site
USD 308,000 - 472,000
Hybrid GPU Fleet Operations Engineer
Hybrid GPU Fleet Operations Engineer

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
GPU Compute Engineer — Fleet Reliability & Automation
GPU Compute Engineer — Fleet Reliability & Automation

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Engineering Manager, GPU Fleet & Infra (Hybrid)
Engineering Manager, GPU Fleet & Infra (Hybrid)

Socket.dev • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
401k plan with company match (USA)
GPU Systems Engineer
GPU Systems Engineer

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
GPU Systems Engineer - Low-Latency HPC & AI Clusters
GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4