Senior GPU & Inference Systems Engineer

Poolside

United States

Remote

USD 140,000 - 200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Poolside is seeking an engineer to optimize GPU utilization and deliver a stable, scalable inference serving stack for researchers. You will partner with the inference and scalability teams to improve throughput, latency, and fault tolerance across large-scale training environments.

You’ll collaborate closely with the infra team to ensure GPU nodes are healthy and fully utilized, enabling faster experimentation and pushing Poolside toward its AGI-focused mission.

Qualifications

  • Strong programming skills in Go, or other similar languages.
  • Strong systems engineering background: distributed systems, schedulers, control planes, or high-throughput data planes.
  • Production experience with Kubernetes internals - controllers, informers, operators - not just deploying to it.
  • Bias toward observability and debuggability: building a system that is easy to navigate when debugging production issues
  • Plus: experience in systems serving large scale inference requests

Responsibilities

  • Design and develop internal scheduling system to maximize GPU utilization
  • Build API and tooling to help manage the lifecycle of GPU workloads and troubleshoot failures
  • Design and improve inference control plane to speed up model deployment and inference request serving
  • Collaborate with research to improve research velocity continuously

Skills

Go
Distributed systems
Kubernetes
Observability
Large-scale inference

Tools

Kubernetes controllers
Go tooling

Job description

Poolside is seeking an engineer to optimize GPU utilization and deliver a stable, scalable inference serving stack for researchers. You will partner with the inference and scalability teams to improve throughput, latency, and fault tolerance across large-scale training environments.

You’ll collaborate closely with the infra team to ensure GPU nodes are healthy and fully utilized, enabling faster experimentation and pushing Poolside toward its AGI-focused mission.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Inference Engineer for Real-Time AI
Senior GPU Inference Engineer for Real-Time AI

Cerebras • United States

On-site
USD 150,000 - 210,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Senior GPU HPC Infrastructure Engineer for ML Pipelines
Senior GPU HPC Infrastructure Engineer for ML Pipelines

Jaide Health • United States

Hybrid
USD 180,000 - 250,000
Weekly lunch stipend
Health and dental benefits
Parental leave
+3
GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior GPU Systems Engineer – Large-Scale AI & HPC
Senior GPU Systems Engineer – Large-Scale AI & HPC

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1