Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras

United States

Remote

USD 150,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Cerebras Systems is hiring a Software Engineer to productionize and optimize our GPU serving stack, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, and rack-scale AMD GPU infrastructure.

This hands-on role requires deep debugging and optimization across application, runtime, distributed systems, and hardware layers to improve time to first token, throughput, tail latency, and capacity efficiency.

Qualifications

  • Experience building production-grade GPU serving systems.
  • Strong debugging and optimization across software and hardware.
  • Familiarity with model serving and distributed computation.

Responsibilities

  • Productionize the GPU inference stack. Design, build, deploy, and maintain the complete GPU prefill path, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • Own GPU operational readiness. Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for the AMD GPU fleet.
  • Drive reliability in production. Define service-level indicators and objectives for GPU-backed inference; improve fault isolation, automated recovery, incident response, and remediation.
  • Improve inference performance. Profile and optimize time to first token, throughput, tail latency, and GPU utilization.
  • Optimize model-serving behavior. Tune scheduling, batching, KV-cache management, parallelism, and distributed communication.
  • Debug across system layers. Diagnose failures across application code, vLLM, PyTorch, ROCm/HIP, and hardware.

Skills

GPU inference optimization
Debugging across system layers
Distributed systems
Python/C++ production code

Tools

PyTorch
vLLM
ROCm/HIP
GPU tooling

Job description

Cerebras Systems is hiring a Software Engineer to productionize and optimize our GPU serving stack, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, and rack-scale AMD GPU infrastructure.

This hands-on role requires deep debugging and optimization across application, runtime, distributed systems, and hardware layers to improve time to first token, throughput, tail latency, and capacity efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Inference Engineer for Real-Time AI
Senior GPU Inference Engineer for Real-Time AI

Cerebras • United States

On-site
USD 150,000 - 210,000
Software Engineer, GPU Inference
Software Engineer, GPU Inference

Cerebras • United States

On-site
USD 150,000 - 210,000
Real-Time GPU Optimization Engineer - Inference
Real-Time GPU Optimization Engineer - Inference

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Staff Software Engineer, Inference Cloud
Staff Software Engineer, Inference Cloud

Cerebras • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 200,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Staff GPU AI Software Engineer: Vision and ML Ops
Staff GPU AI Software Engineer: Vision and ML Ops

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Benefits at a glance
Staff Software Engineer - Real-Time AI Inference Infra
Staff Software Engineer - Real-Time AI Inference Infra

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 110,000 - 140,000
AI Inference Engineer – High-Performance GPU Systems
AI Inference Engineer – High-Performance GPU Systems

Perplexity • California (MO)

On-site
USD 120,000 - 170,000