Senior GPU Inference Engineer for Real-Time AI

Cerebras

United States

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cerebras Systems is seeking a Software Engineer to productionize and optimize our GPU serving stack. You will work across custom inference APIs, vLLM, ROCm, PyTorch, and AMD GPU infrastructure to ensure reliable, numerically correct, observable, and high-performance serving.

You will write production code, establish operational practices for a new accelerator fleet, and push improvements in time to first token, throughput, tail latency, and capacity efficiency.

Qualifications

  • 5+ years of software engineering experience with ownership of complex production systems.
  • Experience building or optimizing production inference systems for large GPU workloads.
  • Strong programming in C++ and Python, with multithreading and memory management.

Responsibilities

  • Productionize the GPU inference stack across APIs, vLLM, PyTorch, ROCm, and GPU infrastructure.
  • Own GPU operational readiness including deployment, upgrades, health checks, and rollback strategies.
  • Drive reliability with SLIs/SLAs, automated recovery, incident response, and post-incident remediation.

Skills

C++
Python
Distributed systems
Performance optimization
GPU inference

Education

Bachelor's degree in CS/CE/EE

Tools

vLLM
PyTorch
ROCm
TensorRT-LLM
Kubernetes

Job description

Cerebras Systems is seeking a Software Engineer to productionize and optimize our GPU serving stack. You will work across custom inference APIs, vLLM, ROCm, PyTorch, and AMD GPU infrastructure to ensure reliable, numerically correct, observable, and high-performance serving.

You will write production code, establish operational practices for a new accelerator fleet, and push improvements in time to first token, throughput, tail latency, and capacity efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
Software Engineer, GPU Inference
Software Engineer, GPU Inference

Cerebras • United States

On-site
USD 150,000 - 210,000
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Real-Time GPU Optimization Engineer - Inference
Real-Time GPU Optimization Engineer - Inference

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
AI Inference Engineer – High-Performance GPU Systems
AI Inference Engineer – High-Performance GPU Systems

Perplexity • California (MO)

On-site
USD 120,000 - 170,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Staff Software Engineer, Inference Cloud
Staff Software Engineer, Inference Cloud

Cerebras • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Senior GPU & Inference Systems Engineer
Senior GPU & Inference Systems Engineer

Poolside • United States

Remote
USD 140,000 - 200,000
Senior GPU AI Software Engineer — Kernels to Scale AI
Senior GPU AI Software Engineer — Kernels to Scale AI

Advanced Micro Devices, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 240,000