Senior GPU Inference Systems Engineer

Cerebras Systems

California (MO)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cerebras Systems is hiring a Software Engineer to productionize and optimize the GPU serving stack across custom APIs, vLLM, ROCm, and rack-scale GPU infrastructure.

You will write production code, establish operational practices, and drive improvements in time to first token, throughput, tail latency, and capacity efficiency. This hands-on role spans application, runtime, distributed systems, and hardware layers.

Qualifications

  • 8+ years of software engineering experience in production systems.
  • Experience with production inference systems for large language models or similar GPU workloads.
  • Strong programming ability in C++ and Python with multithreading and memory management.
  • Hands-on experience with a high-performance model-serving framework (vLLM, SGLang, TensorRT-LLM, Triton or equivalent).
  • Strong understanding of GPU execution, profiling, and asynchronous operations.
  • Experience debugging distributed systems across multiple layers.
  • Experience with Linux, containers, Kubernetes, CI/CD, and latency-sensitive services.
  • Ability to design benchmarks and interpret noisy results for production improvements.
  • Strong communication and technical leadership skills.
  • Bachelor’s degree in Computer Science/Engineering or equivalent.

Responsibilities

  • Productionize the GPU inference stack across API services, vLLM, ROCm, and hardware stack.
  • Own GPU operational readiness: deployment, upgrades, health checks, capacity management.
  • Drive reliability in production: define SLIs/OKRs, incident response, automated recovery.
  • Improve inference performance: reduce time to first token, increase throughput, lower tail latency.
  • Tune model-serving behavior: scheduling, continuous batching, KV-cache, parallelism, quantization.
  • Debug across system layers: diagnose failures across apps, libraries, drivers and hardware.
  • Ensure numerical correctness with validation and regression infrastructure.
  • Build benchmarks, profiling automation, and release gates.

Skills

C++
Python
Distributed systems
GPU/ML serving
Linux

Education

Bachelor's degree

Tools

vLLM
SGLang
TensorRT-LLM
Triton
ROCm

Job description

Cerebras Systems is hiring a Software Engineer to productionize and optimize the GPU serving stack across custom APIs, vLLM, ROCm, and rack-scale GPU infrastructure.

You will write production code, establish operational practices, and drive improvements in time to first token, throughput, tail latency, and capacity efficiency. This hands-on role spans application, runtime, distributed systems, and hardware layers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Inference Systems Engineer
Senior GPU Inference Systems Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 280,000
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 280,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras Systems • California (MO)

On-site
USD 180,000 - 260,000
Staff Software Engineer — Real-Time Inference Systems
Staff Software Engineer — Real-Time Inference Systems

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Comprehensive benefits
Senior GPU & Inference Systems Engineer
Senior GPU & Inference Systems Engineer

Poolside • United States

Remote
USD 140,000 - 200,000
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Senior System Software Engineer - GPU AI Inference Equity
Senior System Software Engineer - GPU AI Inference Equity

NVIDIA • United States

On-site
USD 152,000 - 241,500
Eligible for equity
Additional benefits
Senior System Software Engineer — GPU AI Inference (Triton)
Senior System Software Engineer — GPU AI Inference (Triton)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Equity
Benefits