Senior GPU Inference Systems Engineer

Cerebras

Sunnyvale (CA)

On-site

USD 180,000 - 280,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cerebras Systems is hiring a Software Engineer to productionize and optimize the GPU serving stack across custom inference APIs, vLLM, and ROCm-based infrastructure. This hands-on role emphasizes performance, observability, and reliability in a disaggregated AI inference environment.

You will work across application, runtime, distributed systems, and hardware layers to boost time-to-first-token, throughput, tail latency, and capacity efficiency on the Cerebras Wafer-Scale Engine.

Qualifications

  • Minimum 8+ years of software engineering experience in complex production systems.
  • Experience building/operating/optimizing production inference systems for large language or GPU workloads.
  • Strong programming in C++ and Python with attention to concurrency and performance.

Responsibilities

  • Produce and optimize the GPU inference stack from API services to model-serving workers.
  • Ensure GPU fleet operational readiness with deployment, upgrades and rollback strategies.
  • Define SLIs/SLAs and improve fault isolation, automated recovery, and incident response.

Skills

C++
Python
Multithreading
Distributed systems
GPU performance
Linux
Kubernetes
Benchmarks
Communication
Leadership

Education

Bachelor’s degree in CS/CE/EE

Tools

vLLM
TensorRT-LLM
Triton Inference Server
HIP/ROCm
Docker

Job description

Cerebras Systems is hiring a Software Engineer to productionize and optimize the GPU serving stack across custom inference APIs, vLLM, and ROCm-based infrastructure. This hands-on role emphasizes performance, observability, and reliability in a disaggregated AI inference environment.

You will work across application, runtime, distributed systems, and hardware layers to boost time-to-first-token, throughput, tail latency, and capacity efficiency on the Cerebras Wafer-Scale Engine.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Inference Systems Engineer
Senior GPU Inference Systems Engineer

Cerebras Systems • California (MO)

On-site
USD 180,000 - 260,000
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 280,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras Systems • California (MO)

On-site
USD 180,000 - 260,000
Staff Software Engineer — Real-Time Inference Systems
Staff Software Engineer — Real-Time Inference Systems

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff (Software Engineer)
Member of Technical Staff (Software Engineer)

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 110,000 - 140,000
Member of Technical Staff (Software Engineer)
Member of Technical Staff (Software Engineer)

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
New Grad Kernel Engineer for AI Hardware & HPC
New Grad Kernel Engineer for AI Hardware & HPC

Cerebras Systems • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Comprehensive benefits