Staff GPU Inference Engineer - Real-Time AI at Scale

Cerebras Systems

Toronto

On-site

CAD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cerebras Systems is hiring a Software Engineer to productionize and optimize our GPU serving stack across vLLM, PyTorch, ROCm, and rack-scale GPU infrastructure. You will write production-grade code, improve time to first token, throughput, and capacity efficiency, and drive reliability in a hands-on role spanning application, runtime, distributed systems, and hardware layers.

The role emphasizes deep debugging, numerical correctness, and automated workflows to make the serving path robust and

Qualifications

  • 8+ years of software engineering experience in production systems.
  • Experience with large language models or GPU workloads.
  • Strong programming in C++ and Python with concurrency.
  • Hands-on with a high-performance model-serving framework.
  • Understanding of GPU execution and profiling.
  • Experience debugging distributed systems.
  • Linux, containers, Kubernetes, observability, CI/CD.
  • Ability to design benchmarks and interpret results.

Responsibilities

  • Productionize GPU inference stack end-to-end.
  • Own GPU operational readiness and automation.
  • Improve production reliability and incident response.
  • Profile and optimize inference time, throughput, tail latency.
  • Debug across application to hardware layers.
  • Ensure numerical correctness and testing.
  • Build benchmarks and regression gates.

Skills

C++
Python
Multithreading
Performance optimization
Distributed systems
Linux

Education

Bachelor’s degree in CS/EE

Tools

vLLM
PyTorch
ROCm
TensorRT-LLM

Job description

Cerebras Systems is hiring a Software Engineer to productionize and optimize our GPU serving stack across vLLM, PyTorch, ROCm, and rack-scale GPU infrastructure. You will write production-grade code, improve time to first token, throughput, and capacity efficiency, and drive reliability in a hands-on role spanning application, runtime, distributed systems, and hardware layers.

The role emphasizes deep debugging, numerical correctness, and automated workflows to make the serving path robust and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Toronto

On-site
CAD 150,000 - 210,000
Staff Software Engineer, Inference API
Staff Software Engineer, Inference API

Engg • Toronto

On-site
CAD 140,000 - 200,000
Staff AI Inference Systems Engineer
Staff AI Inference Systems Engineer

Engg • Toronto

On-site
CAD 140,000 - 200,000
Staff Software Engineer, Inference API
Staff Software Engineer, Inference API

Cerebras • Toronto

Hybrid
CAD 120,000 - 180,000
Sr. Staff Software Engineer, Inference Platform
Sr. Staff Software Engineer, Inference Platform

Cerebras Systems • Toronto

On-site
CAD 170,000 - 250,000
Sr. Staff Software Engineer, Inference Platform
Sr. Staff Software Engineer, Inference Platform

Cerebras • Toronto

On-site
CAD 140,000 - 210,000
Distributed Software Engineer
Distributed Software Engineer

Cerebras • Toronto

On-site
CAD 140,000 - 190,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems • Lower Sackville

On-site
CAD 120,000 - 180,000
Distributed Software Engineer
Distributed Software Engineer

Engg • Toronto

On-site
CAD 140,000 - 210,000
FPGA Engineer
FPGA Engineer

Cerebras Systems • Lower Sackville

On-site
CAD 100,000 - 170,000