Staff GPU Inference QA & Reliability Lead

Cerebras

United States

Remote

USD 180,000 - 250,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cerebras Systems is hiring a Staff GPU Inference SDET to lead the validation and reliability of our GPU inference stack and rack-scale systems. You will design automated test ecosystems for multi-node clusters, ensure numerical correctness, and drive production-grade reliability for high-speed inference workloads.

You will architect and implement validation pipelines spanning API services, serving workers, runtimes, and firmware, while measuring TTFT, ITL, throughput, and tail latency to guide

Responsibilities

  • Build GPU Release Qualification Systems: automated test frameworks, regression gates, and release pipelines for the GPU inference stack.
  • Inference Serving Workload Validation: benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs decode, batching, caching, and parallelism.
  • Performance Modeling Verification: automate workload replay/benchmarking to validate GPU performance models; track TTFT, ITL, throughput, and P99 latency.
  • Quality Gates: ensure model accuracy, precision stability (FP16/FP8/quantization), determinism, and output correctness across updates.
  • Fleet Resilience: chaos engineering and fault-injection for multi-node GPU clusters.
  • Observability: CI/CD Integration fidelity across the stack; monitor, alert, and report health

Job description

Cerebras Systems is hiring a Staff GPU Inference SDET to lead the validation and reliability of our GPU inference stack and rack-scale systems. You will design automated test ecosystems for multi-node clusters, ensure numerical correctness, and drive production-grade reliability for high-speed inference workloads.

You will architect and implement validation pipelines spanning API services, serving workers, runtimes, and firmware, while measuring TTFT, ITL, throughput, and tail latency to guide

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff GPU Inference QA & Reliability Lead
Staff GPU Inference QA & Reliability Lead

Engg • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Staff GPU Inference SDET
Staff GPU Inference SDET

Cerebras • United States

Remote
USD 180,000 - 250,000
Staff GPU Inference SDET
Staff GPU Inference SDET

Engg • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior GPU Inference Systems Engineer
Senior GPU Inference Systems Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 280,000
Senior GPU Inference Systems Engineer
Senior GPU Inference Systems Engineer

Cerebras Systems • California (MO)

On-site
USD 180,000 - 260,000
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
Senior SDET – PCIe GPU Bring-Up & AI QA (Equity)
Senior SDET – PCIe GPU Bring-Up & AI QA (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits
Senior SDET: PCIe GPU Validation & AI-Driven Automation
Senior SDET: PCIe GPU Validation & AI-Driven Automation

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits
SDET Technical Lead - AI Inference Core
SDET Technical Lead - AI Inference Core

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 280,000