Staff GPU Inference QA & Reliability Lead

Engg

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Cerebras Systems in Sunnyvale, CA, seeks a Staff GPU Inference SDET to lead test automation for the GPU inference stack and multi-node clusters. You will design end-to-end release qualification, validate prefill, and ensure numerical correctness and performance stability.

You will build automated tests, integrate with CI/CD, and collaborate across engineering to drive reliability, fault injection, and observability using Prometheus and Grafana.

Qualifications

  • 8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer.
  • GPU & Cluster Infrastructure Expertise: Hands-on experience bringing up, provisioning, and validating multi-node GPU clusters (NVIDIA or AMD) across public cloud infrastructure or enterprise data center environments.
  • Inference Stack Knowledge: Deep understanding of LLM serving engines and distributed runtimes, including prefill vs. decode disaggregation, KV-cache management, and dynamic batching.
  • Automation & Scripting: Expert-level Python programming skills with extensive experience designing custom test automation frameworks, diagnostic tooling, and CI/CD integration.
  • Orchestration & Networking: Strong proficiency with container orchestration tools (e.g., Kubernetes, Slurm, Ray) and high-performance cluster interconnects (e.g., InfiniBand, RoCE, NCCL).
  • Failure Analysis & Debugging: Proven background in root-cause analysis across software/hardware boundaries, stress testing, and node failure simulation in distributed systems.

Responsibilities

  • Build GPU Release Qualification Systems: Design and implement automated test automation frameworks, regression gates, and release qualification pipelines for the complete GPU inference stack—spanning custom API services, model-serving workers, container runtimes, serving engines, driver stacks, and firmware.
  • Inference Serving & Workload Validation: Benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs. decode worker performance, continuous batching, prefix caching, KV-cache efficiency, and tensor/expert parallelism.
  • Performance & Performance Modeling Verification: Build automated workload replay and benchmarking tools to validate GPU performance models. Track critical serving metrics including TTFT, ITL, throughput, tail latency (P99), and capacity efficiency.
  • Numerical Correctness & Quality Gates: Build validation infrastructure to ensure model accuracy, precision stability (FP16/FP8/quantization), determinism, and output correctness across software updates, kernel fusions, and hardware revisions.
  • Fault Injection & Fleet Resilience: Engineer chaos engineering and fault-injection suites to simulate node failures, inter-node network degradation, GPU memory leaks, driver/firmware mismatches, and automated recovery paths for multi-node GPU clusters.
  • Observability & CI/CD Integration: Integrate automated test pipelines with telemetry tools to turn investigations into repeatable gates and continuous performance monitoring.

Skills

Python programming
Test automation
Kubernetes
Distributed systems
Root-cause analysis
CI/CD
Performance benchmarking
Networking

Tools

Prometheus
Grafana
PyTorch Profiler
NVTX
ROCm profilers
C++

Job description

Cerebras Systems in Sunnyvale, CA, seeks a Staff GPU Inference SDET to lead test automation for the GPU inference stack and multi-node clusters. You will design end-to-end release qualification, validate prefill, and ensure numerical correctness and performance stability.

You will build automated tests, integrate with CI/CD, and collaborate across engineering to drive reliability, fault injection, and observability using Prometheus and Grafana.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff GPU Inference QA & Reliability Lead
Staff GPU Inference QA & Reliability Lead

Cerebras • United States

Remote
USD 180,000 - 250,000
Staff GPU Inference SDET
Staff GPU Inference SDET

Cerebras • United States

Remote
USD 180,000 - 250,000
Staff GPU Inference SDET
Staff GPU Inference SDET

Engg • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior SDET – PCIe GPU Bring-Up & AI QA (Equity)
Senior SDET – PCIe GPU Bring-Up & AI QA (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits
Senior GPU Inference Systems Engineer
Senior GPU Inference Systems Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 280,000
Senior SDET: PCIe GPU Validation & AI-Driven Automation
Senior SDET: PCIe GPU Validation & AI-Driven Automation

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits
Senior GPU Inference Systems Engineer
Senior GPU Inference Systems Engineer

Cerebras Systems • California (MO)

On-site
USD 180,000 - 260,000
Senior SDET: AI Infrastructure & Distributed Systems Testing
Senior SDET: AI Infrastructure & Distributed Systems Testing

Cerebras • United States

Remote
USD 180,000 - 230,000
SDET Technical Lead - AI Inference Core
SDET Technical Lead - AI Inference Core

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000