Staff AI Inference Systems Engineer

Engg

Toronto

On-site

CAD 140,000 - 200,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Cerebras Systems is building a new generation of disaggregated AI inference systems. We are hiring a Software Engineer to evolve the ML API layer for a heterogeneous serving system across models and accelerator backends.

You will work across inference APIs, model integration, vLLM-based runtime, and Cerebras runtime components to deliver a consistent experience for developers and customers. This role is hands-on and focuses on production-quality APIs.

Qualifications

  • 5+ years of software engineering experience in production software or distributed systems.
  • Strong programming in Python and performance-sensitive services in C++/Go or similar.
  • Understanding of modern LLM inference concepts: tokenization, sampling, streaming, KV-cache, model configuration.
  • Experience integrating software across service, framework, runtime, and infrastructure boundaries.
  • Experience building stable APIs with validation, error handling, observability, compatibility, and versioning.

Responsibilities

  • Build production ML inference APIs.
  • Design, implement, and maintain APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs.
  • Deliver a unified serving experience across GPU prefill, Cerebras decode, and other backends.
  • Enable new models and capabilities by integrating foundation models, tokenizers, prompts, sampling, and model-specific features.
  • Own API compatibility and evolution with clear versioning, deprecation, validation, and backward compatibility.
  • Integrate with model-serving runtimes (vLLM, PyTorch, Hugging Face, ROCm, Cerebras runtime).
  • Support disaggregated inference: coordination of request routing, state transfer, errors, retries, lifecycle.
  • Improve performance: streaming, latency, throughput, batching, tokenization, scheduling.
  • Ensure correctness: validation for tokenization, sampling, logits, outputs, precision changes.
  • Strengthen reliability and observability: logging, tracing, metrics, dashboards, health checks.
  • Develop testing and qualification infrastructure: conformance tests, workload replay, model validation, performance benchmarks.
  • Improve developer experience: configuration, SDKs, docs, examples, debugging tools.
  • Collaborate across stack: translate model/customer requirements into scalable serving capabilities.

Skills

Python
C++/Go
Distributed systems
API design
Linux & Containers
Observability
Communication

Education

Bachelor's degree in CS/CE/EE

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
Hugging Face Text Generation Inference

Job description

Cerebras Systems is building a new generation of disaggregated AI inference systems. We are hiring a Software Engineer to evolve the ML API layer for a heterogeneous serving system across models and accelerator backends.

You will work across inference APIs, model integration, vLLM-based runtime, and Cerebras runtime components to deliver a consistent experience for developers and customers. This role is hands-on and focuses on production-quality APIs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Inference API
Staff Software Engineer, Inference API

Cerebras • Toronto

Hybrid
CAD 120,000 - 180,000
Staff Software Engineer, Inference API
Staff Software Engineer, Inference API

Engg • Toronto

On-site
CAD 140,000 - 200,000
Sr. Staff Software Engineer, Inference Platform
Sr. Staff Software Engineer, Inference Platform

Cerebras Systems • Toronto

On-site
CAD 170,000 - 250,000
Sr. Staff Software Engineer, Inference Platform
Sr. Staff Software Engineer, Inference Platform

Cerebras • Toronto

On-site
CAD 140,000 - 210,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras Systems • Lower Sackville

On-site
CAD 167,000 - 278,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Toronto

On-site
CAD 110,000 - 180,000
Full Stack LLM Engineer
Full Stack LLM Engineer

Cerebras • Toronto

On-site
CAD 80,000 - 120,000
Competitive salary and benefits package
Opportunities for professional growth
Dynamic and innovative work environment
AI Inference Core - Senior SW Engineer for Platform & DevOps
AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras • Toronto

On-site
CAD 120,000 - 180,000
AI Hardware Systems Integration Engineer
AI Hardware Systems Integration Engineer

Cerebras Systems • Lower Sackville

On-site
CAD 167,000 - 278,000
Staff GPU Inference Engineer - Real-Time AI at Scale
Staff GPU Inference Engineer - Real-Time AI at Scale

Cerebras Systems • Toronto

On-site
CAD 180,000 - 240,000