Director, AI Inference & Model Scaling

Cerebras

Sunnyvale, Northern (CA, KY)

Hybrid

USD 250,000 - 450,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cerebras Systems is seeking an engineering leader to build and scale the Inference Model Scaling organization. You will define the technical vision and execution roadmap for a globally distributed team enabling latest foundation models on Cerebras hardware, leading ML model compilation and optimization and high-performance kernel development.

Work across compiler, runtime, cloud, hardware, product management, and AI research to shape the future of AI inference at Cerebras.

Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering or related field.
  • 12+ years building compiler, ML systems, or infrastructure software.
  • 5+ years leading engineering teams.
  • Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
  • Strong understanding of graph compilation and optimization.
  • Experience with Python and C++.
  • Experience delivering production-quality software.
  • Strong communication and cross-functional leadership skills.

Responsibilities

  • Define the technical roadmap and strategy for the team.
  • Establish technical direction across multiple teams and engineering leaders.
  • Lead design reviews and establish engineering standards.
  • Drive support for emerging LLM architectures and inference workloads.
  • Hire, mentor, and grow a high-performing engineering team.
  • Develop future technical leaders and managers.
  • Scale engineering processes while maintaining execution velocity.
  • Partner with Cloud Platform, ML, and Hardware teams in planning and delivering for end-to-end service enablement in Cloud and On-Premise settings.

Skills

Compiler infrastructure
ML/AI systems
Leadership
Python
C++
Distributed systems
Communication

Education

BS/MS/PhD in CS/CE

Tools

LLVM
MLIR
XLA
TVM
Torch FX

Job description

Cerebras Systems is seeking an engineering leader to build and scale the Inference Model Scaling organization. You will define the technical vision and execution roadmap for a globally distributed team enabling latest foundation models on Cerebras hardware, leading ML model compilation and optimization and high-performance kernel development.

Work across compiler, runtime, cloud, hardware, product management, and AI research to shape the future of AI inference at Cerebras.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director/Sr. Manager, AI Inference Model Scaling
Director/Sr. Manager, AI Inference Model Scaling

Cerebras • Sunnyvale (CA), Northern (KY)

Hybrid
USD 250,000 - 450,000
Principal Engineer, Inference Cloud
Principal Engineer, Inference Cloud

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Job stability with startup vitality
Open-source AI research
Simple, non-corporate work culture
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 200,000
Staff Software Engineer, Inference Cloud
Staff Software Engineer, Inference Cloud

Cerebras • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Senior Architect, Scaled AI Inference & Orchestration
Senior Architect, Scaled AI Inference & Orchestration

NVIDIA • California (MO)

On-site
USD 320,000 - 489,000
Equity
Benefits
Lead Principal Engineer, Inference Cloud Platform
Lead Principal Engineer, Inference Cloud Platform

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Opportunity to work on cutting-edge AI technology
Inclusive and supportive work environment
Job stability with startup vitality
Principal Engineer, Inference Cloud
Principal Engineer, Inference Cloud

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Opportunity to work on cutting-edge AI technology
Inclusive and supportive work environment
Job stability with startup vitality
Advanced Technology: AI/ML Research Scientist
Advanced Technology: AI/ML Research Scientist

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Senior Architect, Scaled AI Inference & Systems — Equity
Senior Architect, Scaled AI Inference & Systems — Equity

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
AI Inference Quality Engineer: Scale-Ready Model Performance
AI Inference Quality Engineer: Scale-Ready Model Performance

Cerebras • Sunnyvale (CA)

On-site
USD 120,000 - 160,000