ML Algorithm Mapping and Performance Engineer, Core ML

Cerebras

United States

Remote

USD 140,000 - 220,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cerebras Systems builds the world's largest AI chip and leads in fast ML training and inference. We seek an engineer to map advanced algorithms to Cerebras architecture, benchmark against GPUs, and understand trade-offs as models scale.

You will develop performance models, run experiments, and help push frontier-grade efficiency across our systems. You will combine analytical modeling with hands-on prototyping, spanning kernel-level to end-to-end performance, while evaluating new techniques and

Qualifications

  • Experience building analytical and empirical performance models.
  • Experience in ML training/inference workloads and performance analysis.
  • Ability to design and run benchmarks across large-scale models.

Responsibilities

  • Build analytical and empirical performance models for state-of-the-art ML training and inference algorithms.
  • Characterize asymptotic behavior and how trade-offs scale with model size, sequence length, batch size, parallelism, and hardware.
  • Construct Pareto frontiers across model quality, latency, throughput, memory, communication, and compute cost.
  • Develop prototype implementations and benchmarks for Cerebras WSE and relevant GPU/software baselines.
  • Analyze system behavior to identify kernel, compiler, runtime, communication, and algorithmic bottlenecks.
  • Evaluate emerging techniques in parallel token processing and related areas.

Skills

Performance modeling
Benchmarking
Prototyping
Kernel-level optimization

Tools

GPU baselines

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About The Role

The Core ML team develops novel algorithms for efficient large-scale training and inference. We are looking for an engineer who can determine how these algorithms should be mapped to the Cerebras architecture, when they outperform competing approaches, and how their advantages change as models, workloads, and hardware systems scale.

You will combine analytical performance modeling, empirical benchmarking, and hands-on prototyping to characterize the efficiency frontiers of emerging ML algorithms. Your work will span kernel-level and end-to-end performance, helping the team reason about trade-offs among model quality, latency, throughput, memory, communication, and compute utilization.

Responsibilities
  • Build analytical and empirical performance models for state-of-the‑art ML training and inference algorithms.
  • Characterize asymptotic behavior and identify how algorithmic trade‑offs change with model size, sequence length, batch size, parallelism, and hardware scale.
  • Construct Pareto frontiers across model quality, latency, throughput, memory footprint, communication, and compute cost.
  • Develop prototype implementations and benchmarks for the Cerebras WSE and relevant GPU or software baselines.
  • Analyze system behavior to identify kernel, compiler, runtime, communication, and algorithmic bottlenecks.
  • Evaluate emerging techniques in areas such as parallel token
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Algorithm Mapping and Performance Engineer, Core ML
ML Algorithm Mapping and Performance Engineer, Core ML

Cerebras Systems • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
ML Performance Engineer: Algorithm Mapping for AI Accelerators
ML Performance Engineer: Algorithm Mapping for AI Accelerators

Cerebras • United States

Remote
USD 140,000 - 220,000
ML Runtime and Kernel Engineer - Core ML
ML Runtime and Kernel Engineer - Core ML

Cerebras Systems • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
ML Inference Performance Architect
ML Inference Performance Architect

Foundation Capital • United States

On-site
USD 100,000 - 130,000
Advanced Technology: AI/ML Research Scientist
Advanced Technology: AI/ML Research Scientist

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip

Foundation Capital • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture
High-Performance ML Runtime & Kernel Engineer
High-Performance ML Runtime & Kernel Engineer

Cerebras Systems • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Kernel Engineer
Kernel Engineer

Foundation Capital • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture
ML Systems Performance Engineer — Hardware Co-Design
ML Systems Performance Engineer — Hardware Co-Design

Cerebras Systems • Sunnyvale (CA)

On-site
USD 180,000 - 240,000