Senior Inference ML Engineer: Sparse Attention, Pruning

Cerebras

United States

Ibrido

USD 150.000 - 190.000

Tempo pieno

6 giorni fa
Candidati tra i primi
Generatore di candidature

Trasforma questa posizione in un colloquio — un curriculum e una lettera di presentazione creati in base a ciò questo datore di lavoro sta cercando.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

0

Descrizione del lavoro

Cerebras Systems seeks a Senior Research Engineer to adapt language and vision models for our AI accelerator. You will work with ML researchers to prototype, validate, and optimize models on Cerebras hardware for end-to-end inference at scale.

You will explore speculative decoding, pruning/compression, sparse attention, and sparsity-driven techniques to deliver low latency and high throughput. Collaborate across teams to push the frontier of inference.

Competenze

  • Combination of advanced degrees and deep ML software experience.
  • Experience in implementing and evaluating transformer-based models.
  • Demonstrated ability to translate research into production‑quality code.
  • Strong background in performance optimization on specialized hardware.

Mansioni

  • Design, implement, and optimize transformer architectures on Cerebras hardware.
  • Research and prototype inference algorithms and novel model architectures.
  • Train models to convergence, run hyperparameter sweeps, analyze results.
  • Bring up models on Cerebras system, validate correctness, debug issues.
  • Profile and optimize code to maximize throughput and minimize latency.
  • Develop tooling to surface bottlenecks and guide inference optimization.
  • Collaborate across software, hardware, and product teams from inception to delivery.

Conoscenze

Python/C++ programming
ML software development
Transformer models
Performance optimization
Research impact

Formazione

Bachelor's + 7+ years ML software
Master's + 4+ years software
PhD + 2+ years research
Equivalent practical experience

Strumenti

PyTorch
Transformers
vLLM
SGLang
Triton/CUDA

Descrizione del lavoro

Cerebras Systems seeks a Senior Research Engineer to adapt language and vision models for our AI accelerator. You will work with ML researchers to prototype, validate, and optimize models on Cerebras hardware for end-to-end inference at scale.

You will explore speculative decoding, pruning/compression, sparse attention, and sparsity-driven techniques to deliver low latency and high throughput. Collaborate across teams to push the frontier of inference.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior Research Engineer - Inference ML
Senior Research Engineer - Inference ML

Cerebras • Stati Uniti

In loco
USD 150.000 - 190.000
0
Principal ML Investigator - AI Systems on Fast Hardware
Principal ML Investigator - AI Systems on Fast Hardware

Cerebras • Sunnyvale (CA)

In loco
USD 250.000 - 450.000
Principal ML Investigator
Principal ML Investigator

Cerebras • Sunnyvale (CA)

In loco
USD 250.000 - 450.000
Principal ML Investigator
Principal ML Investigator

Cerebras Systems • Sunnyvale (CA)

In loco
USD 180.000 - 240.000
Director, AI Inference & Model Scaling
Director, AI Inference & Model Scaling

Cerebras • Sunnyvale (CA), Northern (KY)

Ibrido
USD 250.000 - 450.000
Director/Sr. Manager, AI Inference Model Scaling
Director/Sr. Manager, AI Inference Model Scaling

Cerebras • Sunnyvale (CA), Northern (KY)

In loco
USD 250.000 - 450.000
Applied Machine Learning Research Scientist
Applied Machine Learning Research Scientist

Cerebras • Stati Uniti

In loco
USD 120.000 - 160.000
Equal opportunity employer
Inclusive work environment
Continuous learning and growth opportunities
Advanced Technology: AI/ML Research Scientist
Advanced Technology: AI/ML Research Scientist

Cerebras Systems, Inc. • Sunnyvale (CA)

In loco
USD 120.000 - 160.000
Principal ML Investigator: AI Systems & Acceleration
Principal ML Investigator: AI Systems & Acceleration

Cerebras Systems • Sunnyvale (CA)

In loco
USD 180.000 - 240.000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Ibrido
USD 195.200 - 262.200
Health insurance
401(k) plan
Parental leave
+2