Senior Researcher, Efficient Inference for Production ML

MLSys 2020

San Francisco (CA)

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MLSys 2020 in San Francisco is looking for a Senior Researcher specialized in machine learning efficiency. The role involves designing and researching methods for quantization, speculative decoding, and other efficiency techniques, ensuring that research translates into production.

The ideal candidate will have over 5 years of experience in related fields, strong statistical expertise, and proficiency in tools like PyTorch or Jax. Join our small team for impactful work on autonomous research agents.

Qualifications

  • 5+ years of hands-on research experience in machine learning.
  • Strong understanding of training and inference performance.
  • Experience in translating research into production.

Responsibilities

  • Research and develop quantization methods.
  • Design and evaluate speculative decoding approaches.
  • Investigate training-time efficiency methods.
  • Run controlled experiments at production scale.
  • Co-design methods with inference engineering.

Skills

Machine Learning Research
Quantization
Speculative Decoding
Distillation
Efficient Algorithms
Statistical Expertise
Written Communication

Education

PhD in ML or related field

Tools

PyTorch
Jax

Job description

MLSys 2020 in San Francisco is looking for a Senior Researcher specialized in machine learning efficiency. The role involves designing and researching methods for quantization, speculative decoding, and other efficiency techniques, ensuring that research translates into production.

The ideal candidate will have over 5 years of experience in related fields, strong statistical expertise, and proficiency in tools like PyTorch or Jax. Join our small team for impactful work on autonomous research agents.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000
Research Scientist: Efficient AI Inference
Research Scientist: Efficient AI Inference

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 220,000
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Model Efficiency Research Scientist — AI Inference
Model Efficiency Research Scientist — AI Inference

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
Model Efficiency Research Scientist — Inference Optimization
Model Efficiency Research Scientist — Inference Optimization

Bitdeer Technologies Group • Austin (TX)

On-site
USD 120,000 - 230,000
Training & mentoring
Open workspaces
Research Member of Technical Staff- Efficient Modeling
Research Member of Technical Staff- Efficient Modeling

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
Staff Research Engineer, LLM Inference & Efficiency Remote
Staff Research Engineer, LLM Inference & Efficiency Remote

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
Parental leave top‑up
+3
Staff ML Research Engineer — From Prototype to Production
Staff ML Research Engineer — From Prototype to Production

Autonomous Technologies Group • New York (NY)

On-site
USD 110,000 - 150,000
Senior Applied Scientist — Efficient LLM Inference & Optimization
Senior Applied Scientist — Efficient LLM Inference & Optimization

Nebius • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4