Model Efficiency Research Scientist — AI Inference

Bitdeer (NASDAQ: BTDR)

San Jose (CA)

On-site

USD 150,000 - 190,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer AI Lab is seeking a talented ML systems engineer to optimize large language model inference and deployment. You will push efficiency through quantization, sparsity, pruning, and related techniques, while maintaining accuracy and throughput.

You will collaborate on production-ready ML systems, apply advanced benchmarks, and contribute to cutting‑edge AI infrastructure within a fast-growing, globally engaged team.

Qualifications

  • Hands‑on experience in LLM inference, model optimization, or ML systems.
  • Strong programming ability in Python and deep familiarity with PyTorch.
  • Experience with inference engines and deployment considerations.
  • Deep understanding of transformer internals and model behavior.

Responsibilities

  • Make models cheaper and faster to serve without sacrificing quality.
  • Implement and adapt published efficiency methods and develop new optimizations.
  • Build evaluation discipline to defend performance claims.

Skills

LLM inference
Model optimization
ML systems
Python
PyTorch
C++
CUDA
Triton
Transformer internals
Evaluation discipline

Education

Bachelor’s/Master’s/PhD in CS/EE

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Bitdeer AI Lab is seeking a talented ML systems engineer to optimize large language model inference and deployment. You will push efficiency through quantization, sparsity, pruning, and related techniques, while maintaining accuracy and throughput.

You will collaborate on production-ready ML systems, apply advanced benchmarks, and contribute to cutting‑edge AI infrastructure within a fast-growing, globally engaged team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist: Efficient AI Inference
Research Scientist: Efficient AI Inference

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 220,000
Model Efficiency Research Scientist — Inference Optimization
Model Efficiency Research Scientist — Inference Optimization

Bitdeer Technologies Group • Austin (TX)

On-site
USD 120,000 - 230,000
Training & mentoring
Open workspaces
Remote Research Engineer: Real-Time AI Inference
Remote Research Engineer: Real-Time AI Inference

ElevenLabs • Maine

Hybrid
USD 140,000 - 190,000
Annual discretionary stipend
Annual company offsite
Co-working stipend
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer Technologies Group • Austin (TX)

On-site
USD 120,000 - 230,000
Training & mentoring
Open workspaces
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 220,000
Senior ML Infrastructure Engineer — High-Throughput AI Research
Senior ML Infrastructure Engineer — High-Throughput AI Research

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Senior Researcher, Efficient Inference for Production ML
Senior Researcher, Efficient Inference for Production ML

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000