Research Scientist: Efficient AI Inference

Bitdeer (NASDAQ: BTDR)

Austin (TX)

On-site

USD 140,000 - 220,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bitdeer AI Lab is seeking an ML/LLM inference engineer to optimize models for cost and speed. You will implement quantization, sparsity, and other efficiency techniques, and build robust evaluation to support performance claims.

Ideal candidates have strong Python/PyTorch skills, experience with C++, CUDA or Triton, and a track record in bringing efficiency methods to production or research depth. Austin-based role with growth in AI cloud infrastructure.

Qualifications

  • Hands-on experience in LLM inference, model optimization, or ML systems.
  • Strong programming ability in Python and deep familiarity with PyTorch; experience with C++, CUDA, or Triton is a plus.
  • Depth in at least one area of model efficiency such as quantization, sparsity and pruning, speculative decoding and MTP, or serving-time attention and KV-cache methods.
  • Understanding transformer internals and where accuracy loss shows up in model behavior.
  • Rigorous evaluation practices with task-level metrics, honest baselines, and clear statements about what numbers prove.
  • Experience taking efficiency methods into production serving or equivalent research depth.
  • Familiarity with inference engines like vLLM, SGLang, or TensorRT-LLM and their efficiency features.
  • Publications at top-tier venues or substantial opensource contributions are welcome.
  • Strong enthusiasm for cutting-edge AI infrastructure and ownership mentality.

Responsibilities

  • Make models cheaper and faster to serve without sacrificing important quality.
  • Implement and adapt published optimization methods on our models and hardware.
  • Build an evaluation discipline to defend claims like 2× throughput.
  • Bring production-ready efficiency methods into serving pipelines.

Skills

Python
PyTorch
C++
CUDA
Triton
Model efficiency
Transformer internals
Evaluation practice
Production experience
Inference engines
Publications/Open source
Ownership mindset

Education

Bachelor's/Master's/PhD in CS or Electrical Engineering or related field

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Bitdeer AI Lab is seeking an ML/LLM inference engineer to optimize models for cost and speed. You will implement quantization, sparsity, and other efficiency techniques, and build robust evaluation to support performance claims.

Ideal candidates have strong Python/PyTorch skills, experience with C++, CUDA or Triton, and a track record in bringing efficiency methods to production or research depth. Austin-based role with growth in AI cloud infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Efficiency Research Scientist — AI Inference
Model Efficiency Research Scientist — AI Inference

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
Adaptive AI Scientist: LLM Evaluation & Routing
Adaptive AI Scientist: LLM Evaluation & Routing

Bitdeer Technologies Group • Austin (TX)

On-site
USD 140,000 - 210,000
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 220,000
RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000
Applied Scientist: Agent Evaluation & Adaptive Routing
Applied Scientist: Agent Evaluation & Adaptive Routing

Bitdeer • Austin (TX), Northern (KY)

Hybrid
USD 125,000 - 170,000
Mentoring program
Training opportunities
Competitive benefits
Staff ML Engineer - Efficient Production Inference
Staff ML Engineer - Efficient Production Inference

ATBF Labs • San Francisco (CA)

Hybrid
USD 215,000 - 285,000
Equity
Health benefits
LLM Inference Performance Engineer: Speed & Efficiency
LLM Inference Performance Engineer: Speed & Efficiency

Baseten • United States

Remote
USD 150,000 - 210,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break