Research Scientist: Efficient AI Inference

Bitdeer (NASDAQ: BTDR)

Austin (TX)

On-site

USD 140,000 - 220,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer AI Lab is seeking an ML/LLM inference engineer to optimize models for cost and speed. You will implement quantization, sparsity, and other efficiency techniques, and build robust evaluation to support performance claims.

Ideal candidates have strong Python/PyTorch skills, experience with C++, CUDA or Triton, and a track record in bringing efficiency methods to production or research depth. Austin-based role with growth in AI cloud infrastructure.

Qualifications

  • Hands-on experience in LLM inference, model optimization, or ML systems.
  • Strong programming ability in Python and deep familiarity with PyTorch; experience with C++, CUDA, or Triton is a plus.
  • Depth in at least one area of model efficiency such as quantization, sparsity and pruning, speculative decoding and MTP, or serving-time attention and KV-cache methods.
  • Understanding transformer internals and where accuracy loss shows up in model behavior.
  • Rigorous evaluation practices with task-level metrics, honest baselines, and clear statements about what numbers prove.
  • Experience taking efficiency methods into production serving or equivalent research depth.
  • Familiarity with inference engines like vLLM, SGLang, or TensorRT-LLM and their efficiency features.
  • Publications at top-tier venues or substantial opensource contributions are welcome.
  • Strong enthusiasm for cutting-edge AI infrastructure and ownership mentality.

Responsibilities

  • Make models cheaper and faster to serve without sacrificing important quality.
  • Implement and adapt published optimization methods on our models and hardware.
  • Build an evaluation discipline to defend claims like 2× throughput.
  • Bring production-ready efficiency methods into serving pipelines.

Skills

Python
PyTorch
C++
CUDA
Triton
Model efficiency
Transformer internals
Evaluation practice
Production experience
Inference engines
Publications/Open source
Ownership mindset

Education

Bachelor's/Master's/PhD in CS or Electrical Engineering or related field

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Bitdeer AI Lab is seeking an ML/LLM inference engineer to optimize models for cost and speed. You will implement quantization, sparsity, and other efficiency techniques, and build robust evaluation to support performance claims.

Ideal candidates have strong Python/PyTorch skills, experience with C++, CUDA or Triton, and a track record in bringing efficiency methods to production or research depth. Austin-based role with growth in AI cloud infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model Efficiency Research Scientist — Inference Optimization
Model Efficiency Research Scientist — Inference Optimization

Bitdeer Technologies Group • Austin (TX)

On-site
USD 120,000 - 230,000
Training & mentoring
Open workspaces
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer Technologies Group • Austin (TX)

On-site
USD 120,000 - 230,000
Training & mentoring
Open workspaces
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 220,000
RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 240,000
RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000
Senior Researcher, Efficient Inference for Production ML
Senior Researcher, Efficient Inference for Production ML

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000
LLM/VLM Inference Optimization Research Engineer
LLM/VLM Inference Optimization Research Engineer

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+1
GPU Kernel Architect for High-Performance AI Inference
GPU Kernel Architect for High-Performance AI Inference

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 120,000 - 180,000
Senior Quantized Inference Engineer - Accelerate LLMs
Senior Quantized Inference Engineer - Accelerate LLMs

NVIDIA AI • Redmond (WA)

On-site
USD 184,000 - 287,500
Equity
Benefits
LLM Inference Library Engineer - High-Performance AI
LLM Inference Library Engineer - High-Performance AI

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental