Applied Scientist, Agent Evaluation & Adaptive Model Routing

Bitdeer

Austin, Northern (TX, KY)

Hybrid

USD 125,000 - 170,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Mentoring program
Training opportunities
Competitive benefits

Job summary

Bitdeer AI Lab in Austin is seeking a senior researcher to build evaluation pipelines for agentic inference and to extend our LLM and agent evaluation frameworks. You will design task-level and trajectory-level metrics, prototype routing strategies, and drive production-ready research artifacts.

The role requires deep expertise in Python, PyTorch, and evaluating multi-turn, tool-using agents, with a track record of robust experimental design and open-source contributions.

Qualifications

  • Hands-on experience in LLM evaluation and adaptive inference.
  • Strong Python programming and PyTorch experience; able to build reliable experimental pipelines.
  • Experience designing evaluation pipelines for LLMs or agents, including task-level and trajectory-level metrics.
  • Depth in model selection and routing, uncertainty estimation/calibration, cascading/escalation, or stage-aware inference.

Responsibilities

  • Build evaluation and decision systems for agentic inference; own LLM/agent evaluation pipeline.
  • Research and prototype adaptive model-routing strategies and escalation policies.
  • Collaborate with MaaS and platform teams on production integration.

Skills

Python programming
LLM evaluation
Agentic systems
Uncertainty estimation

Education

Bachelor’s/Master’s/PhD in CS/ML/Statistics/EE

Tools

PyTorch
vLLM
SGLang

Job description

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About Bitdeer AI Lab:

Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.

Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built at all. We work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens - and we continue to invest in and expand the infrastructure behind it

What you will be responsible for:
  • This role builds the evaluation and decision systems that make agentic inference measurable, reliable, and adaptive. You will own and extend our LLM and agent evaluation pipeline, develop representative task suites, and measure behavior at the trajectory level across task success, tool use, quality, cost, latency, and token consumption. You will research and prototype adaptive model-routing strategies inside a shared agent harness, including rule-based baselines, cascades, stage-aware routing, uncertainty-aware selection, and escalation and recovery policies. You will take promising methods from offline evaluation through traffic replay, shadow testing, and internal pilots, while working closely with the MaaS and platform engineering teams on production integration. The role owns evaluation methodology, routing policy, and research prototypes; production gateway reliability, billing, access control, and SLA engineering are collaborative responsibilities with the MaaS and platform teams
How you will stand out:
  • Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field, with substantial hands-on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference
  • Strong programming ability in Python and practical experience with PyTorch and modern data and evaluation tooling; able to independently build reliable experimental pipelines and internal research prototypes
  • Demonstrated experience designing or operating evaluation pipelines for LLMs or agents, including task-level and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis
  • In addition to hands-on LLM or agent evaluation experience, candidates should have implementation-level depth in at least one core area: model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware agent inference. Experience with preference modeling, contextual bandits, online learning, or broader adaptive inference methods is a plus.
  • Hands-on experience evaluating multi-turn or tool-using agents, including task completion, tool-call correctness, planning failures, recovery behavior, and cost and latency trade-offs
  • Rigorous experimental practice - controlled comparisons, honest baselines, statistical analysis, and the ability to explain clearly what an evaluation result does and does not prove
  • Ability to analyze per-request and per-trajectory quality, cost, latency, and token usage, and construct cost-quality Pareto frontiers rather than relying only on aggregate model scores
  • Experience with multi-model APIs, agent harnesses, traffic replay, shadow evaluation, A/B testing, or production model monitoring is highly preferred
  • Familiarity with model-specific differences in tool calling, context windows, prompt caching, reasoning controls, and inference systems such as vLLM or SGLang is a plus
  • Publications at top-tier ML, NLP, or systems venues, or substantial open-source contributions in evaluation, agents, routing, or inference, are welcome
  • Strong ownership and product judgment, with a track record of taking ambiguous research questions from problem definition through a working 0-to-1 prototype and measurable internal validation
What you will experience working with us:
  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer Technologies Group • Austin (TX)

On-site
USD 120,000 - 230,000
Training & mentoring
Open workspaces
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 220,000
Applied Scientist: Agent Evaluation & Adaptive Routing
Applied Scientist: Agent Evaluation & Adaptive Routing

Bitdeer • Austin (TX), Northern (KY)

Hybrid
USD 125,000 - 170,000
Mentoring program
Training opportunities
Competitive benefits
Research Kernel Engineer
Research Kernel Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 120,000 - 180,000
Agentic AI Engineer (Xora Portfolio Company)
Agentic AI Engineer (Xora Portfolio Company)

Xora Innovation • San Diego (CA)

Hybrid
USD 140,000 - 210,000
Research Scientist, Agentic Data & Benchmarking
Research Scientist, Agentic Data & Benchmarking

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 180,000 - 250,000
Senior AI Engineer
Senior AI Engineer

Quartile LLC • United States

On-site
USD 180,000 - 280,000
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

Hybrid
USD 140,000 - 190,000
Senior AI Engineer (Xora Portfolio Company)
Senior AI Engineer (Xora Portfolio Company)

Xora Innovation • San Diego (CA)

Hybrid
USD 150,000 - 210,000