Applied Scientist, Agent Evaluation & Adaptive Model Routing

Bitdeer (NASDAQ: BTDR)

Austin (TX)

Presencial

USD 150.000 - 230.000

Jornada completa

14 días+
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Diversity culture
Open workspaces
Fast-growing environment
Impact on industry
Project exposure
Autonomy and growth
Training and mentoring

Descripción de la vacante

Bitdeer AI Lab seeks a capable researcher to advance evaluation frameworks for agentic inference and LLM systems. You will own the evaluation pipeline, develop task suites, and study routing strategies while coordinating with MaaS and platform teams to deploy reliable, cost-aware solutions.

The role emphasizes rigorous experimentation, cross-disciplinary collaboration, and contribution to the cutting edge of AI-powered mining and cloud platforms.

Formación

  • Advanced degrees or equivalent hands-on work in ML/AI research and evaluation.
  • Experience designing or operating evaluation pipelines for LLMs or agents.
  • Proven ability to build reliable experimental pipelines and research prototypes.

Responsabilidades

  • Build evaluation and decision systems for agentic inference; own and extend LLM evaluation pipeline; measure behavior across task success, tool use, cost, latency and tokens; prototype routing strategies; collaborate on production integration.

Conocimientos

Python
PyTorch
LLM evaluation
Agentic systems
Experiment design
Cost analysis
A/B testing

Educación

Bachelor’s/Master’s/PhD in CS/ML/EE

Herramientas

Multi-model APIs
Traffic replay
Shadow evaluation

Descripción del empleo

About Bitdeer

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

About Bitdeer

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About Bitdeer AI Lab

Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.

Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built at all. We work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens — and we continue to invest in and expand the infrastructure behind it

What You Will Be Responsible For
  • This role builds the evaluation and decision systems that make agentic inference measurable, reliable, and adaptive. You will own and extend our LLM and agent evaluation pipeline, develop representative task suites, and measure behavior at the trajectory level across task success, tool use, quality, cost, latency, and token consumption. You will research and prototype adaptive model-routing strategies inside a shared agent harness, including rule-based baselines, cascades, stage-aware routing, uncertainty-aware selection, and escalation and recovery policies. You will take promising methods from offline evaluation through traffic replay, shadow testing, and internal pilots, while working closely with the MaaS and platform engineering teams on production integration. The role owns evaluation methodology, routing policy, and research prototypes; production gateway reliability, billing, access control, and SLA engineering are collaborative responsibilities with the MaaS and platform teams
How You Will Stand Out
  • Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field, with substantial hands‑on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference
  • Strong programming ability in Python and practical experience with PyTorch and modern data and evaluation tooling; able to independently build reliable experimental pipelines and internal research prototypes
  • Demonstrated experience designing or operating evaluation pipelines for LLMs or agents, including task-level and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis
  • In addition to hands‑on LLM or agent evaluation experience, candidates should have implementation‑level depth in at least one core area: model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage‑aware agent inference. Experience with preference modeling, contextual bandits, online learning, or broader adaptive inference methods is a plus.
  • Hands‑on experience evaluating multi‑turn or tool‑using agents, including task completion, tool‑call correctness, planning failures, recovery behavior, and cost and latency trade‑offs
  • Rigorous experimental practice — controlled comparisons, honest baselines, statistical analysis, and the ability to explain clearly what an evaluation result does and does not prove
  • Ability to analyze per-request and per-trajectory quality, cost, latency, and token usage, and construct cost‑quality Pareto frontiers rather than relying only on aggregate model scores
  • Experience with multi‑model APIs, agent harnesses, traffic replay, shadow evaluation, A/B testing, or production model monitoring is highly preferred
  • Familiarity with model‑specific differences in tool calling, context windows, prompt caching, reasoning controls, and inference systems such as vLLM or SGLang is a plus
  • Publications at top‑tier ML, NLP, or systems venues, or substantial open‑source contributions in evaluation, agents, routing, or inference, are welcome
  • Strong ownership and product judgment, with a track record of taking ambiguous research questions from problem definition through a working 0-to-1 prototype and measurable internal validation
What You Will Experience Working With Us
  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start‑up spirit;
  • Fast‑growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Applied Scientist, Agent Evaluation & Adaptive Model Routing
Applied Scientist, Agent Evaluation & Adaptive Model Routing

Bitdeer • Austin (TX), Northern (KY)

Presencial
USD 125.000 - 170.000
Mentoring program
Training opportunities
Competitive benefits
Applied Scientist, Agent Evaluation & Adaptive Model Routing
Applied Scientist, Agent Evaluation & Adaptive Model Routing

Bitdeer Technologies Group • Austin (TX)

Presencial
USD 140.000 - 210.000
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

Presencial
USD 150.000 - 190.000
Research Scientist-Model Efficiency
Research Scientist-Model Efficiency

Bitdeer (NASDAQ: BTDR) • Austin (TX)

Presencial
USD 140.000 - 220.000
Applied Scientist: Adaptive Model Routing & Agent Evaluation
Applied Scientist: Adaptive Model Routing & Agent Evaluation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

Presencial
USD 150.000 - 230.000
Diversity culture
Open workspaces
Fast-growing environment
+4
Applied Scientist: Agent Evaluation & Adaptive Routing
Applied Scientist: Agent Evaluation & Adaptive Routing

Bitdeer • Austin (TX), Northern (KY)

Híbrido
USD 125.000 - 170.000
Mentoring program
Training opportunities
Competitive benefits
Adaptive AI Scientist: LLM Evaluation & Routing
Adaptive AI Scientist: LLM Evaluation & Routing

Bitdeer Technologies Group • Austin (TX)

Presencial
USD 140.000 - 210.000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

United States Digital Space LLC • San Francisco (CA)

Presencial
USD 190.000 - 230.000
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

Presencial
USD 140.000 - 190.000
Applied Research - Forward-Deployed
Applied Research - Forward-Deployed

Prime Intellect, Inc. • San Francisco (CA)

Híbrido
USD 150.000 - 300.000
Cash compensation range: $150-300k
Flexible work (San Francisco or hybrid
Visa sponsorship & relocation
+2