Member of Technical Staff, Model Routing

Dipp AI Technologies, Inc.

New York (NY)

On-site

USD 180,000 - 230,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Health benefits
401(k) match
Hardware budget
Compute credits
Publish methodology

Job summary

Dipp AI Technologies seeks a senior ML/engineering leader to own the dynamic model routing layer. You will classify tasks, govern evaluation across frontier, open-weight, and private models, and ensure budgets and hard stops are enforced before inference.

You will also govern model entry, shadow-test new models, and document decisions for auditability. The role demands hands-on production experience with LLM systems, strong ML/systems background, and fluency in Python with a systems language for

Qualifications

  • Hands-on production experience with LLM systems beyond prototypes.
  • Experience building evaluation infrastructure that teams trust.
  • Strong applied ML or systems background with rigorous measurement habits.

Responsibilities

  • Build task classification and routing policy across frontier, open-weight, and private models.
  • Own the evaluation harness that scores models per task class on quality, latency, and cost.
  • Implement pre-inference cost governance: budgets, ceilings, degradation paths, and hard stops.
  • Enforce enterprise data-boundary rules in routing so no restricted payload reaches an ineligible provider.
  • Build shadow evaluation and staged promotion for new models entering the fleet.
  • Instrument routing decisions so each one is explainable in the audit ledger.

Skills

LLM systems
ML engineering
Evaluation infrastructure
Python proficiency
Systems programming
Inference economics
Benchmark skepticism

Education

BS/MS/PhD in CS/ML/Statistics
Equivalent applied experience

Tools

LLM serving platforms
Evaluation tooling

Job description

Own dynamic model routing: the right model for each task, under a hard cost and policy ceiling.


Dipp AI's strategy is decoupled orchestration — no single frontier model, but the right model for each task under governed cost, latency, and data-boundary constraints. You will own the routing layer that makes that real.


You will build task classification, capability profiles, evaluation harnesses, and the cost governor that enforces spend ceilings before inference happens rather than after the invoice arrives. Routing decisions must be explainable and reproducible, because they end up in the audit ledger alongside the action they enabled.


You will also own how new models enter the fleet: how they are evaluated, gated, shadow-tested, and promoted or rejected on evidence.


What you will do


  • Build task classification and routing policy across frontier, open-weight, and private models.

  • Own the evaluation harness that scores models per task class on quality, latency, and cost.

  • Implement pre-inference cost governance: budgets, ceilings, degradation paths, and hard stops.

  • Enforce enterprise data-boundary rules in routing so no restricted payload reaches an ineligible provider.

  • Build shadow evaluation and staged promotion for new models entering the fleet.

  • Instrument routing decisions so each one is explainable in the audit ledger.


What we look for


  • Hands-on production experience with LLM systems beyond prototypes: serving, evaluation, and cost control.

  • Strong applied ML or systems background with rigorous measurement habits.

  • Experience building evaluation infrastructure that teams actually trusted.

  • Fluency in Python plus a systems language for the serving path.

  • Understanding of inference economics: tokens, context, caching, batching, and provider pricing behaviour.

  • Skepticism about benchmarks and the discipline to construct better ones.


Experience


  • 7+ years in software or ML engineering, including 2+ years shipping LLM-backed systems in production.

  • Has owned a routing, serving, or evaluation platform used by other teams.

  • Experience with self-hosted or private-tenant model deployment is a strong plus.


Education


  • BS, MS, or PhD in Computer Science, Machine Learning, Statistics, or a related field.

  • Equivalent applied experience shipping ML systems is accepted in place of a degree.


Compensation and benefits


  • Meaningful equity with a standard four-year vest.

  • Full medical, dental, and vision coverage.

  • 401(k) with company contribution, hardware budget, and generous compute for evaluation work.

  • Support for publishing evaluation methodology and results.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MTS, Post-Training (Enterprise)
MTS, Post-Training (Enterprise)

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 300,000 - 350,000
Location Mountain View, CA (Onsite)
Base salary $300,000–$350,000 USD/year
25% performance-based bonus
+3
Applied Scientist, Agent Evaluation & Adaptive Model Routing
Applied Scientist, Agent Evaluation & Adaptive Model Routing

Bitdeer Technologies Group • Austin (TX)

On-site
USD 140,000 - 210,000
ML Engineer, Post-Training
ML Engineer, Post-Training

Zoro • Northern (KY)

On-site
USD 120,000 - 180,000
Senior Data Scientist, NLP & LLM Fine-Tuning
Senior Data Scientist, NLP & LLM Fine-Tuning

Xenoss • New York (NY)

Remote
USD 160,000 - 210,000
Applied Scientist, Agent Evaluation & Adaptive Model Routing
Applied Scientist, Agent Evaluation & Adaptive Model Routing

Bitdeer • Austin (TX), Northern (KY)

On-site
USD 125,000 - 170,000
Mentoring program
Training opportunities
Competitive benefits
Software Engineer
Software Engineer

Emissary • San Francisco (CA)

On-site
USD 120,000 - 180,000
Engineering Manager, Provider Ecosystem Remote (US)
Engineering Manager, Provider Ecosystem Remote (US)

OpenRouter, Inc • Northern (KY)

Hybrid
USD 180,000 - 250,000
Restaurant example not provided
Engineering Manager, Provider Ecosystem
Engineering Manager, Provider Ecosystem

OpenRouter, Inc • New York (NY)

On-site
USD 180,000 - 260,000
Applied Scientist, Agent Evaluation & Adaptive Model Routing
Applied Scientist, Agent Evaluation & Adaptive Model Routing

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 150,000 - 230,000
Diversity culture
Open workspaces
Fast-growing environment
+4
ML Software Engineer
ML Software Engineer

Humble Robotics • San Francisco (CA)

On-site
USD 100,000 - 300,000