Mountain View, USA Machine Learning Engineer - Text and Evals

S27a

San Mateo (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Deccan AI is seeking a Machine Learning Engineer to build and scale text and language-eval capabilities in the Bay Area. You will design evaluation pipelines, create prompt sets and rubrics, and run evals across frontier and open-source models, working with language experts and SMEs to surface linguistically meaningful failures.

You will own model-run infrastructure, scoring, and reproducible analysis, and help shape Deccan’s stance on multilingual evaluation and domain-specific benchmarks, in a

Qualifications

  • Build model-evaluation pipelines and deterministic/non-deterministic scoring systems.
  • Identify subtle model failures in language, reasoning, and domain contexts.
  • Produce fast studies and research reports on hot topics in ML evals.
  • Work with SMEs and language experts without outsourcing judgment.
  • Explain methodology, limitations, and findings to technical and non-technical audiences.

Responsibilities

  • Build and maintain evaluation pipelines for text, language, and domain-specific model failures.
  • Design fast benchmark workflows that deliver high-signal eval reports on a weekly cadence.
  • Create prompt sets, rubrics, failure taxonomies, judge/verifier logic, and benchmark charts.
  • Run evals across frontier and open-source models; analyze failure patterns with nuance for model teams.
  • Work with language experts, SMEs, delivery, and operations to collect prompts and surface meaningful failures.
  • Build tooling to speed eval production: model-run scripts, scoring pipelines, data validation, and reproducible notebooks.
  • Help define Deccan’s perspective on languages, evals, localized reasoning, and model failure modes.
  • Partner with GTM/Engagement Managers so findings support customer conversations without shallow marketing.
  • Train operators and analysts to run repeatable eval workflows while preserving quality.

Skills

Python programming
LLM evaluation
Benchmark design
NLP / ML engineering
Data pipelines
Communication

Tools

LLM tooling
Eval harnesses
Notebooks
ML tools

Job description

Machine Learning Engineer - Text and Evals

Location: Bay Area / San Mateo, CA
Employment Type: Full-Time
Department: Engineering

About Us

Deccan AI is a model training and evaluation startup headquartered in the Bay Area, with a delivery office in Hyderabad. We are founded by IIT Bombay, IIM Ahmedabad, and ex-Google alumni, and we work with leading AI labs and enterprise teams on high-quality human data, evaluations, and AI-first scaled operations.

We believe frontier AI systems are only as good as the data, evaluations, and human judgment behind them. Our work sits at the intersection of customer needs, research ambition, and operational execution. That makes the work high-touch, technical, ambiguous, and execution-heavy.

We are looking for people who can operate with urgency, write clearly, build trust, and turn ambiguity into outcomes.

About the Role

We are hiring a Machine Learning Engineer - Text and Evals to build Deccan’s fast-evals and text-evaluation capability.

Frontier models still fail in important text workflows, especially in languages, domains, and cultural contexts that are underrepresented in training data. Deccan is building a fast eval motion: short, high-signal benchmark studies that identify where models break, publish useful findings, and create a path into deeper training-data and model-improvement work.

What You’ll Do
  • Build and maintain evaluation pipelines for text, language, and domain-specific model failures.
  • Design fast benchmark workflows that can produce high-signal eval reports on a weekly or near-weekly cadence.
  • Create prompt sets, rubrics, failure taxonomies, judge / verifier logic, and benchmark charts that are technically defensible.
  • Run evals across frontier proprietary models and open-source models, then analyze failure patterns with enough nuance to be useful to model teams.
  • Work with language experts, SMEs, delivery, and operations to collect prompts, validate responses, and surface culturally or linguistically meaningful failures.
  • Build tooling that makes eval production faster: model-run scripts, scoring pipelines, data validation, report templates, and reproducible analysis notebooks or scripts.
  • Help define Deccan’s point of view on languages, evals, localized reasoning, code-mixing, domain-specific text workflows, and model failure modes.
  • Partner with GTM and Engagement Managers so eval findings can become credible customer conversations without turning the work into shallow marketing.
  • Train operators and analysts to run repeatable eval workflows without lowering the technical quality bar.
What You Will Own
  • Technical systems for Deccan’s text and language evaluation practice.
  • Prompt-set design, rubric structure, failure taxonomy, and scoring methodology.
  • Model-run infrastructure and reproducible benchmark analysis.
  • Judge, verifier, and quality-control logic for text evals.
  • Technical review of eval outputs before they are shared externally.
  • Playbooks that let Deccan repeat fast evals across languages, domains, and customer contexts.
  • The bridge between SME-generated content and model-facing evaluation rigor.
What We’re Looking For:
  • able build model-evaluation pipelines, deterministic and non-deterministic scoring systems, llm-judges, verifiers etc
  • able to identify subtle model failures in language, reasoning, instruction-following, domain knowledge, and cultural context;
  • produce fast studies and research like reports on relevant and hot topics
  • comfortable working with SMEs and language experts without outsourcing your judgment to them;
  • explain methodology, limitations, and findings to both technical and non-technical audiences.
Preferred Qualifications
  • 3+ years of experience in ML engineering, data science, model evaluation, NLP, applied AI, LLM tooling, or related work; exceptional earlier-career candidates with strong evidence of ability will also be considered.
  • Hands-on experience with LLM evaluation, benchmark design, prompt engineering, automated grading, rubric design, data quality, or model-failure analysis.
  • Strong programming ability in Python and modern ML / data tooling. Sharp at thinking algorithimically.
  • Experience with multilingual NLP, text classification, information extraction, instruction-following evals, safety evals, or domain-specific benchmarks is a plus.
  • Familiarity with frontier and open-source model APIs, eval harnesses, data pipelines, notebooks, and reproducible analysis workflows.
  • Strong written communication; you can describe methodology and failure examples without hiding behind vague metrics.
  • Experience working with language experts, SMEs, annotators, or human-in-the-loop data workflows.
Why Join Deccan AI

At Deccan AI, you will work close to the frontier of AI model training, evaluation, and scaled human-data operations. The best evals do more than rank models: they reveal where systems break, why they break, and what data or workflow could improve them.

This is a high-agency role for someone who wants to build that machine. You will help Deccan move quickly, publish useful findings, and turn text and language evals into a durable technical capability.

If you are excited by LLM evals, multilingual NLP, model failure analysis, and the chance to build a fast-evals practice from zero to one, we would like to meet you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Text & Language Model Evaluation Engineer
Text & Language Model Evaluation Engineer

S27a • San Mateo (CA)

On-site
USD 140,000 - 190,000
Mountain View, USA Founding Engineer - Reinforcement Learning
Mountain View, USA Founding Engineer - Reinforcement Learning

S27a • San Mateo (CA)

On-site
USD 150,000 - 210,000
Mountain View, USA Strategic Project Lead
Mountain View, USA Strategic Project Lead

S27a • San Mateo (CA)

On-site
USD 150,000 - 210,000
Senior Software Engineer - Model Evaluation & AI Systems
Senior Software Engineer - Model Evaluation & AI Systems

Worky • California (MO)

On-site
USD 180,000 - 230,000
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 230,000
Equity
Frontier AI exposure
Senior Software Engineer - Model Evaluation & AI Systems
Senior Software Engineer - Model Evaluation & AI Systems

Deepgram • United States

Remote
USD 140,000 - 210,000
Mountain View, USA Founding Engineer - Robotics
Mountain View, USA Founding Engineer - Robotics

S27a • San Mateo (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Software Engineer, Evaluation Platform / Infra
Software Engineer, Evaluation Platform / Infra

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Forward Deployed Engineer - Language Models
Forward Deployed Engineer - Language Models

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including **e