Text & Language Model Evaluation Engineer

S27a

San Mateo (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Deccan AI is seeking a Machine Learning Engineer to build and scale text and language-eval capabilities in the Bay Area. You will design evaluation pipelines, create prompt sets and rubrics, and run evals across frontier and open-source models, working with language experts and SMEs to surface linguistically meaningful failures.

You will own model-run infrastructure, scoring, and reproducible analysis, and help shape Deccan’s stance on multilingual evaluation and domain-specific benchmarks, in a

Qualifications

  • Build model-evaluation pipelines and deterministic/non-deterministic scoring systems.
  • Identify subtle model failures in language, reasoning, and domain contexts.
  • Produce fast studies and research reports on hot topics in ML evals.
  • Work with SMEs and language experts without outsourcing judgment.
  • Explain methodology, limitations, and findings to technical and non-technical audiences.

Responsibilities

  • Build and maintain evaluation pipelines for text, language, and domain-specific model failures.
  • Design fast benchmark workflows that deliver high-signal eval reports on a weekly cadence.
  • Create prompt sets, rubrics, failure taxonomies, judge/verifier logic, and benchmark charts.
  • Run evals across frontier and open-source models; analyze failure patterns with nuance for model teams.
  • Work with language experts, SMEs, delivery, and operations to collect prompts and surface meaningful failures.
  • Build tooling to speed eval production: model-run scripts, scoring pipelines, data validation, and reproducible notebooks.
  • Help define Deccan’s perspective on languages, evals, localized reasoning, and model failure modes.
  • Partner with GTM/Engagement Managers so findings support customer conversations without shallow marketing.
  • Train operators and analysts to run repeatable eval workflows while preserving quality.

Skills

Python programming
LLM evaluation
Benchmark design
NLP / ML engineering
Data pipelines
Communication

Tools

LLM tooling
Eval harnesses
Notebooks
ML tools

Job description

Deccan AI is seeking a Machine Learning Engineer to build and scale text and language-eval capabilities in the Bay Area. You will design evaluation pipelines, create prompt sets and rubrics, and run evals across frontier and open-source models, working with language experts and SMEs to surface linguistically meaningful failures.

You will own model-run infrastructure, scoring, and reproducible analysis, and help shape Deccan’s stance on multilingual evaluation and domain-specific benchmarks, in a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Mountain View, USA Machine Learning Engineer - Text and Evals
Mountain View, USA Machine Learning Engineer - Text and Evals

S27a • San Mateo (CA)

On-site
USD 140,000 - 190,000
Senior AI Model Evaluation & Systems Engineer
Senior AI Model Evaluation & Systems Engineer

Doist • United States

Remote
USD 140,000 - 210,000
Frontier Language Model Evaluation Engineer
Frontier Language Model Evaluation Engineer

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 230,000
Equity
Frontier AI exposure
Senior AI Systems Engineer — Model Evaluation
Senior AI Systems Engineer — Model Evaluation

Worky • California (MO)

On-site
USD 180,000 - 230,000
Senior AI Systems Engineer: Model Evaluation & QA
Senior AI Systems Engineer: Model Evaluation & QA

Deepgram • United States

Remote
USD 140,000 - 210,000
Frontline Language Model Engineer (Client-Facing)
Frontline Language Model Engineer (Client-Facing)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including **e
NLP Research Engineer: Multilingual Language Models
NLP Research Engineer: Multilingual Language Models

NLP PEOPLE • United States

Remote
USD 143,000 - 264,000
Stock options
RSUs program
Medical & dental coverage
+1
Senior Software Engineer - Model Evaluation & AI Systems
Senior Software Engineer - Model Evaluation & AI Systems

Worky • California (MO)

On-site
USD 180,000 - 230,000
Senior Software Engineer - Model Evaluation & AI Systems
Senior Software Engineer - Model Evaluation & AI Systems

Deepgram • United States

Remote
USD 140,000 - 210,000
AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home