ML Evaluation Engineer

Mployee.me

Bengaluru

On-site

INR 2,400,000 - 4,000,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Lunch provided at the office
Flexible working hours
Comprehensive health insurance for you
Zomato meal benefits
Remote Wednesdays

Job summary

Triomics in Bengaluru is seeking an ML Evaluation Engineer to own model quality, regression testing, release validation, and production impact analysis for clinical AI systems. This IC role sits between applied ML, clinical data, MLOps, and production operations, ensuring every release is measurable and stable.

You will maintain evaluation datasets, create hidden test sets, run regression checks, and produce release-readiness reports to guide critical deployment decisions.

Qualifications

  • 3–6+ years in ML engineering or related field.
  • Strong Python and data analysis skills.
  • Experience with evaluation pipelines, regression frameworks, and dashboards.

Responsibilities

  • Build and maintain evaluation frameworks for clinical NLP, LLM, RAG, and information extraction systems.
  • Create hidden test datasets to reduce overfitting and ensure robust releases.
  • Define release metrics, regression thresholds, and coverage across data sources.
  • Compare model versions and document performance across segments and data types.
  • Collaborate with ML, data, and MLOps teams to monitor production behavior.
  • Produce release-readiness reports before deployment.

Skills

Python
Data analysis
ML QA
NLP evaluation
Statistical testing

Tools

MLflow
Weights & Biases
Great Expectations
pytest
Airflow
Prefect

Job description

Model Quality · Release Validation · Production Impact Analysis

Location: On-site, HSR Layout, Bengaluru

Level: Mid/Senior Individual Contributor / Lead-track IC

About Triomics

Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained staff manually reading unstructured patient records - pathology reports, clinical notes, genomic panels - to power workflows like trial matching, registry curation, visit prep, and quality reporting. We replace that manual work with task-driven AI agents that sit inside the EMR and process records at scale, in real time.

Our platform is trusted by leading cancer centers including Memorial Sloan Kettering, Mount Sinai, and Yale Cancer Center. We have grown 10x in the last year and process millions of oncology medical documents monthly.

Our investors include Battery Ventures, Lightspeed, General Catalyst, Nexus Venture Partners, and Y Combinator.

GROWTH PATH

This is an individual contributor role with strong ownership expectations. High performers may be considered for workstream lead or functional lead responsibilities after approximately 12 months, based on demonstrated ownership, delivery, technical judgment, mentoring, cross-functional influence, and ability to reduce dependency on the Director of ML.

ABOUT THE ROLE

We are looking for an ML Evaluation Engineer to own model quality, regression testing, release validation, and production impact analysis for clinical AI systems. This role sits between applied ML, clinical data, MLOps, and production operations.

Your job is to ensure that every model or workflow release is measurable, stable, and not degrading important clinical behavior. You will maintain evaluation datasets, create hidden test sets, run regression checks, analyze production issues, and produce release-readiness reports.

WHAT YOU WILL DO
  • Build and maintain evaluation frameworks for clinical NLP, LLM, RAG, information extraction, and structured abstraction systems.
  • Create and manage hidden test datasets that are not directly visible to model developers, reducing overfitting risk.
  • Define release metrics, regression thresholds, slice-based evaluation, failure-mode tracking, and release/blocker criteria.
  • Compare model versions and identify performance degradation across clinical segments, document types, clients, data sources, labels, and edge cases.
  • Work with Clinical AI Data Specialists to design gold sets, hidden test sets, adjudication workflows, and label quality checks.
  • Work with Research Engineers to understand model changes, expected behavior, and evaluation risks without compromising test-set independence.
  • Work with MLOps/Data Engineering to monitor production behavior, triage bugs, analyze incident impact, and prioritize fixes.
  • Create release-readiness reports before production deployment.
  • Build dashboards, scripts, and automated checks for evaluation, monitoring, regression testing, and model comparison.
  • Prioritize model bugs based on clinical severity, user impact, frequency, regression risk, and operational urgency.
WHAT WE EXPECT
  • 3-6+ years of experience in ML engineering, data science, model evaluation, ML QA, applied NLP evaluation, or data-heavy quality engineering.
  • Strong Python and data analysis skills.
  • Strong understanding of precision/recall/F1, calibration, confidence thresholds, dataset splits, leakage, overfitting, statistical testing, and error analysis.
  • Experience building evaluation pipelines, benchmark suites, test harnesses, dashboards, or regression frameworks.
  • Ability to work with imperfect labels, annotation disagreement, clinical ambiguity, and hidden evaluation sets.
  • Strong independence and judgment; ability to challenge releases when evidence is weak.
  • Clear written communication for release reports, incident analysis, and quality decisions.
NICE TO HAVE
  • Experience with LLM evaluation, RAG evaluation, extraction evaluation, clinical NLP, or healthcare ML.
  • Experience with model monitoring, production incident analysis, data drift, or observability.
  • Experience with MLflow, Weights & Biases, Evidently, Great Expectations, DeepEval, Ragas, pytest, Airflow, Prefect, or similar tools.
  • Clinical or biomedical NLP exposure.
SUCCESS IN 6 MONTHS
  • Establishes a repeatable release validation process.
  • Maintains hidden evaluation datasets and prevents overfitting to test data.
  • Produces release reports that leadership, ML, and engineering can trust.
  • Catches meaningful regressions before release.
  • Provides reliable impact analysis for production issues and helps prioritize fixes.
Why Join Triomics
  • Impact at scale. The systems your teams build directly power AI workflows that accelerate cancer research and improve patient outcomes.
  • Cutting-edge problems. Hard, data-intensive systems at the intersection of AI, healthcare, and scale - in a highly regulated industry where reliability is non-negotiable.
  • World-class team. Work alongside top talent across AI, engineering, and product, with best-in-industry compensation.
  • Culture that ships. Fast-paced, ownership-driven, with company-sponsored workations.
Perks & Benefits
  • Lunch provided at the office - one less daily decision.
  • Flexible working hours - we care about output, not clock-ins.
  • Comprehensive health insurance for you and your family.
  • Zomato meal benefits for early starts and late nights.
  • Remote Wednesdays
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Model Evaluation Engineer
ML Model Evaluation Engineer

Triomics • Varanasi

On-site
INR 1,200,000 - 2,400,000
Lunch provided at the office.
Flexible working hours.
Comprehensive health insurance for you
+1
Senior ML Research Engineer
Senior ML Research Engineer

Mployee.me • Bengaluru

Hybrid
INR 450,000 - 750,000
Lunch provided
Flexible working hours
Comprehensive health insurance for you
+2
Research Engineer, Applied ML
Research Engineer, Applied ML

Triomics • Varanasi

On-site
INR 900,000 - 1,500,000
Lunch provided at the office
Flexible working hours
Comprehensive health insurance for you
+1
Site Reliability Engineer
Site Reliability Engineer

Triomics • Varanasi

On-site
INR 350,000 - 600,000
Lunch Provided at Office
Flexible Working Hours
Health Insurance
+1
Implementation Engineer
Implementation Engineer

Triomics • Varanasi

On-site
INR 900,000 - 1,700,000
Lunch provided at the office
Flexible working hours
Comprehensive health insurance for you
+1
Technical Support Engineer
Technical Support Engineer

Mployee.me • Bengaluru

On-site
INR 900,000 - 1,300,000
Lunch provided at the office
Flexible working hours
Comprehensive health insurance
+1
Clinical Data Abstractor
Clinical Data Abstractor

Triomics • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Technical Support Engineer — US Shift
Technical Support Engineer — US Shift

Triomics • India

On-site
INR 900,000 - 1,300,000
Lunch provided
Flexible working hours
Comprehensive health insurance
+1
Implementation Engineer
Implementation Engineer

Triomics • Karnataka

On-site
INR 600,000 - 900,000
Lunch provided at office
Flexible working hours
Health insurance for you and family
+1
Senior Medical Data Abstractor (Immediate Joiners Only)
Senior Medical Data Abstractor (Immediate Joiners Only)

Triomics • Bengaluru

On-site
INR 900,000 - 1,300,000