LLM Evaluation Engineer — Quality & Automation

Grid Dynamics

United States

On-site

USD 140,000 - 170,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flexible schedule
Medical insurance
Vision and dental
Corporate social events
Professional development opportunities
Well-equipped office

Job summary

Grid Dynamics in the United States is seeking a Machine Learning Engineer specializing in LLM‑based evaluation to design and build automated systems that measure and improve the quality of model outputs.

You will lead evaluation pipelines, develop metrics, and deliver actionable insights for continuous product improvements while collaborating with cross‑functional teams across data science, engineering, and product.

Qualifications

  • 5+ years of experience in ML engineering, NLP, or AI/ML automation.
  • Hands‑on experience in prompt engineering and designing LLM‑based evaluation systems is preferred.
  • Strong understanding of machine learning principles with a focus on NLP and advanced LLM capabilities (e.g., Chain‑of‑Thought, agentic workflows).

Responsibilities

  • Design and implement automated systems and pipelines for evaluating LLM outputs.
  • Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM‑based evaluations.
  • Collaborate with Engineering teams to create automated logic checks and validation tools.
  • Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures.
  • Provide feedback loops to ensure evaluation guidelines align with LLM‑based assessments.
  • Investigate how LLM‑derived evaluations can enhance product reliability and user experience.
  • Recommend refinements to prompt engineering, evaluation strategies, and automation tools.
  • Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains.
  • Continuously improve and expand automated evaluation processes based on industry best practices.

Skills

Python
SQL
Prompt engineering
LLM evaluation
NLP

Education

Bachelor’s/Master’s in Computer Science/Engineering

Tools

PySpark

Job description

Grid Dynamics in the United States is seeking a Machine Learning Engineer specializing in LLM‑based evaluation to design and build automated systems that measure and improve the quality of model outputs.

You will lead evaluation pipelines, develop metrics, and deliver actionable insights for continuous product improvements while collaborating with cross‑functional teams across data science, engineering, and product.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer - LLM Evaluation & Automation
ML Engineer - LLM Evaluation & Automation

Grid Dynamics • United States

On-site
USD 140,000 - 170,000
Flexible schedule
Medical insurance
Vision and dental
+3
Senior ML Engineer: RAG/LLM Systems & Rapid Prototyping
Senior ML Engineer: RAG/LLM Systems & Rapid Prototyping

Grid Dynamics • United States

On-site
USD 140,000 - 190,000
Flexible schedule
Medical insurance
Vision and dental
+2
Senior AI Engineer - Production LLM EvalOps
Senior AI Engineer - Production LLM EvalOps

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
LLM Evaluation & Verification Lead
LLM Evaluation & Verification Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
Data Quality Engineer — LLM Post-Training & QA Pipelines
Data Quality Engineer — LLM Post-Training & QA Pipelines

Reflection AI Ltd • New York (NY)

On-site
USD 140,000 - 190,000
Top-tier compensation
Stock options
Health & wellness
+5
Senior AI Engineer: LLM Evaluation, Production & Optimization
Senior AI Engineer: LLM Evaluation, Production & Optimization

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
LLM Evaluation & Engineering Simulation Lead (Freelance)
LLM Evaluation & Engineering Simulation Lead (Freelance)

Zettamine Labs • United States

Remote
USD 165,000 - 276,000
LLM Evaluation & Benchmarking Engineer
LLM Evaluation & Benchmarking Engineer

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000