ML Engineer - LLM Evaluation & Automation

Grid Dynamics

United States

On-site

USD 140,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible schedule
Medical insurance
Vision and dental
Corporate social events
Professional development opportunities
Well-equipped office

Job summary

Grid Dynamics in the United States is seeking a Machine Learning Engineer specializing in LLM‑based evaluation to design and build automated systems that measure and improve the quality of model outputs.

You will lead evaluation pipelines, develop metrics, and deliver actionable insights for continuous product improvements while collaborating with cross‑functional teams across data science, engineering, and product.

Qualifications

  • 5+ years of experience in ML engineering, NLP, or AI/ML automation.
  • Hands‑on experience in prompt engineering and designing LLM‑based evaluation systems is preferred.
  • Strong understanding of machine learning principles with a focus on NLP and advanced LLM capabilities (e.g., Chain‑of‑Thought, agentic workflows).

Responsibilities

  • Design and implement automated systems and pipelines for evaluating LLM outputs.
  • Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM‑based evaluations.
  • Collaborate with Engineering teams to create automated logic checks and validation tools.
  • Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures.
  • Provide feedback loops to ensure evaluation guidelines align with LLM‑based assessments.
  • Investigate how LLM‑derived evaluations can enhance product reliability and user experience.
  • Recommend refinements to prompt engineering, evaluation strategies, and automation tools.
  • Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains.
  • Continuously improve and expand automated evaluation processes based on industry best practices.

Skills

Python
SQL
Prompt engineering
LLM evaluation
NLP

Education

Bachelor’s/Master’s in Computer Science/Engineering

Tools

PySpark

Job description

We are seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will design and build systems that automatically measure and improve the accuracy, relevance, and consistency of model outputs. You will lead initiatives to create evaluation pipelines, develop metrics, and deliver actionable insights for continuous improvements. This position requires strong technical expertise, analytical problem‑solving abilities, and the capacity to manage projects across multiple cross‑functional teams.

Responsibilities
  • Design and implement automated systems and pipelines for evaluating LLM outputs.
  • Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM‑based evaluations.
  • Collaborate with Engineering teams to create automated logic checks and validation tools.
  • Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures.
  • Provide feedback loops to ensure evaluation guidelines align with LLM‑based assessments.
  • Investigate how LLM‑derived evaluations can enhance product reliability and user experience.
  • Recommend refinements to prompt engineering, evaluation strategies, and automation tools.
  • Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains.
  • Continuously improve and expand automated evaluation processes based on industry best practices.
Requirements
  • 5+ years of experience in ML engineering, NLP, or AI/ML automation.
  • Hands‑on experience in prompt engineering and designing LLM‑based evaluation systems is preferred.
  • Strong understanding of machine learning principles with a focus on NLP and advanced LLM capabilities (e.g., Chain‑of‑Thought, agentic workflows).
  • Expertise in building automated evaluation or QA pipelines.
  • Excellent analytical and problem‑solving skills with experience in root cause and error pattern analysis.
  • Proven project management and cross‑functional collaboration experience.
  • Excellent communication skills to convey complex insights to technical and non‑technical audiences.
  • Detail‑oriented mindset with a focus on evaluation metrics, prompt design, and automation.
  • Ability to quickly adapt to new business rules and evaluation guidelines across diverse product domains.
  • Strong programming skills in Python and SQL.
  • Experience with big data technologies like PySpark for data aggregation and sampling is a strong plus.
  • Bachelor’s/Master’s degree in Computer Science/Engineering or a related field.
We offer
  • Opportunity to work on cutting‑edge projects.
  • Work with a highly motivated and dedicated team.
  • Competitive salary.
  • Flexible schedule.
  • Benefits package – medical insurance, vision, dental, etc.
  • Corporate social events.
  • Professional development opportunities.
  • Well‑equipped office.
About Us

Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI, supported by profound expertise and ongoing investment in data, analytics, cloud & DevOps, application modernization and customer experience. Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer
Senior ML Engineer

Grid Dynamics • United States

On-site
USD 140,000 - 190,000
Flexible schedule
Medical insurance
Vision and dental
+2
LLM Evaluation Engineer — Quality & Automation
LLM Evaluation Engineer — Quality & Automation

Grid Dynamics • United States

On-site
USD 140,000 - 170,000
Flexible schedule
Medical insurance
Vision and dental
+3
ML Engineer — LLM Evaluation
ML Engineer — LLM Evaluation

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Principal AI Engineer
Principal AI Engineer

Stellantis NV • Auburn (AL)

On-site
USD 150,000 - 230,000