Distinguished, Data Scientist - Quality & LLM Judging Systems in Conversational Commerce - Walmart

OpenTalent

Fremont (CA)

On-site

USD 230,000 - 300,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Walmart’s Next Gen Commerce team is shaping the future of conversational shopping by building intelligent agents. A Distinguished Data Scientist for Quality & LLM Judging Systems will lead model development for evaluation methodologies, ensuring high-quality judgments for AI-powered conversations and tools.

You will design prompts, validate agreement with human judgment, and collaborate with product and platform teams to drive trustworthy AI and reliable metrics across the system.

Qualifications

  • 7+ years of experience in data science or machine learning, preferably in evaluation, NLP, or conversational AI.
  • Hands-on experience with large language models, including prompt engineering, response grading, and structured generation tasks.
  • Familiarity with both human annotation workflows and automated evaluation strategies using LLMs.
  • Deep understanding of metric design, evaluation reliability, and statistical validity.
  • Strong software engineering fundamentals and ability to own end-to-end pipelines.

Responsibilities

  • Design evaluation pipelines for conversational agents and their tool outputs using LLM-as-a-judge, human annotation, and hybrid methods.
  • Develop high-quality prompts for structured evaluation tasks and iterate based on inter-rater reliability with human judges.
  • Develop novel techniques to assess non-textual or subjective outputs—such as recommendations, summaries, and agent-driven actions—where standard metrics fall short.
  • Guide the modeling team to distill or fine-tune smaller LLMs to act as scalable evaluation proxies.
  • Work with engineering partners to integrate evaluation hooks into model training, validation, and production workflows.
  • Conduct in-depth failure mode analysis and define actionable quality signals that inform model and production iteration.
  • Uphold statistical rigor in metric design, validation, and experimental analysis to ensure reliable and interpretable results.
  • Foster a culture of principled measurement and trustworthy AI throughout the organization.
  • 7+ years of experience in data science or machine learning, preferably in evaluation, NLP, or conversational AI.
  • Hands-on experience with large language models, including prompt engineering, response grading, and structured generation tasks.
  • Familiarity with both human annotation workflows and automated evaluation strategies using LLMs.
  • Deep understanding of metric design, evaluation reliability, and statistical validity.
  • Strong software engineering fundamentals and ability to own end-to-end pipelines.

Job description

Position Summary... What you'll do... About the Role Walmart's Next Gen Commerce team is shaping the future of conversational shopping by building intelligent agents that not only respond, but reason, recommend, and proactively assist customers. As a Distinguished Data Scientist for Quality & LLM Judging Systems in Conversational Commerce , you will serve as the key IC partner to the Director of Data Science for this space. You will lead the technical vision and model development for cutting-edge evaluation methodologies to measure and improve the quality of AI-powered conversations and tool outputs. You'll help define how we evaluate our agents and their dependent tools using a combination of human-labeled benchmarks, LLM-as-a-judge systems, and scalable automated pipelines. You'll design prompts, validate agreement with human judgment, and develop LLM distillation strategies to replicate high-quality judgment cost-effectively. This is a high-impact, hands-on technical role requiring deep expertise in LLM prompting, evaluation frameworks, and structured experimentation. You will work closely with modeling, product, and platform teams to ensure that measurement drives improvement, and that the agent's behaviors align with quality, safety, and relevance at every step.

Responsibilities
  • Design evaluation pipelines for conversational agents and their tool outputs using LLM-as-a-judge, human annotation, and hybrid methods
  • Develop high-quality prompts for structured evaluation tasks and iterate based on inter-rater reliability with human judges
  • Develop novel techniques to assess non-textual or subjective outputs-such as recommendations, summaries, and agent-driven actions-where standard metrics fall short
  • Guide the modeling team to distill or fine-tune smaller LLMs to act as scalable evaluation proxies
  • Work with engineering partners to integrate evaluation hooks into model training, validation, and production workflows
  • Conduct in-depth failure mode analysis and define actionable quality signals that inform model and production iteration.
  • Uphold statistical rigor in metric design, validation, and experimental analysis to ensure reliable and interpretable results
  • Foster a culture of principled measurement and trustworthy AI throughout the organization
  • 7+ years of experience in data science or machine learning, preferably in evaluation, NLP, or conversational AI
  • Hands-on experience with large language models, including prompt engineering, response grading, and structured generation tasks
  • Familiarity with both human annotation workflows and automated evaluation strategies using LLMs
  • Deep understanding of metric design, evaluation reliability, and statistical validity
  • Strong software engineering fundamentals and ability to own end-to-end pipelines
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Scientist: AI Evaluation & LLM Judging
Senior Data Scientist: AI Evaluation & LLM Judging

OpenTalent • Fremont (CA)

On-site
USD 230,000 - 300,000
Distinguished, Data Scientist - Agent-Led Engagement in Conversational Commerce - Walmart
Distinguished, Data Scientist - Agent-Led Engagement in Conversational Commerce - Walmart

OpenTalent • Fremont (CA)

On-site
USD 180,000 - 240,000
Data Scientist
Data Scientist

Programmers.io • Louisville (KY)

On-site
USD 120,000 - 160,000
Machine Learning Engineer - Agentic AI Evaluation Frameworks
Machine Learning Engineer - Agentic AI Evaluation Frameworks

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
AI Quality Scientist: LLM Evaluation and Governance
AI Quality Scientist: LLM Evaluation and Governance

Programmers.io • Louisville (KY)

On-site
USD 120,000 - 160,000
Applied Scientist/Research Engineer, LLM Training Data
Applied Scientist/Research Engineer, LLM Training Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 180,000
Sr. Machine Learning Engineer, Speech LLM Evaluation
Sr. Machine Learning Engineer, Speech LLM Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Applied Scientist, AGI Quality & LLM Evaluation
Applied Scientist, AGI Quality & LLM Evaluation

Amazon • Factoria (WA)

On-site
USD 136,000 - 184,000
Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E
Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E

Apple • San Diego (CA)

On-site
USD 140,000 - 190,000
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

Remote
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits