ML Engineer: LLM Evaluation & Observability

Gleanwork

Mountain View (CA)

Hybrid

USD 200,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) plan
Home office improvement stipend
Annual education stipend
Wellness stipend
Daily healthy lunches

Job summary

Gleanwork is looking for a software engineer focused on AI assistant evaluation in Mountain View, CA. You will design evaluation datasets and build pipelines to measure assistant quality, collaborating closely with cross-functional teams.

The ideal candidate has strong coding skills in Go and Python, and experience with LLM evaluation. The position offers a competitive salary between $200,000 and $300,000, with hybrid work options and comprehensive benefits including health coverage and stipends for home office improvements and education.

Qualifications

  • 2+ years of software engineering experience with strong coding skills.
  • Strong backend fundamentals in Go and Python.
  • Comfortable with distributed data pipelines.
  • Experience with LLM evaluation or reinforcement learning.
  • Analytically rigorous in predicting user experience.
  • Strong team player in a cross-functional environment.

Responsibilities

  • Design and curate evaluation datasets for assistant behavior.
  • Build large-scale evaluation pipelines for quality measurement.
  • Develop LLM-powered judges for scoring assistant metrics.
  • Evaluate new models and product changes before launches.
  • Build observability infrastructure for AI agents.
  • Collaborate with engineers to enhance evaluation processes.

Skills

Software engineering experience
Backend fundamentals in Go and Python
Experience with LLM evaluation
Analytical rigor
Customer-focused mindset

Job description

Gleanwork is looking for a software engineer focused on AI assistant evaluation in Mountain View, CA. You will design evaluation datasets and build pipelines to measure assistant quality, collaborating closely with cross-functional teams.

The ideal candidate has strong coding skills in Go and Python, and experience with LLM evaluation. The position offers a competitive salary between $200,000 and $300,000, with hybrid work options and comprehensive benefits including health coverage and stipends for home office improvements and education.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer, LLM Eval & Observability (Hybrid)
ML Engineer, LLM Eval & Observability (Hybrid)

Gleanwork • San Francisco (CA)

Hybrid
USD 200,000 - 300,000
Home office improvement stipend
Annual education and wellness stipends
Healthy daily lunches
LLM Evaluation & Verification Lead
LLM Evaluation & Verification Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
ML Engineer: LLMs, VLMs & Reasoning AI | Equity
ML Engineer: LLMs, VLMs & Reasoning AI | Equity

Tensor • San Jose (CA)

On-site
USD 75,000 - 300,000
Competitive compensation package
Participation in discretionary equity incentive plan
Access to comprehensive benefits program
ML Engineer: AI Evaluation & LLM Systems
ML Engineer: AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
ML Engineer — LLM Evaluation
ML Engineer — LLM Evaluation

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Engineer: Build Production‑Ready LLM Features
AI Engineer: Build Production‑Ready LLM Features

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
US$1,000 per month food and commuting allowance
Laptop of choice
ESOP available
Applied AI Software Engineer: LLM Agent Evaluation Lead
Applied AI Software Engineer: LLM Agent Evaluation Lead

Canvas Construction • San Francisco (CA)

Hybrid
USD 300,000 - 400,000
Competitive Salary & Equity Package
Health Insurance
Home Office Stipend
+3
Machine Learning Engineer, LLM Evals & Observability
Machine Learning Engineer, LLM Evals & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Senior LLM Evaluation Infrastructure Engineer
Senior LLM Evaluation Infrastructure Engineer

Inception • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health insurance
Equity
Flexible vacation
+2