ML Engineer, LLM Eval & Observability (Hybrid)

Gleanwork

San Francisco (CA)

Hybrid

USD 200,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Home office improvement stipend
Annual education and wellness stipends
Healthy daily lunches

Job summary

Gleanwork is seeking a Software Engineer to enhance their AI assistant's quality through rigorous evaluation and measurement systems. You'll design datasets and build pipelines while collaborating with engineers across the company.

The ideal candidate has expertise in Go and Python, and at least 2 years of experience. This hybrid role is based in the San Francisco Bay Area, offering competitive compensation and a comprehensive benefits package, including stipends for home office improvements.

Qualifications

  • 2+ years of software engineering experience with strong coding skills.
  • Strong backend fundamentals in Go and Python; comfortable with distributed data pipelines.
  • Experience working with LLM evaluation or reinforcement learning from human feedback.

Responsibilities

  • Design and curate evaluation datasets for assistant quality measurement.
  • Build and maintain large-scale evaluation pipelines.
  • Build LLM-powered judges that score metrics like correctness and completeness.
  • Evaluate new models and product changes before they ship.
  • Build observability infrastructure for AI agents.

Skills

Software engineering
Go
Python
Machine learning
Natural language processing

Job description

Gleanwork is seeking a Software Engineer to enhance their AI assistant's quality through rigorous evaluation and measurement systems. You'll design datasets and build pipelines while collaborating with engineers across the company.

The ideal candidate has expertise in Go and Python, and at least 2 years of experience. This hybrid role is based in the San Francisco Bay Area, offering competitive compensation and a comprehensive benefits package, including stipends for home office improvements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer: LLM Evaluation & Observability
ML Engineer: LLM Evaluation & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
Machine Learning Engineer, LLM Evals & Observability
Machine Learning Engineer, LLM Evals & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
Machine Learning Engineer, LLM Evals & Observability
Machine Learning Engineer, LLM Evals & Observability

Gleanwork • San Francisco (CA)

Hybrid
USD 200,000 - 300,000
Home office improvement stipend
Annual education and wellness stipends
Healthy daily lunches
Production ML Engineer — AI Assistants (Hybrid SF)
Production ML Engineer — AI Assistants (Hybrid SF)

Glean • San Francisco (CA)

Hybrid
USD 180,000 - 205,000
Medical coverage
Vision coverage
Dental coverage
+6
ML Engineer - LLM Evaluation & Automation
ML Engineer - LLM Evaluation & Automation

Grid Dynamics • United States

On-site
USD 140,000 - 170,000
Flexible schedule
Medical insurance
Vision and dental
+3
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
Staff Software Engineer, AI & LLM Features
Staff Software Engineer, AI & LLM Features

Monograph • San Francisco (CA)

Hybrid
USD 227,000 - 284,000
Equity (RSUs)
Competitive base pay
Comprehensive benefits
Principal AI Engineer
Principal AI Engineer

People In AI • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior AI Engineer - LLM Platform Lead
Senior AI Engineer - LLM Platform Lead

ZS • South San Francisco (CA)

Hybrid
USD 120,000 - 150,000
ML Engineer — LLMs, KYC/KYB, Hybrid SF
ML Engineer — LLMs, KYC/KYB, Hybrid SF

Baselayer • San Francisco (CA)

Hybrid
USD 150,000 - 225,000