Engineering Manager, AI Evaluation Systems

Cursor

San Francisco (CA)

On-site

USD 130,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cursor in San Francisco is seeking an Engineering Manager to lead the Evals team responsible for building high-quality evaluation datasets for coding agents. You will set the roadmap for evaluations and guide a high-impact team of engineers.

The ideal candidate has experience leading engineering teams, strong leadership skills, and the ability to align research and product metrics effectively. If you are passionate about coding and enjoy building assessment tools, this role is for you.

Qualifications

  • Experience in leading engineering teams shipping production systems.
  • Strong people leadership and coaching skills.
  • Good taste and strong opinions on model and agent behaviors.
  • Strong data acumen and ability to collaborate effectively with scientists.

Responsibilities

  • Set and lead the eval roadmap from end to end.
  • Grow a high-impact team building eval datasets and tools.
  • Define online quality signals and turn regressions into guardrails.
  • Integrate evals into decision-making for launches and training.

Skills

Leadership
Engineering
Data collaboration
AI evaluation systems

Job description

Cursor in San Francisco is seeking an Engineering Manager to lead the Evals team responsible for building high-quality evaluation datasets for coding agents. You will set the roadmap for evaluations and guide a high-impact team of engineers.

The ideal candidate has experience leading engineering teams, strong leadership skills, and the ability to align research and product metrics effectively. If you are passionate about coding and enjoy building assessment tools, this role is for you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Evals
Engineering Manager, Evals

Socket.dev • San Francisco (CA)

On-site
USD 130,000 - 160,000
Engineering Manager, AI Prompts & Eval Platform
Engineering Manager, AI Prompts & Eval Platform

Anthropic • San Francisco (CA)

Hybrid
USD <1,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
AI Evaluation Engineer: Coding Task Architect
AI Evaluation Engineer: Coding Task Architect

United States Digital Space LLC • United States

Remote
USD 55,000 - 69,000
Engineering Manager - AI Evaluation & Test Tooling
Engineering Manager - AI Evaluation & Test Tooling

Apple Inc. • Cupertino (CA)

On-site
USD 237,000 - 357,000
Stock programs
Discretionary bonuses
Relocation
+2
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior AI Evaluation Engineer: Pipelines & Metrics
Senior AI Evaluation Engineer: Pipelines & Metrics

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Benchmarking & Evaluation Engineer
AI Benchmarking & Evaluation Engineer

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
AI Coding Evaluator & Testing Architect (Python)
AI Coding Evaluator & Testing Architect (Python)

Mindrift • United States

On-site
USD 55,000 - 69,000
AI Engineer – Algorithm Evaluation & Agentic Systems
AI Engineer – Algorithm Evaluation & Agentic Systems

Socket.dev • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Staff Engineer — Evals & Post-Training Product
Staff Engineer — Evals & Post-Training Product

Fireworks AI • San Mateo (CA)

On-site
USD 120,000 - 160,000
Work with cutting-edge technology
Collaborative and inclusive environment
Opportunities for growth and learning