Product Manager, Claude Code Model Performance

United States Digital Space LLC

Washington

On-site

USD 305,000 - 460,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity donation matching
Generous vacation and parental leave
Flexible working hours
SF office

Job summary

Claude Code is seeking a Product Manager to drive end-to-end model launches, build evals that measure meaningful coding performance, and collaborate with researchers and engineers to translate model improvements into developer-facing outcomes.

You will be the bridge between frontier research and the millions of developers who rely on Claude Code to do their best work. Expect a mission-driven, highly collaborative environment in one of our US offices.

Qualifications

  • Have personally built agentic evals (e.g. SWE-bench-style task suites).
  • Are a daily Claude Code user and can articulate what behaviors you’d want to change or add to the model.
  • Have an engineering background and 2+ years in product management, or equivalent experience driving product direction as an engineer.
  • Have a deep grasp of AI concepts and are comfortable going deep on model behavior, prompt engineering, and evaluation methodology.
  • Are a systems thinker: when you find a problem, you build the infrastructure that prevents its whole class.
  • Have launched products or capabilities in ambiguous, research-adjacent environments.
  • Have a creative, hacker spirit and love solving puzzles.

Responsibilities

  • Own model launch planning and execution for Claude Code: define readiness criteria, coordinate across research and product engineering, and ensure launches land cleanly with developers.
  • Design and implement agentic evals that measure real-world coding performance.
  • Drive the engineering team's eval roadmap.
  • Partner with researchers working on coding capabilities to define target behaviors and influence model development with evidence from real usage.
  • Talk with users and analyze transcripts to understand capability gaps and turn research progress into shipped improvements.
  • Synthesize signal from internal users, external developers, and competitive benchmarks into clear priorities.

Skills

Product management
Agentic evals
Systems thinking
Prompt engineering
Communication

Education

Bachelor’s degree or equivalent

Job description

About the company

the company’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

As a Product Manager on Claude Code's model performance team, you will drive model launches end-to-end, build evals that measure what matters, and partner directly with researchers and product engineers to translate model improvements into developer-facing outcomes.

Claude Code is the most capable coding agent in the world but there’s much more we can do to extract the maximum performance from our models. We're looking for a PM who has personally built agentic evals, thinks in systems, uses Claude Code every day, and has refined model taste. You should be as comfortable influencing our research team as you are getting in the weeds of transcripts. You will be the connective tissue between frontier research and the millions of developers who depend on Claude Code to do their best work.

Responsibilities
  • Own model launch planning and execution for Claude Code: define readiness criteria, coordinate across research and product engineering, and ensure launches land cleanly with developers
  • Design and implement agentic evals that measure real-world coding performance
  • Drive the engineering team's eval roadmap
  • Partner with researchers working on coding capabilities to define target behaviors and influence model development with evidence from real usage
  • Talk with users and analyze transcripts to understand capability gaps and turn research progress into shipped improvements
  • Synthesize signal from internal users, external developers, and competitive benchmarks into clear priorities
You might be a good fit if you
  • Have personally built agentic evals (e.g. SWE-bench-style task suites)
  • Are a daily Claude Code user and can articulate what behaviors you’d want to change or add to the model
  • Have an engineering background and 2+ years in product management, or equivalent experience driving product direction as an engineer
  • Have a deep grasp of AI concepts and are comfortable going deep on model behavior, prompt engineering, and evaluation methodology
  • Are a systems thinker: when you find a problem, you build the infrastructure that prevents its whole class
  • Have launched products or capabilities in ambiguous, research-adjacent environments
  • Have a creative, hacker spirit and love solving puzzles

San Francisco and Seattle only

The annual compensation range for this role is listed below.

For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$305,000—$460,000 USD

LogisticsMinimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experienceRequired field of study:A field relevant to the role as demonstrated through coursework, training, or professional experienceMinimum years of experience: Years of experience required will correlate with the internal job level requirements for the positionLocation-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.Visa sponsorship:We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

How we're different

We believe that the highest-impact AI research will be big science. At the company we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.

The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to the company, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.

Come work with us!

the company is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage:Learn aboutour policy for using AI in our application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Manager, Claude Code Model Performance
Product Manager, Claude Code Model Performance

Visa Hunt • San Francisco (CA)

Hybrid
USD 305,000 - 460,000
Product Manager, Claude Code Model Performance
Product Manager, Claude Code Model Performance

Anthropic • Seattle (WA)

Hybrid
USD 305,000 - 460,000
Product Manager, Claude Code Model Performance
Product Manager, Claude Code Model Performance

Anthropic • San Francisco (CA)

Hybrid
USD 305,000 - 460,000
Product Manager, Claude Code Model Performance San Francisco, CA | New York City, NY
Product Manager, Claude Code Model Performance San Francisco, CA | New York City, NY

Anthropic • San Francisco (CA)

Hybrid
USD 305,000 - 460,000
Staff+ Software Engineer, Enterprise AI Products
Staff+ Software Engineer, Enterprise AI Products

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Research Product Manager, Model Behaviors
Research Product Manager, Model Behaviors

United States Digital Space LLC • New York (NY), San Francisco (CA)

Hybrid
USD 385,000 - 460,000
Product Manager, Claude Tag
Product Manager, Claude Tag

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 385,000 - 460,000
Staff Software Engineer, Claude Code
Staff Software Engineer, Claude Code

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 320,000 - 625,000
Staff+ Software Engineer, Claude Science
Staff+ Software Engineer, Claude Science

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Engineering Manager, Agent Prompts & Evals
Engineering Manager, Agent Prompts & Evals

Anthropic • San Francisco (CA)

On-site
USD <1,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave