Research Engineer, Knowledge & Evaluation Systems

Neura Market

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 850,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a Research Engineer on the Knowledge team to design and run experiments that improve Claude's search, retrieval, and reasoning over information at scale.

You will build training environments, curate data, and contribute to evaluation suites while collaborating with researchers across RL Data, post-training, and product teams. This hybrid role emphasizes reliable, observable systems and scalable tooling.

Qualifications

  • Experienced Python engineer with a track record of shipping reliable, well-instrumented code.
  • Experience designing, running, and analyzing ML experiments.
  • Ability to work across the stack—from data pipelines to model training to evaluation.
  • 5+ years of experience operating ML or distributed systems at scale.
  • Clear written and verbal communication across time zones.
  • Finds genuine satisfaction and impact in making critical systems dependable.

Responsibilities

  • Design, build, and iterate on training environments and data pipelines that improve Claude's ability to reason over knowledge-intensive tasks
  • Run experiments end-to-end: form a hypothesis, build the infrastructure, train models, analyze results, and decide what to try next
  • Develop evaluations that meaningfully capture progress on search, retrieval, and reasoning quality
  • Identify failure modes in current model behavior and translate them into concrete training signals
  • Collaborate closely with researchers across RL Data, post-training, and product teams to align on priorities and ship improvements
  • Contribute to shared infrastructure and tooling that compounds the team's velocity over time
  • Own a clean, canonical set of evaluation tools and processes for Knowledge Work capabilities, including the process used for model releases
  • Build and automate observability, dashboards, and operational tooling for our training environments and evaluation systems, with an emphasis on high signal-to-noise

Skills

Python
ML experiments
Distributed systems
Data pipelines
Model training
Cross-time-zone communication

Job description

Anthropic is seeking a Research Engineer on the Knowledge team to design and run experiments that improve Claude's search, retrieval, and reasoning over information at scale.

You will build training environments, curate data, and contribute to evaluation suites while collaborating with researchers across RL Data, post-training, and product teams. This hybrid role emphasizes reliable, observable systems and scalable tooling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - AI Knowledge Architect
Research Engineer - AI Knowledge Architect

Anthropic • United States

Hybrid
USD 350,000 - 850,000
Research Engineer - AI Knowledge Architect
Research Engineer - AI Knowledge Architect

SignalAI • New York (NY)

Hybrid
USD 350,000 - 850,000
Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Eval Engineer for AI Model Evaluations
Eval Engineer for AI Model Evaluations

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Research Engineer, Knowledge Team
Research Engineer, Knowledge Team

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Anthropic • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space for collaboration
Research Engineer, Knowledge Team
Research Engineer, Knowledge Team

Anthropic • United States

Hybrid
USD 350,000 - 850,000
Research Engineer: AI Computer-Use & Vision RL
Research Engineer: AI Computer-Use & Vision RL

Menlo Ventures • New York (NY)

Hybrid
USD 500,000 - 850,000
Competitive compensation and benefits
Equity donation matching (optional)
Generous vacation and parental leave
+2
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Research Engineer - Knowledge Architect for LLM Data
Research Engineer - Knowledge Architect for LLM Data

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
+2