Research Engineer, Model Evaluations - Remote-Friendly Impact
Menlo Ventures
San Francisco (CA)
On-site
USD 320,000 - 485,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Generous vacation and parental leave
Flexible working hours
Lovely office space for collaboration
Job summary
Anthropic in New York City is seeking a Research Engineer to develop evaluations for Claude’s capabilities. The ideal candidate should have strong Python programming skills, experience with distributed systems, and the ability to communicate technical results effectively. Responsibilities include designing evaluations, building infrastructure for running evaluations, and debugging results during training runs. The role offers a hybrid work model and competitive compensation ranging from $320,000 to $485,000 USD.
Qualifications
Strong Python programming skills, including production or research infrastructure.
Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale.
Clear written and verbal communication, especially when explaining technical results to non-specialists.
Responsibilities
Design and run new evaluations of Claude’s capabilities.
Build the distributed eval execution platform.
Debug anomalous eval results mid-training-run.
Communicate evaluations and their results to stakeholders.
Skills
Python programming skills
Experience building distributed systems
Clear communication
Experience with large language models
Experience with data visualization
Education
Bachelor’s degree or equivalent
Tools
Data pipelines
Job description
Anthropic in New York City is seeking a Research Engineer to develop evaluations for Claude’s capabilities. The ideal candidate should have strong Python programming skills, experience with distributed systems, and the ability to communicate technical results effectively. Responsibilities include designing evaluations, building infrastructure for running evaluations, and debugging results during training runs. The role offers a hybrid work model and competitive compensation ranging from $320,000 to $485,000 USD.