AI Benchmarking Engineer — Evaluations & Failure Analysis
Mercor
San Francisco (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
Free Equinox membership
Health insurance
Job summary
A cutting-edge AI firm in San Francisco is seeking a Research Engineer to develop evaluation systems and benchmarking pipelines for language models. Candidates should have a strong background in applied research, coding skills, and familiarity with ML models. You will work collaboratively with cross-functional teams and operate in a fast-paced environment, conducting failure analyses and contributing to the improvement of AI tools. The role requires working onsite five days a week, emphasizing strong ownership and high-intensity collaboration.
Qualifications
Experience in model evaluation, benchmarking, or failure analysis is expected.
Hands-on coding experience with machine learning models and evaluation code is crucial.
Comfortable with cloud services for deploying and storing evaluation results.
Responsibilities
Design, implement, and maintain benchmarking systems.
Build LLM evaluation systems with scoring and tracking capabilities.
Conduct systematic failure analyses to refine training processes.
Collaborate across teams to align benchmarks with training objectives.
Skills
Applied research background
Strong coding skills
Data structures and algorithms
Comfort with APIs
Tools
ML models
SQL/NoSQL
Cloud platforms
Job description
A cutting-edge AI firm in San Francisco is seeking a Research Engineer to develop evaluation systems and benchmarking pipelines for language models. Candidates should have a strong background in applied research, coding skills, and familiarity with ML models. You will work collaboratively with cross-functional teams and operate in a fast-paced environment, conducting failure analyses and contributing to the improvement of AI tools. The role requires working onsite five days a week, emphasizing strong ownership and high-intensity collaboration.