Researcher, Evaluations

Epoch Ai

Deutschland

Vor Ort

EUR 100.542 - 174.856

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Comprehensive health insurance
Generous paid time off
Flexible expense policy for productivity tools
Paid work trips and conferences

Zusammenfassung

Inclusively Remote is currently seeking a Researcher to evaluate frontier AI models on real-world tasks. This role involves curating evaluation suites, assessing AI performance quantitatively and qualitatively, and creating public reports.

Ideal candidates will have analytical thinking skills, experience with AI tools, and strong written communication abilities. The position offers a fully remote environment with competitive benefits, including unlimited paid time off.

Qualifikationen

  • Experience testing frontier models and writing assessments.
  • Ability to conduct rigorous experiments with evidence-supported findings.
  • Comfort analyzing results with some coding skills.

Aufgaben

  • Create and curate an evaluation suite with real-world AI challenges.
  • Regularly evaluate AI systems and update tasks and rubrics.
  • Communicate research findings in public reports and visualizations.
  • Conduct data analysis on evaluation results.
  • Automate workflow parts and build standalone benchmarks.

Kenntnisse

Analytical thinking
Grounded, skeptical mentality
Comfort with AI agents and tools
Familiarity with AI benchmarks
Research and data-analysis experience
Strong written communication skills

Tools

Python

Jobbeschreibung

About the role

Epoch AI is looking for a researcher to evaluate frontier AI models on hard-to-grade tasks drawn from real‑world scenarios. We’re seeking a Researcher to lead a new effort evaluating how well frontier models perform on the kinds of open‑ended tasks that make up real office work. You will curate a suite of realistic tasks to serve as a benchmark, design the grading rubrics for AI performance, and run newly‑released models through the suite, assessing their performance both quantitatively and qualitatively. The focus is on how models handle messy, real‑world work rather than on scientific knowledge or programming ability. The role makes heavy use of AI tools, but strong software engineering experience is not required. Comfort setting up AI‑assisted automated workflows is a plus. If this role sounds interesting, we are also looking for researchers on multiple other teams.

Key Responsibilities
  • Create and curate an evaluation suite. Find real‑world tasks that serve as challenging tests for practical AI capabilities, and update the tasks over time as AI capabilities evolve. Devise rubrics for evaluating AI performance.
  • Evaluate AI systems. Regularly evaluate new, notable AI models and products on the task suite. Update tasks and rubrics to reflect the changing landscape of AI capabilities.
  • Communicate your research. Create public‑facing reports, blog posts, and data visualizations with your observations. Ensure the evaluations feed into our other research topics and help keep our team informed.
  • Conduct data analysis. Analyze evaluation results and compare models across tasks.
  • Improve the process. You might automate parts of the workflow, and build out parts of the evaluation into standalone benchmarks.
What we are looking for
  • Analytical thinking. You conduct experiments with rigor and care, making sure that findings are well‑supported by evidence.
  • Grounded, skeptical mentality. You form your own well‑reasoned view of what an AI system can do, distinguishing practical capabilities from hype.
  • Comfort with AI agents and tools. You have experience working with AI agents in the course of your own work, and are comfortable delegating tasks.
  • Familiarity with AI benchmarks and evaluations. You follow AI capabilities at least casually and have opinions on what benchmarks do and don’t tell us.
  • Research and data‑analysis experience, including enough comfort with light coding to analyze your own results.
  • Strong written communication skills: you can convey nuanced observations clearly and precisely.
Nice to have
  • Experience testing frontier models and writing assessments of their capabilities.
  • Coding skills, including python proficiency.
Compensation & Benefits

Annual salary between $115,000 – $200,000 USD, depending on location and experience.

Salaries are not restricted to USD, and contracts and payments are usually in local currencies. Conversions are based on one‑year average exchange rates.

Fully remote environment, including flexible work hours.

Competitive global benefits program, including a comprehensive health insurance program—supplemental benefits specific to a local country, as available and mandated by local law—and life insurance and a pension plan, if applicable in your country.

Generous paid time off (PTO), including no specific annual limit, with 30 days PTO per year protected, unlimited personal and sick leave, and 4 months paid parental leave for permanent staff with at least 12 months of tenure (prorated parental leave if less than 12 months).

A flexible and generous expense policy for you to spend on equipment and a large range of productivity tools or learning/development opportunities, including unlimited spending on AI tools, subject to regulations and manager approval.

Paid work trips, including 3 staff retreats per year and relevant conferences.

Access to our very well‑equipped offices in Berkeley, California, including paid meals, snacks, gym, and more. All staff, independently of where they are based, have access to the office for at least 20 days each year.

Inclusion Statement

Epoch is committed to building an inclusive, equitable, and supportive community for you to thrive and do your best work. We’re committed to finding the best people for our team, so please don’t hesitate to apply for a role regardless of your age, gender identity/expression, political identity, personal preferences, physical abilities, veteran status, neurodiversity or any other background.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior+ Software Engineer - Research Platform, Consumer Devices
Senior+ Software Engineer - Research Platform, Consumer Devices

United States Digital Space LLC • Deutschland

Hybrid
EUR 105.000 - 141.000
AI Deployment Engineer, Startups
AI Deployment Engineer, Startups

United States Digital Space LLC • Deutschland

Hybrid
EUR 75.000 - 90.000
Relocation assistance
Research Scientist / Research Engineer
Research Scientist / Research Engineer

adaption • Berlin

Vor Ort
EUR 90.000 - 130.000
Flexible work
Travel stipend for exploration
Lunch stipend
+1
Senior Research Scientist
Senior Research Scientist

adaption • Berlin

Hybrid
EUR 90.000 - 130.000
Flexible work
Travel stipend
Lunch stipend
+2
Senior AI Agent Evaluation Engineer
Senior AI Agent Evaluation Engineer

aitrainer • Deutschland

Remote
EUR 34.000 - 61.000
Recruiter, AI/ML Research EMEA
Recruiter, AI/ML Research EMEA

United States Digital Space LLC • Deutschland

Hybrid
EUR 65.000 - 85.000
Relocation assistance
Well-stocked kitchens and in-house meals
Wellness rooms and outdoor spaces
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Berlin

Remote
Competitive hourly rates
Flexible schedule
Experience in advanced AI projects
Research Engineer (Agentic Models)
Research Engineer (Agentic Models)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 90.000 - 120.000
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·

Resaro • München

Hybrid
EUR 90.000 - 130.000
Principal AI Architect (m/f/d)
Principal AI Architect (m/f/d)

Futurice GmbH • München

Vor Ort
EUR 90.000 - 120.000
Sports sponsorship
Health and mobility benefits
Diverse laptop choices
+1