AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor

Miami (FL)

Sur place

USD 83 000 - 138 000

Temps partiel

14 jours+
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Résumé du poste

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will develop original, executable research problems that frontier models cannot solve.

Expect depth in multiple subdomains, such as numerical linear algebra, computational mechanics, or computational finance, with strong Python and Docker/GitHub fluency to ensure high-quality, reproducible runs within a PR workflow.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsabilités

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Connaissances

Python for scientific computing
Git/GitHub proficiency
Docker workflows

Formation

PhD in mathematics, applied mathematics, computational mathematics, or closely related field
Master's degree in a related field

Outils

Docker
GitHub
Kaggle

Description du poste

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will develop original, executable research problems that frontier models cannot solve.

Expect depth in multiple subdomains, such as numerical linear algebra, computational mechanics, or computational finance, with strong Python and Docker/GitHub fluency to ensure high-quality, reproducible runs within a PR workflow.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

Sur place
USD 11 021 000 - 13 776 000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

Sur place
USD 96 000 - 179 000
AI Evaluation Scientist (Math PhD) & Trainer
AI Evaluation Scientist (Math PhD) & Trainer

Mercor • Los Angeles (CA)

Sur place
USD 110 000 - 165 000
AI Benchmark Designer: Math PhD & Trainer
AI Benchmark Designer: Math PhD & Trainer

Obsidian • Los Angeles (CA)

Sur place
USD 83 000 - 124 000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Obsidian • San Francisco (CA)

Sur place
USD 96 000 - 179 000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

Sur place
USD 11 021 000 - 13 776 000
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

Sur place
USD 83 000 - 124 000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Miami (FL)

Sur place
USD 83 000 - 124 000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

Sur place
USD 83 000 - 138 000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Los Angeles (CA)

Sur place
USD 83 000 - 124 000