Staff AI Benchmark Architect

Vals AI, Inc.

San Francisco (CA)

Sur place

USD 150 000 - 230 000

Plein temps

14 jours+
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.

Passez les filtres ATS

Avantages offerts par ce poste

Relocation support
Health/dental insurance
Lunch and dinner provided
Unlimited PTO
Housing stipend (within SF)

Résumé du poste

Vals AI, Inc. is seeking exceptional researchers and research engineers to design and build AI benchmarks that push the boundaries of evaluation for foundation models.

You will lead the development of novel benchmarks and collaborate with labs and enterprises to understand evolving evaluation needs. You will publish findings and contribute to the broader evaluation research community while working with an in-house infrastructure team to scale benchmark deployments.

Qualifications

  • Advanced research experience at graduate or post-graduate level.
  • Publication track record in reputable venues (NeurIPS/ICML/ACL/EMNLP).
  • Strong understanding of experimental design and evaluation frameworks.
  • Proficiency in Python for research and experimentation.
  • Clear communication of complex ideas to technical and non-technical audiences.
  • Experience working in research teams and integrating feedback.

Responsabilités

  • Design and develop novel benchmarks that assess real-world LLM capabilities.
  • Conduct research to ensure benchmarks are valid, reliable, and meaningful.
  • Collaborate with labs and enterprises to understand evolving evaluation needs.
  • Analyze model performance across benchmarks and articulate results.
  • Publish research findings and contribute to the evaluation community.
  • Work with the infrastructure team to implement designs at scale.
  • Stay current with developments in LLM capabilities and evaluation methodologies.

Connaissances

Advanced research
Publication track record
Research methodology
Python for research
Communication
Collaboration
Location in SF
Portfolio

Formation

Master's degree or PhD in CS/NLP/ML
Undergraduate with strong research background

Outils

Python

Description du poste

Vals AI, Inc. is seeking exceptional researchers and research engineers to design and build AI benchmarks that push the boundaries of evaluation for foundation models.

You will lead the development of novel benchmarks and collaborate with labs and enterprises to understand evolving evaluation needs. You will publish findings and contribute to the broader evaluation research community while working with an in-house infrastructure team to scale benchmark deployments.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Head of AI Evaluation & Benchmarks
Head of AI Evaluation & Benchmarks

Vibehackers • San Francisco (CA), Northern (KY)

Hybride
USD 225 000 - 275 000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

Sur place
USD 180 000 - 260 000
Equity
Competitive compensation
Technical Writer, AI Benchmarking & Research
Technical Writer, AI Benchmarking & Research

Vals AI, Inc. • San Francisco (CA)

Sur place
USD 110 000 - 160 000
Relocation support
Transportation support
Health, dental, and vision insurance
+4
AI Evaluations Engineer — Benchmarking Frontiers
AI Evaluations Engineer — Benchmarking Frontiers

Meta • Menlo Park (CA)

Sur place
USD 180 000 - 240 000
Senior AI Benchmarking & Systems Architect
Senior AI Benchmarking & Systems Architect

Aionia Group • San Francisco (CA)

Sur place
USD 130 000 - 220 000
Equity
On-site
Senior Member of Technical Staff, AI Benchmarking
Senior Member of Technical Staff, AI Benchmarking

Artificial Analysis, Inc. • San Francisco (CA)

Sur place
USD 100 000 - 150 000
Competitive compensation including equity
Opportunity to shape AI development
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

Sur place
USD 150 000 - 230 000
Relocation support
Health/dental insurance
Lunch and dinner provided
+2
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Ignite Next GmbH • Palo Alto (CA), Northern (KY)

Hybride
USD 120 000 - 190 000
Remote work
Office visits Palo Alto, Paris, Wroclw
Senior AI Evaluation Scientist — Benchmarks & Systems
Senior AI Evaluation Scientist — Benchmarks & Systems

Oracle • Santa Clara (CA)

Sur place
USD 115 000 - 235 000
Medical, dental, and vision insurance
401(k) Savings with company match
Paid time off and holidays
+1
AI Benchmarks & Evaluations Program Manager
AI Benchmarks & Evaluations Program Manager

Mercor • San Francisco (CA)

Sur place
USD 120 000 - 200 000
Performance bonus structure
Equity grant
$15K relocation bonus
+7