Evaluation Lead

SupportFinity™

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A cutting-edge AI technology firm in San Francisco is seeking an Evaluation Lead to drive the assessment of AI model performance. You will design evaluation methodologies, automate evaluation processes, and oversee various evaluation strategies. The ideal candidate has extensive experience in AI model evaluation and is proficient in Python. This high-impact role demands strong collaboration skills and a startup-ready mindset to thrive in a fast-paced environment.

Qualifications

  • Extensive expertise in evaluating AI and machine learning models, ideally in physical AI.
  • Experience in designing, implementing, and refining evaluation metrics.
  • Deep understanding of machine learning, AI, and generative models.
  • Excellent Python and software engineering skills.
  • Experience designing and building scalable data pipelines and evaluation tools.
  • Experience collaborating closely with key stakeholders from research, engineering, and product teams.
  • Strong communication and documentation skills, with a bias for creating detailed evaluation reports.
  • Startup-ready mindset with the ability to thrive in high-velocity environments.

Responsibilities

  • Design and implement evaluation methodologies and benchmarks for model effectiveness.
  • Build and oversee pipelines and tools that automate model evaluation.
  • Develop strategies for evaluating physical AI models across various use cases.

Skills

Evaluating AI and machine learning models
Designing evaluation metrics
Understanding of machine learning
Python programming
Building scalable data pipelines
Strong communication skills
Collaboration with teams
Startup mindset

Job description

About Archetype AI

Archetype AI is developing the world's first AI platform to bring AI into the real world. Formed by an exceptionally high-caliber team from Google, Archetype AI is building a foundation model for the physical world, a real‑time multimodal LLM for real life, transforming real‑world data into valuable insights and knowledge that people will be able to interact with naturally. It will help people in their real lives, not just online, because it understands the real‑time physical environment and everything that happens in it.

Supported by deep tech venture funds in Silicon Valley, Archetype AI is currently at the Series A stage and is progressing rapidly to develop technology for its next stage. This presents a unique and once‑in‑a‑lifetime opportunity to be part of an exciting AI team at the beginning of their journey, located in the heart of Silicon Valley. Our team is headquartered in San Mateo, California, with team members throughout the US and Europe. We are actively growing, so if you are an exceptional candidate excited to work on the cutting edge of physical AI and don’t see a role that exactly fits you below you can contact us directly with your resume via jobsarchetypeaiio.

About The Role

Archetype AI is seeking a hands‑on Evaluation Lead to build and assess model performance for physical AI. You will design and implement advanced evaluation techniques for assessing the strengths and weaknesses of real‑world AI models, and build and scale evaluation frameworks to rapidly test and generate reports on model performance. Responsibilities include partnering closely with research and engineering teams to develop evaluation methodologies, analytically assessing and improving test datasets, uncovering model weaknesses or risks, and tracking competitive industry benchmarks. This is a high‑impact role for someone who thrives in a fast‑paced AI environment and wants to directly influence our path as we scale our AI technologies and business.

Core Responsibilities
  • Drive Benchmarking & Evaluation
    • Design and implement rigorous evaluation methodologies and benchmarks for measuring model effectiveness, reliability, alignment, and safety
    • Lead evaluation of model performance, ranging from offline experiments to full production model testing
  • Build & Scale Evaluation Frameworks
    • Design and oversee the pipelines, dashboards, and tools that automate model evaluation
    • Design and oversee tools for A/B model testing, regression testing, and production model performance
  • Lead Evaluation Strategy
    • Develop and implement strategies for evaluating physical AI models that can scale across a broad range of real‑world use cases, sensor types, and edge cases
    • Plan, run, and oversee evaluations, across internal teams and external customers
    • Drive edge case discovery, red‑teaming, safety, privacy, and risk evaluation – feeding back knowledge to key stakeholders in research and engineering teams
Key Requirements
  • Extensive expertise in evaluating AI and machine learning models, ideally in physical AI or a related AI field
  • Experience in designing, implementing, and refining evaluation metrics
  • Deep understanding of machine learning, AI, and generative models
  • Excellent Python and software engineering skills
  • Experience designing and building scalable data pipelines and evaluation tools
  • Experience collaborating closely with key stakeholders from research, engineering, and product teams
  • Strong communication and documentation skills, with a bias for creating detailed evaluation reports that help drive model performance
  • Startup‑ready mindset with the ability to thrive in high‑velocity, high‑ambiguity environments
Minimum Qualifications
  • Extensive expertise in evaluating AI and machine learning models, ideally in physical AI or a related AI field
  • Experience in designing, implementing, and refining evaluation metrics
  • Deep understanding of machine learning, AI, and generative models
  • Excellent Python and software engineering skills
  • Experience designing and building scalable data pipelines and evaluation tools
  • Experience collaborating closely with key stakeholders from research, engineering, and product teams
  • Strong communication and documentation skills, with a bias for creating detailed evaluation reports that help drive model performance
  • Startup‑ready mindset with the ability to thrive in high‑velocity, high‑ambiguity environments
What We Would Love To See
  • Experience evaluating real‑world, real‑time algorithms
  • Experience evaluating a broad range of sensor types, such as cameras, LIDAR, physical sensors, RF sensors, and beyond
  • A strong scientific approach to evaluation and understanding model performance
  • Experience in evaluating production algorithms
  • Experience building and curating data campaigns to create extensive test datasets
  • Experience managing internal teams and/or external vendors
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 230,000
Equity
Frontier AI exposure
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Applied AI Researcher
Applied AI Researcher

Morpheus Talent Solutions • United States

On-site
USD 120,000 - 180,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Data Scientist, AI Evaluation & Insights
Data Scientist, AI Evaluation & Insights

Arena • San Francisco (CA)

On-site
Machine Learning Applied Researcher
Machine Learning Applied Researcher

Archetype AI • San Mateo (CA)

On-site
USD 120,000 - 180,000
Pre-training Distributed Systems Tech Lead / Manager
Pre-training Distributed Systems Tech Lead / Manager

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000