AI Evaluation Lead: Real-World Systems Benchmarking
SupportFinity™
San Francisco (CA)
On-site
USD 150,000 - 230,000
Full time
14 days+
Application generator
Turn this role into an interview — a resume and cover letter built around what this employer wants.
Get past ATS filters
Job summary
A cutting-edge AI technology firm in San Francisco is seeking an Evaluation Lead to drive the assessment of AI model performance. You will design evaluation methodologies, automate evaluation processes, and oversee various evaluation strategies. The ideal candidate has extensive experience in AI model evaluation and is proficient in Python. This high-impact role demands strong collaboration skills and a startup-ready mindset to thrive in a fast-paced environment.
Qualifications
Extensive expertise in evaluating AI and machine learning models, ideally in physical AI.
Experience in designing, implementing, and refining evaluation metrics.
Deep understanding of machine learning, AI, and generative models.
Responsibilities
Design and implement evaluation methodologies and benchmarks for model effectiveness.
Build and oversee pipelines and tools that automate model evaluation.
Develop strategies for evaluating physical AI models across various use cases.
Skills
Evaluating AI and machine learning models
Designing evaluation metrics
Understanding of machine learning
Python programming
Building scalable data pipelines
Strong communication skills
Collaboration with teams
Job description
A cutting-edge AI technology firm in San Francisco is seeking an Evaluation Lead to drive the assessment of AI model performance. You will design evaluation methodologies, automate evaluation processes, and oversee various evaluation strategies. The ideal candidate has extensive experience in AI model evaluation and is proficient in Python. This high-impact role demands strong collaboration skills and a startup-ready mindset to thrive in a fast-paced environment.