Applied AI Research Intern: Benchmark & Evaluation

Labelbox

San Francisco (CA)

Hybrid

USD 35,000 - 45,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Career advancement opportunities
Hybrid work model
Fast-paced environment

Job summary

A cutting-edge technology firm seeks an Applied Research intern to design and build evaluation systems for AI models. You'll create post-training datasets and prototype innovative training loops to improve real-world performance. Required qualifications include a strong background in AI, a relevant degree, and proficiency in Python and popular deep learning frameworks. Join a dynamic team focused on pushing the boundaries of AI and contributing to transformative technology in a hybrid work environment.

Qualifications

  • Strong foundation in AI and machine learning.
  • Deep understanding of multimodal models & data strategies.
  • Passion for LLM evaluation and benchmarking.

Responsibilities

  • Design and build evaluation and benchmark suites.
  • Create post-training datasets at scale.
  • Prototype RLHF-style training loops to improve performance.

Skills

Strong foundation in AI and machine learning
Proficiency in Python
Expertise in training data quality construction
Exceptional communication and collaboration skills

Education

Ph.D. or Master's degree in Computer Science, Machine Learning, AI

Tools

PyTorch
JAX
TensorFlow

Job description

A cutting-edge technology firm seeks an Applied Research intern to design and build evaluation systems for AI models. You'll create post-training datasets and prototype innovative training loops to improve real-world performance. Required qualifications include a strong background in AI, a relevant degree, and proficiency in Python and popular deep learning frameworks. Join a dynamic team focused on pushing the boundaries of AI and contributing to transformative technology in a hybrid work environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI Model Evaluation Engineer — Benchmarking & Validation
AI Model Evaluation Engineer — Benchmarking & Validation

SpreeAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options
Remote AI Benchmark Engineer & Researcher
Remote AI Benchmark Engineer & Researcher

Pathway • Palo Alto (CA)

Remote
USD 120,000 - 180,000
AI Benchmarking Engineer — Evaluations & Failure Analysis
AI Benchmarking Engineer — Evaluations & Failure Analysis

Mercor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
+2
AI Benchmarks & Evaluations Program Manager
AI Benchmarks & Evaluations Program Manager

Mercor • San Francisco (CA)

On-site
USD 120,000 - 200,000
Performance bonus structure
Equity grant
$15K relocation bonus
+7
Edge AI Intern: Real-Time, Efficient Foundation Models
Edge AI Intern: Real-Time, Efficient Foundation Models

Mitsubishi Electric Research Laboratories • Cambridge (MA)

On-site
Relocation stipend
Covered travel to and from MERL
Monthly Charlie Card for commuting
+2
Summer Research Intern - Build Open-Source AI Benchmarks
Summer Research Intern - Build Open-Source AI Benchmarks

Abaka AI • Mountain View (CA)

On-site
Paid internship
Onsite in Palo Alto
AI/ML Intern: Hands-on Model Building & Validation
AI/ML Intern: Hands-on Model Building & Validation

cloud-ai-labs • United States

On-site
USD 20,664 - 34,440
AI/ML Intern: Hands-on Model Building & Validation
AI/ML Intern: Hands-on Model Building & Validation

alphasoftwaretech • United States

On-site
USD 20,000 - 40,000