LLM Evaluations Engineer — Remote Benchmarking & Tools

Jaide Health

United States

On-site

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance for you and dependents
Company-provided equipment
Wellbeing and home office allowances
Frequent team get togethers
Diverse & inclusive people-first culture

Job summary

Jaide Health is looking for a full-time remote employee who will design and implement infrastructure used by researchers and engineers. Your mission is to ensure progress leads to meaningful improvements for users. Responsibilities include collaborating on model evaluations, defining metrics, and working closely with peers. Ideal candidates will have experience in Large Language Models, strong programming skills, and a curious mindset. The role offers flexible hours and generous vacation days.

Qualifications

  • Strong understanding and intuition of LLMs and their limitations.
  • Good taste and a curious mindset.
  • Use modern tools and always looking to improve.

Responsibilities

  • Research and implement evaluations and benchmarks for both base and instruction following models.
  • Collaborate to define meaningful metrics and evaluations.
  • Work in team, plan future steps, and communicate clearly.

Skills

Experience with Large Language Models (LLM)
Programming experience in multiple languages
Strong algorithmic skills
Strong engineering background
Familiar with full software development life cycle
Critical thinking

Tools

Linux
Python

Job description

Jaide Health is looking for a full-time remote employee who will design and implement infrastructure used by researchers and engineers. Your mission is to ensure progress leads to meaningful improvements for users. Responsibilities include collaborating on model evaluations, defining metrics, and working closely with peers. Ideal candidates will have experience in Large Language Models, strong programming skills, and a curious mindset. The role offers flexible hours and generous vacation days.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer for Scalable ML Evaluations (Remote)
Platform Engineer for Scalable ML Evaluations (Remote)

Jaide Health • United States

On-site
USD 100,000 - 150,000
Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance
+4
Remote Part-Time Java Engineer — LLM Benchmarking
Remote Part-Time Java Engineer — LLM Benchmarking

Jointaro • United States

On-site
Applied Research Engineer – LLMs & Code Gen (Remote)
Applied Research Engineer – LLMs & Code Gen (Remote)

Jaide Health • United States

On-site
USD 120,000 - 150,000
Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance
+3
Member of Engineering (Evaluations)
Member of Engineering (Evaluations)

B Capital • United States

On-site
USD 90,000 - 130,000
Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance for you and dependents
+4
Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
LLM Evaluations Engineer — Benchmark Leaderboards
LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
ML Engineer: LLM Evaluation & Observability
ML Engineer: LLM Evaluation & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
Senior Model Evaluation Engineer – LLM Benchmarking
Senior Model Evaluation Engineer – LLM Benchmarking

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
A weekly lunch stipend of $75/£75 or?e
Full health and dental benefits
RRSP matching, 401K, Pension Scheme
+2
AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home
Remote AI Math Evaluator & LLM Benchmark Designer
Remote AI Math Evaluator & LLM Benchmark Designer

United States Digital Space LLC • United States

Remote
Fully remote
AI projects
Contract extension potential