AI Evaluation Engineer - Reproducible ML Benchmarks Hybrid

SHRM

Alexandria (VA)

Hybrid

USD 100,000 - 130,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Retirement plan
Bonuses & incentives
Professional growth

Job summary

SHRM seeks a Senior Machine Learning Engineer, AI Evaluation to design and operate the measurement and engineering infrastructure for AAIR, supporting rigorous evaluation of AI models against professional standards.

You will build scalable evaluation systems, versioned experiments, and data pipelines while partnering with HR experts to translate domain criteria into measurable specs; this role serves multiple AI research workstreams with a focus on reproducibility and auditable findings.

Qualifications

  • 7+ years of experience in ML/LLM engineering, applied AI, or research infrastructure.
  • Hands-on experience with multi-model LLM applications and infrastructure.
  • Experience with AI model evaluation/benchmarking, rubric-based scoring, and inter-rater reliability.
  • Cloud-based data/AI infrastructure experience including BigQuery and Vertex AI.

Responsibilities

  • Design, build, and maintain scalable engineering infrastructure for evaluations across multiple AI/LLM families.
  • Develop a provider-agnostic orchestration layer for consistent evaluation across frontier model providers.
  • Design and implement rigorous AI evaluation and scoring frameworks, including rubric-based scoring and model-as-judge methods.
  • Build systems for reproducible experimentation with versioning, logging, and drift detection.
  • Maintain data infrastructure for experiment results, prompts, rubrics, and metadata.
  • Collaborate with HR experts to translate standards into measurable evaluation specifications.
  • Communicate methodologies and findings clearly to technical and non-technical audiences.

Skills

Python
ML/LLM
AI evaluation
Experiment tracking
Cloud platforms
Communication

Education

Bachelor's in Computer Science/Data Science/ML/Engineering
Master's preferred

Tools

BigQuery
Vertex AI
Looker Studio

Job description

SHRM seeks a Senior Machine Learning Engineer, AI Evaluation to design and operate the measurement and engineering infrastructure for AAIR, supporting rigorous evaluation of AI models against professional standards.

You will build scalable evaluation systems, versioned experiments, and data pipelines while partnering with HR experts to translate domain criteria into measurable specs; this role serves multiple AI research workstreams with a focus on reproducibility and auditable findings.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: AI Evaluation & Reproducible Infra
Senior ML Engineer: AI Evaluation & Reproducible Infra

Society for Human Resource Management (SHRM) • Alexandria (VA)

Hybrid
USD 100,000 - 130,000
Health benefits
Dental benefits
Vision benefits
+5
Senior Machine Learning Engineer, AI Evaluation
Senior Machine Learning Engineer, AI Evaluation

SHRM • Alexandria (VA)

Hybrid
USD 100,000 - 130,000
Health benefits
Retirement plan
Bonuses & incentives
+1
Senior Machine Learning Engineer, AI Evaluation
Senior Machine Learning Engineer, AI Evaluation

Society for Human Resource Management (SHRM) • Alexandria (VA)

Hybrid
USD 100,000 - 130,000
Health benefits
Dental benefits
Vision benefits
+5
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Senior AI Evaluation Engineer — Metrics & Data Pipelines
Senior AI Evaluation Engineer — Metrics & Data Pipelines

Sentry • San Francisco (CA)

Hybrid
USD 240,000 - 280,000
Equity grants
Paid time off
Group health insurance coverage
Senior AI/ML Evaluation & Benchmark Engineer
Senior AI/ML Evaluation & Benchmark Engineer

OpenTeams • Colorado

Hybrid
USD 145,000 - 250,000
401(k) Match
Unlimited PTO
Fully Remote Setup
+3
AI Evaluation Scientist - Real-World Model Insights
AI Evaluation Scientist - Real-World Model Insights

arena • United States

On-site
USD 120,000 - 190,000
Competitive compensation
Senior AI Evaluation Engineer - Remote Data Pipelines
Senior AI Evaluation Engineer - Remote Data Pipelines

Socket.dev • United States

On-site
USD 196,000 - 227,000
Senior AI/ML Benchmarking & Evaluation Engineer
Senior AI/ML Benchmarking & Evaluation Engineer

openteams • Washington

Hybrid
USD 145,000 - 250,000
Staff AI Evaluation Architect
Staff AI Evaluation Architect

Uncover • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000