Senior Machine Learning Engineer, AI Evaluation

Society for Human Resource Management (SHRM)

Alexandria (VA)

Hybrid

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
Well-being programs
Health savings
Flexible spending
Retirement plan
Annual discretionary bonus

Job summary

SHRM is seeking a Senior Machine Learning Engineer, AI Evaluation to design and operate the measurement infrastructure for Applied AI Research. You will build scalable evaluation systems, manage model-versioning, and ensure reproducible benchmarks across multiple AI models in a collaborative, cross-functional environment.

You will partner with HR experts to translate professional standards into technically executable evaluation specifications and maintain auditable data infrastructure for

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Engineering, or related field; relevant experience accepted.
  • 7+ years of ML/LLM engineering, applied AI, or research infrastructure experience.
  • Hands-on experience with multi-model LLM applications and evaluation frameworks.
  • Strong knowledge of AI evaluation methodologies and reproducible systems.

Responsibilities

  • Design, build, and maintain scalable engineering infrastructure for structured evaluations across AI/LLM families.
  • Develop a unified orchestration layer for provider-agnostic evaluation across frontier models.
  • Create robust evaluation and scoring frameworks with rubric-based scoring and model-as-judge approaches.
  • Ensure reproducible experimentation with version pinning, logging, tracking, and drift detection.
  • Collaborate with HR SMEs to translate standards into measurable evaluation specifications.

Skills

Python
ML/LLM
Data pipelines
Cloud platforms
Experimentation

Education

Bachelor's degree in CS/DS/Engineering
Master's degree preferred

Tools

BigQuery
Looker Looker Studio
Vertex AI
Git

Job description

Senior Machine Learning Engineer, AI Evaluation
Job Description

Posted Wednesday, August 19, 2026 at 4:00 AM

SHRMis a member-driven catalystfor creating better workplaces wherepeople and businessesthrive together.As thetrusted authority on all things work, SHRM is the foremost expert, researcher, advocate, and thoughtleader on issuesand innovationsimpacting today’s evolving workplaces.With nearly340,000members in180countries, SHRM touches the lives of more than362million workers and their families globally.

Summary

The Senior Machine Learning Engineer, AI Evaluation builds and operates the measurement and engineering infrastructure supporting the organization's Applied AI Research (AAIR) function, a continuous experimental environment designed to evaluate how artificial intelligence models perform real-world HR and workplace-related tasks against established professional standards.

This role is responsible for designing and maintaining the engineering infrastructure used to conduct rigorous, reproducible AI model evaluations and benchmarks. The Senior Machine Learning Engineer develops the systems that run multiple AI models against structured, domain-specific evaluations; builds scoring and evaluation frameworks; maintains reproducibility across model versions; and creates the data infrastructure necessary to analyze and track model performance over time.

Working closely with HR subject matter experts and Applied AI Research colleagues, this position translates expert-defined standards and evaluation criteria into technically rigorous, measurable specifications. Subject matter experts establish the domain-specific ground truth and standards for correctness, while the Senior Machine Learning Engineer owns the technical systems and methodologies used to measure model performance against those standards.

The position serves as a shared technical engineering resource across multiple Applied AI Research workstreams and helps ensure that published findings, benchmarks, and research conclusions are supported by reliable, auditable, and defensible measurement practices.

This is an AI evaluation and engineering infrastructure role rather than a model-training or frontier AI research position . The role is not responsible for developing novel model architectures or training foundation models.

Responsibilities
AI Evaluation Engineering & Infrastructure
  • Design, build, and maintain scalable engineering infrastructure for conducting structured evaluations and experiments across multiple AI and large language model (LLM) families.
  • Develop and maintain a unified, provider-agnostic orchestration layer that enables consistent evaluation across multiple frontier model providers and architectures.
  • Design and implement rigorous AI evaluation and scoring frameworks, including rubric-based scoring, model-as-judge methodologies with appropriate safeguards, partial-credit methodologies, and approaches for managing ambiguity.
  • Build systems and processes that support reproducible experimentation, including model-version pinning, comprehensive run logging, experiment tracking, and drift detection.
  • Maintain portable evaluation architecture across AI model providers to enable consistent and defensible cross-model comparisons as models and technologies evolve.
  • Establish and maintain technical standards and engineering practices that support reliable, repeatable, and auditable AI evaluation.
  • Partner closely with HR subject matter experts to translate professional standards, research criteria, and judgment-based rubrics into measurable and technically executable evaluation specifications.
  • Identify and surface ambiguity, inconsistencies, or measurement limitations within proposed evaluation criteria and collaborate with subject matter experts to strengthen evaluation design.
  • Apply knowledge of AI evaluation methodologies, benchmarking techniques, inter-rater reliability, and known limitations of automated and model-as-judge evaluation approaches.
  • Support the design of measurement methodologies when definitive ground truth is unavailable or requires expert interpretation.
  • Ensure evaluation methodologies align with research-defined validation standards and produce findings that are reproducible, transparent, and defensible.
  • Contribute technical expertise to the design and continuous improvement of AI research experiments, benchmarks, and evaluation methodologies.
Data, Monitoring & Reproducibility
  • Design and maintain structured repositories for experiment results, prompt libraries, scoring rubrics, model metadata, and longitudinal evaluation data using cloud-based data infrastructure.
  • Build and maintain data structures in BigQuery or comparable platforms that enable research results to be queried, analyzed, reproduced, and audited.
  • Develop monitoring, reporting, and visualization capabilities using Looker, Looker Studio, or comparable tools to provide visibility into experiment status, model performance, and performance drift.
  • Establish processes for tracking changes in model behavior across model versions and over time.
  • Maintain complete technical documentation and metadata necessary to reproduce research findings and evaluation results.
  • Ensure appropriate quality controls are incorporated throughout data collection, evaluation, scoring, storage, and reporting processes.
Cross-Functional Technical Partnership
  • Serve as a shared engineering resource across multiple Applied AI Research teams and research workstreams.
  • Collaborate with research leaders, HR subject matter experts, data professionals, engineers, and other internal stakeholders to translate research requirements into scalable technical solutions.
  • Support collaboration with university, affiliate, research, and other external partners when appropriate and within established organizational access controls and data-handling requirements.
  • Communicate technical methodologies, limitations, risks, and findings clearly to both technical and non-technical audiences.
  • Evaluate emerging AI models, tools, technologies, and evaluation methodologies and recommend appropriate applications within the research environment.
  • Contribute to continuous improvement of the Applied AI Research technical environment, engineering practices, and evaluation capabilities.
Data Governance, Security & Responsible AI
  • Ensure AI evaluation systems and workflows comply with organizational requirements for data security, privacy, ownership, access, and responsible AI use.
  • Maintain appropriate controls to protect proprietary, member, research, and other sensitive data from unauthorized access or use.
  • Ensure organizational data is not used to train external shared models except where expressly authorized and appropriately governed.
  • Implement and maintain appropriate access controls and data-handling requirements when working with external research or partner organizations.
  • Partner with appropriate internal stakeholders to ensure evaluation infrastructure aligns with organizational technology, security, privacy, and governance standards.
Education & Experience Requirements
Education
  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Engineering, or a related quantitative or technical field, or relevant equivalent experience in lieu of degree.
  • Master's degree in Computer Science, Data Science, Machine Learning, Artificial Intelligence, or a related field preferred.
  • A Ph.D. is not required; demonstrated expertise in AI/ML evaluation engineering, research infrastructure, and production-grade systems is valued.
Experience
  • Seven (7) or more years of progressively responsible experience in ML/LLM engineering, applied AI, applied data science, or research infrastructure, including experience developing, implementing, and supporting production-grade systems.
  • Demonstrated hands-on experience developing multi-model LLM applications and infrastructure, including provider-agnostic model access, APIs, prompt engineering, and evaluation frameworks.
  • Demonstrated experience with AI model evaluation and benchmarking, including rubric-based scoring, inter-rater reliability, model-as-judge methodologies and their limitations, and approaches for evaluating performance when definitive ground truth may not be readily available.
  • Experience designing and maintaining reproducible technical systems incorporating version control, model-version pinning, comprehensive logging, experiment tracking, and/or drift detection.
  • Experience with cloud-based data and AI infrastructure on a major cloud platform; Google Cloud Platform experience, including BigQuery, Vertex AI, IAM, and audit logging, preferred.
  • Experience developing and supporting data pipelines, structured experiment repositories, dashboards, or monitoring solutions.
  • Experience working with HR, workforce, survey, behavioral, or other professional-domain data preferred.
  • Experience supporting academic, applied research, benchmarking, or peer-reviewed research workflows preferred.
Certifications
Knowledge, Skills & Abilities
  • Advanced proficiency in Python and strong software-engineering fundamentals, including the ability to develop reliable, maintainable, production-quality code.
  • Strong knowledge of machine learning, large language models, generative AI systems, and contemporary AI application architectures.
  • Demonstrated knowledge of AI evaluation and benchmarking methodologies, including rubric-based evaluation, automated scoring, model-as-judge approaches, inter-rater reliability, and measurement design.
  • Strong understanding of the limitations and failure modes of generative AI systems and the ability to design evaluation approaches that appropriately account for those limitations.
  • Demonstrated commitment to reproducibility, including disciplined use of versioning, documentation, logging, experiment tracking, and drift detection.
  • Ability to translate complex, judgment-based requirements from subject matter experts into technically rigorous and measurable evaluation specifications without oversimplifying the underlying domain expertise.
  • Strong analytical and problem-solving skills with the ability to identify technical, methodological, and data-quality issues and develop appropriate solutions.
  • Working knowledge of cloud-based AI and data environments, APIs, data warehouses, access controls, and related technical infrastructure.
  • Ability to effectively communicate complex technical concepts, methodologies, limitations, and findings to technical and non-technical audiences.
  • Strong collaboration and consultation skills, with the ability to work effectively with researchers, engineers, data professionals, subject matter experts, and external partners.
  • Ability to balance technical rigor, research requirements, scalability, and practical implementation considerations.
  • Strong understanding of responsible AI principles, data governance, privacy, security, and appropriate handling of sensitive or proprietary information.
  • Ability to evaluate emerging AI models, technologies, and evaluation methodologies and determine their appropriate application within the organization's research environment.
  • Ability to effectively leverage AI tools and technologies to streamline workflows, enhance productivity, and improve overall work quality.
Physical Requirements

This position operates in a typical office environment (which includes a home office setting) and requires the ability to perform essential job functions with or without reasonable accommodation. Physical requirements may include:

  • Prolonged periods of sitting at a desk and working on a computer.
  • Frequent use of hands and fingers for typing, handling documents, and using office equipment.
  • Occasional standing, walking, bending, and reaching.
  • Ability to lift and carry up to 30 pounds as needed.
  • Clear verbal and written communication skills for effective interaction with colleagues and stakeholders.
Hybrid Schedule (3 Days In-Office/2 Days Remote)

This position follows a hybrid work schedule, with Tuesday through Thursday in office and Monday and Friday remote. Employees must be available during standard business hours, with core hours beginning between 8:00–9:00 a.m. and concluding between 5:00–6:00 p.m. local time.

Travel: Occasional 0 – 10%

#LI

The hiring range for this position is $100,000 to $130,000 per year. This range is an estimate, and the actual salary may vary based on the candidate's experience, skills, and qualifications. SHRM offers a competitive and comprehensive total rewards package. The benefits for this position include professional growth and development, health, dental, vision, well-being, health savings, flexible spending, retirement, open leave, and annual discretionary bonus and incentives.

Our employment practices are in accordance with the laws that prohibit discrimination against qualified individuals on the basis of race, religion, color, gender, age, national origin, physical or mental disability, genetic information, veteran’s status, marital status, gender identity and expression, sexual orientation, or any other status protected by applicable law.

SHRM is an equal opportunity employer (Minority/Female/Disabled/Veteran).

We do not sponsor applicants for work visas.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer, AI Evaluation
Senior Machine Learning Engineer, AI Evaluation

SHRM • Alexandria (VA)

Hybrid
USD 100,000 - 130,000
Health benefits
Retirement plan
Bonuses & incentives
+1
Senior ML Engineer: AI Evaluation & Reproducible Infra
Senior ML Engineer: AI Evaluation & Reproducible Infra

Society for Human Resource Management (SHRM) • Alexandria (VA)

Hybrid
USD 100,000 - 130,000
Health benefits
Dental benefits
Vision benefits
+5
Senior Specialist, HR Tech Content
Senior Specialist, HR Tech Content

SHRM • Alexandria (VA)

Hybrid
USD 85,000 - 100,000
Hybrid Schedule
Health benefits
Retirement plan
+3
Senior Specialist, HR Tech Content
Senior Specialist, HR Tech Content

Society for Human Resource Management (SHRM) • Alexandria (VA)

Hybrid
USD 85,000 - 100,000
Health benefits
Dental
Vision
+4
Director, Operations Center for Inclusion & Diversity
Director, Operations Center for Inclusion & Diversity

Society for Human Resource Management (SHRM) • Alexandria (VA)

Hybrid
USD 150,000 - 175,000
Professional growth and development
Health, dental, vision benefits
Retirement benefits
Senior Machine Learning Engineer, Analytics Center of Excellence (Remote/WFH)
Senior Machine Learning Engineer, Analytics Center of Excellence (Remote/WFH)

Latitude • Durham (NC), Northern (KY)

Hybrid
USD 111,000 - 279,000
Senior Staff Software Engineer, AI Infrastructure
Senior Staff Software Engineer, AI Infrastructure

LinkedIn • Sunnyvale (CA)

Hybrid
USD 198,000 - 326,000
Machine Learning Engineer III
Machine Learning Engineer III

Workday • San Francisco (CA)

Hybrid
USD 160,000 - 240,000
Senior Staff Machine Learning Engineer, Data & Eval
Senior Staff Machine Learning Engineer, Data & Eval

Traveltechessentialist • United States

On-site
USD 150,000 - 210,000
Bonus
Equity
Employee Travel Credits
Machine Learning Scientist - Open Source Lead
Machine Learning Scientist - Open Source Lead

LMArena • San Francisco (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Comprehensive health and wellness benefits
Opportunity to work on cutting-edge AI
+1