Global AI Evaluation Engineer - Simulation & Design

Dover

United States

Remote

USD 1,200 - 2,500

Part time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Careerflow Human Data Labs is hiring experienced AI Evaluation Engineers to author and validate simulation-based engineering design problems that train and evaluate state‑of‑the‑art AI agents. You will operate across Electrical, Mechanical, Aerospace, and related domains, creating multi‑constraint tasks and automated graders to push frontier models.

The role requires weekend on‑call availability as a contractor, with 30‑40 hours per week and overlap with PST.

Qualifications

  • Education & Expertise: Master’s degree or PhD in Electrical, Mechanical, Aerospace with 10+ years of hands-on engineering design experience.
  • Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package with strong Python scripting.
  • AI Evaluation & Failure Diagnostics: Hands‑on experience with modern LLMs/coding agents and evaluation concepts, ability to audit trajectory logs.
  • Domain Rigor & Precision: Attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and documentation.
  • Availability & Commitment: Weekend on‑call availability (part‑time engagement acceptable).
  • Technical Infrastructure: Remote setup with a stable internet connection.

Responsibilities

  • Model‑Breaking Problem Design: Author self‑contained engineering design tasks with constraints, targets, and autograders.
  • Environment & Simulation Integration: Build, run, and validate environments using open‑source tools and Python test benches.
  • Trajectory Analysis & Failure Mode Taxonomy: Evaluate agent outputs and logs to identify failure modes.
  • Difficulty Calibration & Benchmark Refinement: Refine problem difficulty based on performance data.
  • Cross‑Functional Collaboration: Work with AI researchers and domain experts to integrate benchmarks into evaluation pipelines.

Skills

Engineering design
LLM evaluation concepts
Python scripting
Simulation interpretation

Education

Master’s degree or PhD in Electrical, Mechanical, Aerospace

Tools

ngspice
PySpice
OpenFOAM
FEniCSx
CalculiX
python-control
CadQuery
build123d
OpenModelica
Cantera
Gmsh

Job description

Careerflow Human Data Labs is hiring experienced AI Evaluation Engineers to author and validate simulation-based engineering design problems that train and evaluate state‑of‑the‑art AI agents. You will operate across Electrical, Mechanical, Aerospace, and related domains, creating multi‑constraint tasks and automated graders to push frontier models.

The role requires weekend on‑call availability as a contractor, with 30‑40 hours per week and overlap with PST.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer: Simulation & Design (Contract)
AI Evaluation Engineer: Simulation & Design (Contract)

Turing Global India • United States

Remote
USD 124,000 - 193,000
Remote AI Evaluation Engineer for Simulations & Design
Remote AI Evaluation Engineer for Simulations & Design

Benture • San Francisco (CA)

Remote
USD 83,000 - 138,000
Engineering Expert
Engineering Expert

Turing Global India • United States

Remote
USD 124,000 - 193,000
AI Evaluation Engineer · Turing Turing · TBD · remote · 3w ago TBD 3w ago
AI Evaluation Engineer · Turing Turing · TBD · remote · 3w ago TBD 3w ago

Benture • San Francisco (CA)

Remote
USD 83,000 - 138,000
LLM Engineering Expert- Simulation & Design
LLM Engineering Expert- Simulation & Design

Dover • United States

Remote
USD 1,200 - 2,500
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

EPAM Systems Inc • United States

Remote
USD 140,000 - 190,000
Senior Civil Engineer (Remote) for AI Evaluation Tasks
Senior Civil Engineer (Remote) for AI Evaluation Tasks

YO AI Labs • Maryland

Remote
USD 70,000 - 110,000
Remote Senior Civil Engineer - AI Evaluation Specialist
Remote Senior Civil Engineer - AI Evaluation Specialist

YO AI Labs • Houston (TX)

Remote
USD 55,000 - 124,000
AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Civil Engineer - Remote AI Evaluation Expert
Senior Civil Engineer - Remote AI Evaluation Expert

YO AI Labs • California (MO)

Remote
USD 10,000 - 26,000