AI Evaluation Engineer

Crossing Hurdles

United States

Remote

USD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Crossing Hurdles seeks a senior engineer to design and validate complex simulation problems for AI agents. You will create tasks with multiple constraints, run open-source tools in Python, and test solutions for accuracy and physical plausibility.

The role requires advanced engineering degrees and 10+ years hands-on design experience, plus strong documentation and analytical skills. You will analyze model outputs and iterate benchmarks to improve evaluation results.

Qualifications

  • Masters or PhD in EE/ME/AE with 10+ years hands-on design experience.
  • Proficient with engineering simulation tools and Python scripting.
  • Experience with LLMs, coding agents, AI evaluation, or failure-mode analysis.
  • Strong grasp of engineering principles, unit consistency, and validation.

Responsibilities

  • Design and validate complex engineering simulations for AI agents.
  • Create constrained optimization tasks with validated solutions.
  • Build and test problems using open-source tools and Python.
  • Analyze agent outputs and logs to identify failures and improve benchmarks.
  • Ensure accuracy and physical plausibility of benchmark tasks.

Skills

Engineering simulation
Python scripting
AI evaluation
Failure-mode analysis

Education

Master's degree in EE/ME/AE
PhD in related field

Tools

Open-source simulation tools
Python tooling

Job description

Role & responsibilities
  • Design and validate complex engineering simulation problems for AI agents.
  • Create engineering design tasks with multiple constraints, optimization goals, and validated solutions.
  • Use open-source simulation tools and Python to build and test engineering problems.
  • Analyze AI agent outputs and execution logs to identify reasoning, coding, and simulation failures.
  • Refine benchmark tasks based on model performance and ensure technical accuracy and physical plausibility.

Preferred candidate profile
  • Masters degree or PhD in Electrical, Mechanical, or Aerospace Engineering with 10+ years of hands‑on engineering design experience.
  • Strong experience with engineering simulation tools and Python scripting.
  • Experience with LLMs, coding agents, AI evaluation, or failure‑mode analysis.
  • Strong understanding of engineering principles, physical constraints, unit consistency, and simulation validation.
  • Excellent analytical, problem‑solving, and technical documentation skills.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Expert
Engineering Expert

Turing Global India • United States

Remote
USD 124,000 - 193,000
LLM Engineering Expert (Freelancing)
LLM Engineering Expert (Freelancing)

Zettamine Labs • United States

Remote
USD 165,000 - 276,000
AI Benchmark & Simulation Engineer
AI Benchmark & Simulation Engineer

Crossing Hurdles • United States

Remote
USD 120,000 - 180,000
Freelancing | LLM - Engineering Expert | Remote | Immediate Joiner
Freelancing | LLM - Engineering Expert | Remote | Immediate Joiner

PGC Digital (America) Inc: CMMI Level 3 Company • United States

Remote
USD 138,000 - 248,000
Remote AI Evaluation Engineer for Simulations & Design
Remote AI Evaluation Engineer for Simulations & Design

Benture • San Francisco (CA)

Remote
USD 83,000 - 138,000
Global AI Evaluation Engineer - Simulation & Design
Global AI Evaluation Engineer - Simulation & Design

Dover • United States

Remote
USD 1,200 - 2,500
AI Evaluation Engineer: Model-Breaking Simulation Design
AI Evaluation Engineer: Model-Breaking Simulation Design

PGC Digital (America) Inc: CMMI Level 3 Company • United States

Remote
USD 138,000 - 248,000
Senior Software Engineer, Agent Simulation and Evaluation
Senior Software Engineer, Agent Simulation and Evaluation

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Equity
AI Evaluation Engineer: Simulation & Design (Contract)
AI Evaluation Engineer: Simulation & Design (Contract)

Turing Global India • United States

Remote
USD 124,000 - 193,000
AI/ML Engineer
AI/ML Engineer

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 260,000