AI Evaluation Engineer · Turing Turing · TBD · remote · 3w ago TBD 3w ago

Benture

San Francisco (CA)

Remote

USD 83,000 - 138,000

Part time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Turing, based in San Francisco, seeks experienced AI Evaluation Engineers specializing in engineering simulation and design to author and validate complex problems that train and evaluate state-of-the-art AI agents. This high-impact contractor role operates across Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics Engineering.

Responsibilities include designing model-breaking tasks, integrating environments, analyzing trajectories, calibrating difficulty, and collaborating

Qualifications

  • Master’s degree or PhD in Electrical, Mechanical, Aerospace, Control Systems, Systems Engineering, Robotics, or an applied science discipline.
  • 3+ years of hands-on engineering design experience.
  • Proficiency with at least one domain-relevant open-source simulation package and strong Python scripting skills.
  • Hands-on experience with modern LLMs/coding agents and evaluation concepts.
  • Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
  • Weekend on-call availability (part-time engagement acceptable).
  • Personal desktop or laptop with a stable, high-speed internet connection.

Responsibilities

  • Author original engineering design tasks with constraints, targets, reference solutions, and autograders.
  • Build, run, and validate problem environments using open-source simulation tools and Python test benches.
  • Evaluate agent outputs and logs to identify systemic failure modes.
  • Iteratively refine problem difficulty based on performance data.
  • Collaborate with AI researchers and domain experts to integrate benchmarks into evaluation pipelines.

Skills

Python scripting
Engineering design
LLM evaluation
Domain rigor
Cross-disciplinary collaboration

Education

Master’s or PhD in Electrical/Mechanical/Aerospace/Control Systems/Robotics

Tools

ngspice
PySpice
OpenFOAM
FEniCSx
CalculiX
python-control
CadQuery
OpenModelica
Cantera
Gmsh

Job description

AI Evaluation Engineer (Engineering Simulation & Design) | Contractor | 40 hrs/week | Worldwide Remote

Turing is seeking experienced AI Evaluation Engineers specializing in engineering simulation and design to author and validate complex, model-breaking problems that train and evaluate state-of-the-art AI agents. This is a high-impact contractor role working at the frontier of AI research across disciplines including Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics Engineering.

About Turing

Based in San Francisco, Turing is the world's leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing accelerates frontier research through high-quality data, advanced training pipelines, and top AI researchers specializing in coding, reasoning, STEM, multilinguality, multimodality, and agents.

Key Responsibilities
  • Model-Breaking Problem Design: Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders.
  • Environment & Simulation Integration: Build, run, and validate problem environments using open-source simulation tools and custom Python test benches.
  • Trajectory Analysis & Failure Mode Taxonomy: Evaluate coding agent outputs and execution logs across repeated trials to identify systemic failure modes such as misinterpreted simulator feedback, premature design convergence, or physically impossible geometries.
  • Difficulty Calibration & Benchmark Refinement: Iteratively refine problem difficulty based on empirical model performance data without introducing ambiguity.
  • Cross-Functional Collaboration: Partner with AI researchers, pod leads, and domain experts to integrate high-rigor benchmarks into the model evaluation pipeline.
Requirements
  • Education: Master's degree or PhD in Electrical, Mechanical, Aerospace, Control Systems, Systems Engineering, Robotics, or an applied science discipline.
  • Experience: 3+ years of hands-on engineering design experience.
  • Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package (e.g., ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, OpenModelica, Cantera, Gmsh) and strong Python scripting skills.
  • AI Evaluation: Hands-on experience with modern LLMs/coding agents and evaluation concepts (pass@k, failure-mode analysis, nondeterministic behavior), with the ability to audit trajectory logs and isolate core reasoning and tool-use failures.
  • Domain Rigor: Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
  • Availability: Weekend on-call availability required (part-time engagement acceptable).
  • Infrastructure: Personal desktop or laptop with a stable, high-speed internet connection.
Target Engineering Domains
  • Electrical Engineering
  • Mechanical Engineering
  • Aerospace Engineering
Engagement Details
  • Commitment: 40 hours per week with 4 hours of overlap with PST
  • Engagement Type: Contractor
  • Duration: Up to 24 weeks
Evaluation Process
  • Shortlisted candidates will receive a Job Interest Form.
  • Finalized candidates will undergo a delivery review before proceeding further.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Expert
Engineering Expert

Turing Global India • United States

Remote
USD 124,000 - 193,000
Remote Software Engineer
Remote Software Engineer

Turing • United States

Remote
USD 83,000 - 138,000
Remote Software Engineer
Remote Software Engineer

turing • San Francisco (CA)

Remote
USD 83,000 - 124,000
Software Developer
Software Developer

turing • San Francisco (CA)

Remote
USD 83,000 - 165,000
Remote Software Developer
Remote Software Developer

Turing • United States

Remote
USD 69,000 - 138,000
Remote Senior Software Developer
Remote Senior Software Developer

turing • United States

Remote
USD 83,000 - 152,000
Java Developer
Java Developer

Turing • San Francisco (CA)

On-site
USD 69,000 - 96,000
Flexible hours
Contractor engagement
Remote Software Engineer - C++
Remote Software Engineer - C++

Turing • United States

Remote
USD 55,000 - 110,000
Java Developer
Java Developer

Turing • New York (NY)

On-site
USD 83,000 - 165,000
AI Evaluation Engineer: Simulation & Design (Contract)
AI Evaluation Engineer: Simulation & Design (Contract)

Turing Global India • United States

Remote
USD 124,000 - 193,000