AI Evaluation Engineer: Model-Breaking Simulation Design

PGC Digital (America) Inc: CMMI Level 3 Company

United States

Remote

USD 138,000 - 248,000

Full time

32 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

PGC Digital (America) Inc: CMMI Level 3 Company seeks experienced AI Evaluation Engineers to author and validate model-breaking, simulation-based design problems across Electrical and Mechanical domains. You will create complex tasks, assess AI agents, and build objective graders to push frontier model performance.

Requirements include a Master’s or PhD in Electrical/Mechanical/Aerospace with 10+ years hands-on design experience, strong Python and Open-Source simulation tooling, and weekend

Qualifications

  • Master’s degree or PhD in Electrical, Mechanical, or Aerospace engineering with 10+ years hands-on engineering design experience.
  • Proficiency with at least one domain-relevant open-source simulation package and strong Python scripting skills.
  • Hands-on experience with modern LLMs/coding agents and evaluation concepts, with ability to audit trajectory logs.
  • Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
  • Talent must have weekend on-call availability (part-time engagement is acceptable).
  • Personal desktop/laptop with stable high-speed internet in a remote setup.

Responsibilities

  • Design original, self-contained engineering tasks with competing constraints and validated reference solutions.
  • Build, run, and validate environments using open-source simulation tools and Python test benches.
  • Evaluate agent outputs and logs to identify systemic failure modes.
  • Iteratively calibrate task difficulty based on empirical model performance data.
  • Collaborate with AI researchers, pod leads, and domain experts to integrate high-rigor benchmarks into the evaluation pipeline.

Skills

Python scripting
AI evaluation concepts
Failure diagnostics
Simulation problem design
Domain rigor

Education

Master’s degree or PhD in Electrical/Mechanical/Aerospace engineering

Tools

ngspice
PySpice
OpenFOAM
FEniCSx
CalculiX
python-control
CadQuery
build123d
OpenModelica
Cantera
Gmsh

Job description

PGC Digital (America) Inc: CMMI Level 3 Company seeks experienced AI Evaluation Engineers to author and validate model-breaking, simulation-based design problems across Electrical and Mechanical domains. You will create complex tasks, assess AI agents, and build objective graders to push frontier model performance.

Requirements include a Master’s or PhD in Electrical/Mechanical/Aerospace with 10+ years hands-on design experience, strong Python and Open-Source simulation tooling, and weekend

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Expert
Engineering Expert

Turing Global India • United States

Remote
USD 124,000 - 193,000
AI Evaluation Engineer
AI Evaluation Engineer

Crossing Hurdles • United States

Remote
USD 120,000 - 180,000
Remote AI Evaluation Engineer for Simulations & Design
Remote AI Evaluation Engineer for Simulations & Design

Benture • San Francisco (CA)

Remote
USD 83,000 - 138,000
AI Evaluation Engineer: Simulation & Design (Contract)
AI Evaluation Engineer: Simulation & Design (Contract)

Turing Global India • United States

Remote
USD 124,000 - 193,000
AI Benchmark & Simulation Engineer
AI Benchmark & Simulation Engineer

Crossing Hurdles • United States

Remote
USD 120,000 - 180,000
Global AI Evaluation Engineer - Simulation & Design
Global AI Evaluation Engineer - Simulation & Design

Dover • United States

Remote
USD 1,200 - 2,500
Freelancing | LLM - Engineering Expert | Remote | Immediate Joiner
Freelancing | LLM - Engineering Expert | Remote | Immediate Joiner

PGC Digital (America) Inc: CMMI Level 3 Company • United States

Remote
USD 138,000 - 248,000
Senior Software Engineer, Agent Simulation and Evaluation
Senior Software Engineer, Agent Simulation and Evaluation

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Equity
AI Evaluation Engineer · Turing Turing · TBD · remote · 3w ago TBD 3w ago
AI Evaluation Engineer · Turing Turing · TBD · remote · 3w ago TBD 3w ago

Benture • San Francisco (CA)

Remote
USD 83,000 - 138,000
Electrical Engineer: AI Model Evaluator & Designer
Electrical Engineer: AI Model Evaluator & Designer

DataAnnotation • Redondo Beach (CA)

Remote
USD 55,000 - 172,000