LLM Engineering Expert- Simulation & Design

Dover

United States

Remote

USD 1,200 - 2,500

Part time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Careerflow Human Data Labs is hiring experienced AI Evaluation Engineers to author and validate simulation-based engineering design problems that train and evaluate state‑of‑the‑art AI agents. You will operate across Electrical, Mechanical, Aerospace, and related domains, creating multi‑constraint tasks and automated graders to push frontier models.

The role requires weekend on‑call availability as a contractor, with 30‑40 hours per week and overlap with PST.

Qualifications

  • Education & Expertise: Master’s degree or PhD in Electrical, Mechanical, Aerospace with 10+ years of hands-on engineering design experience.
  • Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package with strong Python scripting.
  • AI Evaluation & Failure Diagnostics: Hands‑on experience with modern LLMs/coding agents and evaluation concepts, ability to audit trajectory logs.
  • Domain Rigor & Precision: Attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and documentation.
  • Availability & Commitment: Weekend on‑call availability (part‑time engagement acceptable).
  • Technical Infrastructure: Remote setup with a stable internet connection.

Responsibilities

  • Model‑Breaking Problem Design: Author self‑contained engineering design tasks with constraints, targets, and autograders.
  • Environment & Simulation Integration: Build, run, and validate environments using open‑source tools and Python test benches.
  • Trajectory Analysis & Failure Mode Taxonomy: Evaluate agent outputs and logs to identify failure modes.
  • Difficulty Calibration & Benchmark Refinement: Refine problem difficulty based on performance data.
  • Cross‑Functional Collaboration: Work with AI researchers and domain experts to integrate benchmarks into evaluation pipelines.

Skills

Engineering design
LLM evaluation concepts
Python scripting
Simulation interpretation

Education

Master’s degree or PhD in Electrical, Mechanical, Aerospace

Tools

ngspice
PySpice
OpenFOAM
FEniCSx
CalculiX
python-control
CadQuery
build123d
OpenModelica
Cantera
Gmsh

Job description

Careerflow Human Data Labs partners with AI companies to bring real-world professional expertise into their products.

Role Overview We are looking for experienced AI Evaluation Engineers (Engineering Simulation & Design) to author and validate "model-breaking," simulation-based engineering design problems to train and evaluate state-of-the-art AI agents. Operating across major engineering disciplines—including Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics—you will create complex, multi-constraint tasks where AI agents must interpret requirements, navigate trade-offs, configure open-source simulation tools, diagnose failures, and iterate toward valid solutions. You will analyze agent execution logs, expose systemic reasoning gaps, and build automated, objective graders to elevate frontier model performance.

Job Requirements:
  • Education & Expertise: Master’s degree or PhD in Electrical, Mechanical, Aerospace, with 10+ years of hands-on engineering design experience.

  • Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package (e.g., ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, Gmsh) combined with strong Python scripting skills.

  • AI Evaluation & Failure Diagnostics: Hands‑on experience with modern LLMs/coding agents and evaluation concepts (pass@k, failure‑mode analysis, nondeterministic behavior), with the ability to audit trajectory logs and isolate core reasoning/tool‑use failures.

  • Domain Rigor & Precision: Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.

  • Availability & Commitment: Talent must have weekend on‑call availability (part‑time engagement is acceptable).

  • Technical Infrastructure: Personal desktop/laptop equipped with a stable, high‑speed internet connection in a remote setup.

Job Responsibilities:
  • Model‑Breaking Problem Design: Author original, self‑contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders.

  • Environment & Simulation Integration: Build, run, and validate problem environments using open‑source simulation tools and custom Python test benches.

  • Trajectory Analysis & Failure Mode Taxonomy: Evaluate coding agent outputs and execution logs across repeated trials to identify systemic failure modes (e.g., misinterpreting simulator feedback, premature design convergence, physically impossible geometries).

  • Difficulty Calibration & Benchmark Refinement: Iteratively refine problem difficulty based on empirical model performance data without introducing ambiguity or missing information.

  • Cross‑Functional Collaboration: Partner with AI researchers, pod leads, and domain experts to integrate high‑rigor benchmarks into the model evaluation pipeline.

  • Location: Bangladesh, India, Indonesia, Egypt, Ghana, Kenya, Nigeria, Turkey, Brazil, Colombia.

Domains
  • Electrical Engineering
  • Mechanical Engineering
  • Aerospace Engineering
Education & Experience
  • Bachelor's degree or equivalent practical experience in any field.
  • Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required.
Offer Details
  • Commitments Required: 30-40 hours per week with 4 hours of overlap with PST.
  • Engagement type: Contractor
  • Engagement Length: upto 24 weeks
Evaluation Process
  • Shortlisted candidates will be sent a Job Interest Form.
  • Finalized talents will go through delivery review & proceed further accordingly.
Engagement Length
  • Up to 24 weeks, Full‑time (8 hours/day) 40 hours per week. Overlap 4 hours with PST (Weekend on‑call availability required) Rate Range: $500 per task (considering 30-40 hours of AHT per task)

This is a Pay-Per-Task model which means you'll get paid per task for an estimated hourly rate.

--

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Expert
Engineering Expert

Turing Global India • United States

Remote
USD 124,000 - 193,000
LLM Engineering Expert (Freelancing)
LLM Engineering Expert (Freelancing)

Zettamine Labs • United States

Remote
USD 165,000 - 276,000
Global AI Evaluation Engineer - Simulation & Design
Global AI Evaluation Engineer - Simulation & Design

Dover • United States

Remote
USD 1,200 - 2,500
AI Evaluation Engineer: Simulation & Design (Contract)
AI Evaluation Engineer: Simulation & Design (Contract)

Turing Global India • United States

Remote
USD 124,000 - 193,000
AI Software Engineer – LLM Evaluation & Automation (Remote)
AI Software Engineer – LLM Evaluation & Automation (Remote)

Stage 4 Solutions Inc • United States

Remote
USD 99,000 - 108,000
Health benefits
401K
LLM Red Team Specialist - Failure Modes & Edge Cases
LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

On-site
USD 68,880 - 206,640
Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
LLM Evaluation & Engineering Simulation Lead (Freelance)
LLM Evaluation & Engineering Simulation Lead (Freelance)

Zettamine Labs • United States

Remote
USD 165,000 - 276,000
ML Engineer - Remote
ML Engineer - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 83,000