Freelancing | LLM - Engineering Expert | Remote | Immediate Joiner

PGC Digital (America) Inc: CMMI Level 3 Company

United States

Remote

USD 138,000 - 248,000

Full time

37 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

PGC Digital (America) Inc: CMMI Level 3 Company seeks experienced AI Evaluation Engineers to author and validate model-breaking, simulation-based design problems across Electrical and Mechanical domains. You will create complex tasks, assess AI agents, and build objective graders to push frontier model performance.

Requirements include a Master’s or PhD in Electrical/Mechanical/Aerospace with 10+ years hands-on design experience, strong Python and Open-Source simulation tooling, and weekend

Qualifications

  • Master’s degree or PhD in Electrical, Mechanical, or Aerospace engineering with 10+ years hands-on engineering design experience.
  • Proficiency with at least one domain-relevant open-source simulation package and strong Python scripting skills.
  • Hands-on experience with modern LLMs/coding agents and evaluation concepts, with ability to audit trajectory logs.
  • Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
  • Talent must have weekend on-call availability (part-time engagement is acceptable).
  • Personal desktop/laptop with stable high-speed internet in a remote setup.

Responsibilities

  • Design original, self-contained engineering tasks with competing constraints and validated reference solutions.
  • Build, run, and validate environments using open-source simulation tools and Python test benches.
  • Evaluate agent outputs and logs to identify systemic failure modes.
  • Iteratively calibrate task difficulty based on empirical model performance data.
  • Collaborate with AI researchers, pod leads, and domain experts to integrate high-rigor benchmarks into the evaluation pipeline.

Skills

Python scripting
AI evaluation concepts
Failure diagnostics
Simulation problem design
Domain rigor

Education

Master’s degree or PhD in Electrical/Mechanical/Aerospace engineering

Tools

ngspice
PySpice
OpenFOAM
FEniCSx
CalculiX
python-control
CadQuery
build123d
OpenModelica
Cantera
Gmsh

Job description

We are seeking experienced AI Evaluation Engineers (Engineering Simulation & Design) to author and validate "model-breaking," simulation-based engineering design problems to train and evaluate state-of-the-art AI agents. Operating across major engineering disciplines—including Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics—you will create complex, multi-constraint tasks where AI agents must interpret requirements, navigate trade-offs, configure open-source simulation tools, diagnose failures, and iterate toward valid solutions. You will analyze agent execution logs, expose systemic reasoning gaps, and build automated, objective graders to elevate frontier model performance.

Job Requirements:
  • Education & Expertise: Master’s degree or PhD in Electrical, Mechanical, Aerospace, with 10+ years of hands-on engineering design experience.
  • Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package (e.g., ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, Gmsh) combined with strong Python scripting skills.
  • AI Evaluation & Failure Diagnostics: Hands-on experience with modern LLMs/coding agents and evaluation concepts (pass@k, failure-mode analysis, nondeterministic behavior), with the ability to audit trajectory logs and isolate core reasoning/tool-use failures.
  • Domain Rigor & Precision: Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
  • Availability & Commitment: Talent must have weekend on-call availability (part-time engagement is acceptable).
  • Technical Infrastructure: Personal desktop/laptop equipped with a stable, high-speed internet connection in a remote setup.
Job Responsibilities:
  • Model-Breaking Problem Design: Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders.
  • Environment & Simulation Integration: Build, run, and validate problem environments using open-source simulation tools and custom Python test benches.
  • Trajectory Analysis & Failure Mode Taxonomy: Evaluate coding agent outputs and execution logs across repeated trials to identify systemic failure modes (e.g., misinterpreting simulator feedback, premature design convergence, physically impossible geometries).
  • Difficulty Calibration & Benchmark Refinement: Iteratively refine problem difficulty based on empirical model performance data without introducing ambiguity or missing information.
  • Cross-Functional Collaboration: Partner with AI researchers, pod leads, and domain experts to integrate high-rigor benchmarks into the model evaluation pipeline.
Domains:
  • Electrical Engineering
  • Mechanical Engineering
Education & Experience
  • Bachelor's degree or equivalent practical experience in any field.
  • Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required.
  • Commitments Required: 40 hours per week with 4 hours of overlap with PST.
  • Engagement type: Contractor
  • Engagement Length: upto 24 weeks
Evaluation Process -
  • Shortlisted candidates will be sent a Job Interest Form.
  • Finalized talents will go through delivery review & proceed further accordingly.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Engineering Expert- Simulation & Design
LLM Engineering Expert- Simulation & Design

Dover • United States

Remote
USD 1,200 - 2,500
Engineering Expert
Engineering Expert

Turing Global India • United States

Remote
USD 124,000 - 193,000
LLM Engineering Expert (Freelancing)
LLM Engineering Expert (Freelancing)

Zettamine Labs • United States

Remote
USD 165,000 - 276,000
AI Evaluation Engineer · Turing Turing · TBD · remote · 3w ago TBD 3w ago
AI Evaluation Engineer · Turing Turing · TBD · remote · 3w ago TBD 3w ago

Benture • San Francisco (CA)

Remote
USD 83,000 - 138,000
AI Evaluation Engineer
AI Evaluation Engineer

Crossing Hurdles • United States

Remote
USD 120,000 - 180,000
Senior Software Engineer – LLM Evaluation
Senior Software Engineer – LLM Evaluation

Jobgether SRL • United States

Remote
USD 55,000 - 96,000
Fully remote
Part-time
Independent contractor
+4
LLM Red Team Specialist - Failure Modes & Edge Cases
LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday AI • United States

Remote
USD 171,924,000 - 257,887,000
Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based
Senior Software Engineer – LLM Evaluation (Fully Remote)
Senior Software Engineer – LLM Evaluation (Fully Remote)

Partner Company • United States

Remote
USD 83,000 - 165,000
Remote-friendly
Flexible hours
Contractor-friendly
Remote | Senior Software Engineer – LLM Evaluation
Remote | Senior Software Engineer – LLM Evaluation

24-Mag Llc • New York (NY)

Remote
USD 14,000 - 55,000