Job Title: LLM Engineering Expert AI Evaluation / Engineering Simulation (Freelancing)
Job Type: Freelance
Contract Duration: Up to 24 weeks
Experience: 10+ years
Work Mode: Remote
Availability: 40 hours/week with 4 hours overlap with PST
Joining: Immediate
Job Overview
We are looking for experienced Engineering Experts with strong LLM/AI evaluation and engineering simulation experience to work on advanced AI evaluation and benchmarking projects.
The role involves creating and validating complex engineering design problems for AI agents, working with open-source simulation tools, analyzing AI-generated solutions and execution logs, and identifying reasoning, coding, and tool-use failures.
Candidates should have a strong engineering background combined with hands-on experience in Python, engineering simulation, and modern LLM/coding-agent evaluation.
Key Responsibilities
- Create complex, self-contained engineering design and simulation problems for AI evaluation.
- Define technical constraints, optimization objectives, reference solutions, and objective evaluation criteria.
- Build and validate engineering simulation environments using open-source simulation tools.
- Develop Python-based scripts, test benches, and automated graders.
- Analyze AI agent outputs, coding trajectories, and execution logs.
- Identify failure modes such as incorrect simulator interpretation, premature convergence, invalid designs, and tool-use errors.
- Refine task difficulty based on model performance and evaluation results.
- Ensure physical plausibility, unit consistency, boundary conditions, convergence, and technical accuracy.
- Collaborate with AI researchers, domain experts, and engineering teams to improve AI evaluation benchmarks.
Mandatory Skills
- 10+ years of hands-on engineering design experience.
- Master's degree or PhD in Electrical, Mechanical, Aerospace Engineering, or a closely related engineering discipline.
- Strong Python programming/scripting skills.
- Hands-on experience with at least one engineering simulation/tooling platform such as:
- ngspice
- PySpice
- OpenFOAM
- FEniCSx
- CalculiX
- python-control
- CadQuery
- build123d
- OpenModelica
- Cantera
- Gmsh
- Experience working with LLMs, AI coding agents, or AI evaluation.
- Understanding of AI evaluation concepts such as pass@k, failure-mode analysis, and nondeterministic behavior.
- Ability to analyze AI agent trajectories/logs and identify reasoning or tool-use failures.
- Strong understanding of engineering fundamentals, physical plausibility, units, boundary conditions, and convergence criteria.
Preferred Background :
- Masters degree or PhD in Electrical Engineering, Mechanical Engineering, Aerospace Engineering, or a closely related engineering discipline.
- Strong academic or professional background in engineering design, simulation, computational engineering, or related technical domains.
- Candidates with relevant practical engineering expertise and equivalent experience may also be considered, subject to the role requirements.
Important Note
This is a short-term contract/freelance opportunity focused on engineering simulation and AI evaluation rather than a conventional permanent engineering position.