Evaluation Engineer (Consulting / Contract-to-Hire) Seattle/Remote

AI Ethics Network

Northern (KY)

Hybrid

USD 138,000 - 220,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

AI Ethics Network is seeking evaluators to perform high-stakes testing of robotics, teen-facing products, and post-incident scenarios. You will red-team models and agents, run prompt-injection tests, assess sensor and navigation safety, and document findings for executives.

The role requires a Master’s degree (PhD preferred in your track), clear written English, and the ability to defend methods in front of counsel. U.S.

Qualifications

  • Master’s required; PhD preferred in related track.
  • Evidence you personally ran the work; not summarized papers.
  • Authorized to work as a U.S. contractor (no sponsorship).
  • Able to refuse scope outside your qualification.

Responsibilities

  • Red-team models and agents including jailbreaks and prompt injection.
  • Evaluate embodied and robotic systems for safety and risk.
  • Audit companion and youth products for age-gate integrity and data minimization.
  • Assess accessibility and neurodivergent passes in interfaces.
  • Produce board-ready deliverables: severity-ranked findings and executive letters.
  • Join founder calls for scoping and price-setting; comply with NDA and data protection.

Skills

Clear written English

Education

Master’s degree
PhD preferred

Job description

$100–$160 per hour on contract. Rate follows domain depth. Robotics, teen-facing, or post-incident work sits at the top of the band.

Typical assessment: 20–60 hours. If conversion is offered after paid work: $145,000–$185,000 base equivalent plus project bonuses. No equity at the consulting stage.

Why this role exists

Eval Labs does not let the training team grade itself. When a client is weeks from launch — or managing an active incident — the evaluator must hold an advanced degree in a field related to the system under test and be able to defend methods under questioning from counsel, leadership, or insurers. We hire for that standard.

Core responsibilities
  • Red-team models and agents. Jailbreaks, prompt injection, tool-use exploits, multi-step scope creep, over-refusal, and deception.
  • Evaluate embodied and robotic systems. Sensor spoofing, stop-condition failure, actuator boundaries, navigation safety, and physical-action risk.
  • Audit companion and youth products. Parasocial risk, age-gate integrity, unknown-age defaults, and minor-data minimization.
  • Accessibility and neurodivergent passes. Instruction load, caption fidelity, disclosure response, and sensory/interface strain.
  • Board-ready deliverables. Severity-ranked transcripts, credential and privacy hygiene notes, and executive findings letters for non-researcher readers.
  • Scoping. Join founder calls after intake to set boundaries and price the engagement.
  • Data protection. Strict NDA. No reuse of client prompts, logs, or exploit chains.

Skilled in at least one track. Two preferred. Three is priced accordingly.

Track A
Models & agents

Degree: MS or PhD in CS, machine learning, cybersecurity, or equivalent.

LLM/agent evaluation, log analysis, eval harnesses, and failure modes explained without theater.

Track B
Embodied / robotics

Degree: MS or PhD in robotics, mechanical or electrical engineering, controls, or computer vision.

Sensors, actuators, stop conditions, and real-world safety-interlock failure — not only simulators.

Track C
Vulnerable users & accessibility

Degree: MS or PhD in HCI, cognitive science, special education, clinical/counseling with HCI crossover, or documented ND accessibility practice.

Paid reviewer work, instruction-load evaluation, and a method you can defend.

Qualifications

Required. Master’s required; PhD preferred in a field related to your primary track. Evidence you personally ran the work — not summarized papers. Clear written English. Authorized to work as a U.S. contractor (no sponsorship). Able to refuse scope outside your qualification. No active conflict with an Eval Labs prospect or a competing audit of the same client system.

Strongly preferred. Expert-witness writing, incident reports, or board/insurer packets. Public red-team write-ups, eval suites, or robotics safety publications. Companion, teen-facing, or voice-agent experience. Pacific Time hours for client calls.

Hiring process
  • Written application. CV, primary track (A, B, or C), and a short list of bullet points covering the most impressive work you have actually done — systems tested, failures found, papers, products, incidents, harnesses, robots, audits. Concrete. Yours.
  • Interviews. A short series of conversations with the Eval Labs team, including a technical / whiteboard session on a failure mode in your track.
  • Paid working sample. 3–5 hours at the posted rate on a sanitized fixture. You produce a short findings memo with severity ranks.
  • Bench. NDA, then live client assessments.

We do not hire on enthusiasm for “AI ethics” alone. We hire people who can defend a method when a lawyer is in the room.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer – Evals
Research Engineer – Evals

Rnb Consultancy • San Francisco (CA)

On-site
USD 200,000 - 250,000
Equity
Relocation support
Visa/immigration assistance
AI Safety and Red-Team Evaluator
AI Safety and Red-Team Evaluator

AI Trainer Jobs • United States

Remote
USD 83,000 - 131,000
Software Engineering Evaluation Specialist
Software Engineering Evaluation Specialist

AI Trainer Jobs • United States

Remote
USD 158,000 - 200,000
Engineering / manufacturing / technical operations Evaluator
Engineering / manufacturing / technical operations Evaluator

AI Trainer Jobs • United States

Remote
USD 110,000 - 165,000
Remote work
Flexible hours
Control Systems Engineer — AI Evaluation for Automation & Contro
Control Systems Engineer — AI Evaluation for Automation & Contro

AI Trainer Jobs • United States

Remote
USD 55,000 - 110,000
Autonomous Systems Safety Evaluator
Autonomous Systems Safety Evaluator

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Frontend Engineering AI Evaluator
Frontend Engineering AI Evaluator

AI Trainer Jobs • United States

Remote
USD 55,000 - 96,000
Robotics Evaluation Specialist
Robotics Evaluation Specialist

AI Trainer Jobs • United States

Remote
USD 55,000 - 83,000
Evals Engineer, Offensive Cyber
Evals Engineer, Offensive Cyber

Zealot Labs • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineering AI Evaluator
Site Reliability Engineering AI Evaluator

AI Trainer Jobs • United States

Remote
USD 83,000 - 165,000
Remote work
Contractor position
Flexible hours