Site Reliability Engineering AI Evaluator

AI Trainer Jobs

United States

Remote

USD 83,000 - 165,000

Part time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote work
Contractor position
Flexible hours

Job summary

AI Trainer Jobs is seeking a remote Site Reliability Engineering AI Evaluator to analyze production code, debug traces, and model outputs against correctness standards. You will reproduce failures, write unit tests, and explain fixes so the modeling team can target gaps.

The role emphasizes evaluating generated code for compilation, tests, and edge cases, working with experienced engineers, and judging outputs with a rubric. This is a contractor position with flexible hours.

Qualifications

  • Strong day-job engineering experience.
  • Ability to read, run, and debug unfamiliar code.
  • Experience writing concise unit tests that capture a single failure mode.
  • Clear written reasoning to convince a senior engineer.
  • Reliable async availability for at least 10 hours per week.
  • Prior code-review or technical-interview-grading experience is a plus.

Responsibilities

  • Run and reproduce candidate code outputs in a sandboxed environment.
  • Grade cloud & platform engineering solutions for correctness, style, and edge-case handling.
  • Write minimal failing tests that demonstrate the bug a model output missed.
  • Compare paired solutions and rank them with a written rationale tied to the rubric.

Skills

Engineering experience
Unit tests
Concise written reasoning
Async availability
Code review

Tools

pytest
Jest
JUnit

Job description

About the role

Site Reliability Engineering AI Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written, and explain the fix so the modeling team can target the gap.
Engineering model quality lives or dies on whether the generated code actually compiles, passes tests, and handles edge cases. AuraOne pairs experienced engineers with the modeling team to grade outputs the way a code reviewer would.
Judge generated code and software engineering agents. Read their debugging traces.

Responsibilities
  • Run and reproduce candidate code outputs in a sandboxed environment for Site Reliability Engineering AI Evaluator assignments.
  • Grade cloud & platform engineering solutions for correctness, style, and edge-case handling.
  • Write minimal failing tests that demonstrate the bug a model output missed.
  • Compare paired solutions and rank them with a written rationale tied to the rubric.

Role details

Track Code review & evaluation Work model Remote · Independent specialist contractor Compensation Hourly rate confirmed after the interview process. Eligible from US

What you should bring

  • Strong day-job engineering experience — you can read, run, and debug unfamiliar code for Site Reliability Engineering AI Evaluator work.
  • Comfort writing concise unit tests that capture a single failure mode.
  • Familiarity with a testing framework. pytest, Jest, or JUnit. Go test, RSpec, or whatever you use.
  • Clear written reasoning — your review note has to convince another senior engineer.
  • Reliable async availability for at least 10 hours per week.
  • Prior code-review or technical-interview-grading experience is a plus.

Role signals

Example tasks

  • Reproduce a generated engineering solution to a coding task, run the test suite, and grade it.
  • Write the smallest failing test that demonstrates a model's edge-case bug.
  • Compare two paired solutions and rank them with a written rationale tied to the rubric.
  • Triage a security issue surfaced by a model and document the patch the model should have produced.

Useful experience

  • Open-source contributions or a public portfolio that demonstrates production-quality code.
  • Experience with the target language's standard tooling, linters, and idiomatic style guides.
  • Familiarity with security-review checklists (OWASP, CWE) and AppSec patterns.

Compensation and schedule

Hourly rate confirmed after the interview process.
Expected arrangement: contractor , with program-defined task volume and review pacing. Placement depends on current program demand and reviewer confirmation.

Skills used in matching

  • Code review
  • Debugging
  • Unit testing
  • Software engineering judgment
  • Cloud & platform engineering
  • Software engineering and computer use
  • AI evaluation
  • Rubric writing
  • Expert review
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Incident management / reliability / SRE Evaluator
Incident management / reliability / SRE Evaluator

AI Trainer Jobs • United States

Remote
USD 110,000 - 165,000
Site Reliability Engineer
Site Reliability Engineer

AI Trainer Jobs • United States

Remote
USD 83,000 - 179,000
Senior Software Engineer — AI Coding Evaluator
Senior Software Engineer — AI Coding Evaluator

AI Trainer Jobs • United States

Remote
USD 83,000 - 124,000
Frontend Engineering AI Evaluator
Frontend Engineering AI Evaluator

AI Trainer Jobs • United States

Remote
USD 55,000 - 96,000
Software Engineer, Data Platforms
Software Engineer, Data Platforms

AI Trainer Jobs • United States

Remote
USD 165,000 - 193,000
Software Engineering Evaluation Specialist
Software Engineering Evaluation Specialist

AI Trainer Jobs • United States

Remote
USD 158,000 - 200,000
AI Software Engineering Domain Expert
AI Software Engineering Domain Expert

AI Trainer Jobs • United States

Remote
USD 55,000 - 83,000
DevOps and Cloud Infrastructure AI Expert
DevOps and Cloud Infrastructure AI Expert

AI Trainer Jobs • United States

Remote
USD 55,000 - 110,000
Reliability Engineer
Reliability Engineer

AI Trainer Jobs • United States

Remote
USD 110,000 - 124,000
Fullstack Engineer-GEN AI
Fullstack Engineer-GEN AI

AI Trainer Jobs • United States

Remote
USD 165,000 - 193,000