AI Software Engineer – LLM Evaluation & Automation (Remote)

Stage 4 Solutions Inc

United States

Remote

USD 99,000 - 108,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits
401K

Job summary

Stage 4 Solutions Inc. seeks an AI Software Engineer to build evaluation harnesses and automation for AI software use cases. The role focuses on turning real engineering artifacts into repeatable benchmark tasks, with reproducible runs and controlled environments.

The position is remote in the United States, 6 months with potential extension, 40 hours per week, W2 employee benefits including health and 401K. Must align to Pacific Time hours.

Qualifications

  • Strong software engineering background with automation/test/validation experience.
  • Proficient in Python and at least one of Java or JavaScript.
  • Familiar with containers and Git.
  • Experience with APIs, CI/CD, and engineering workflows.
  • Understanding of AI, LLM, or agent evaluation concepts.
  • Able to troubleshoot and analyze results carefully.

Responsibilities

  • Build and integrate evaluation harnesses and automation for software development use cases.
  • Create versioned, repeatable evaluation processes with reproducible environments.
  • Validate evaluation approaches against human judgment for consistency.
  • Support benchmarking across quality, productivity, and efficiency, including cost and latency.
  • Analyze results across repeated runs to improve workflow reliability and automation.
  • Collaborate with engineering and data teams and document methods and findings.

Skills

Python
Java
JavaScript
Git
Docker
APIs
CI/CD
AI evaluation
Troubleshooting
Automation

Tools

Docker
Git

Job description

AI Software Engineer LLM Evaluation & Automation (Remote)

We are looking for an AI Software Engineer for a B2B high-tech company. In this role, you will build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.

This is a 6-month (extensions likely), 40-hour/week, remote role in the US.

This is a W2 role as a Stage 4 Solutions employee. Health benefits and 401K are offered.

Must be able to work Pacific Time (PST) hours.

Responsibilities
  • Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees), so results stay comparable over time.
  • Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
  • Support execution-based benchmarking across quality, productivity, and efficiency measures, including cost and latency.
  • Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
  • Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.
Requirements:
  • Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
  • Proficient in Python, and comfortable in at least one of Java, JavaScript, or a similar language.
  • Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
  • Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
  • Some familiarity with how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.
  • Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
Preferred:
  • Hands-on work with AI-powered coding tools and agentic applications, such as Claude Code, Devin, or OpenCode.
  • Experience designing benchmarks or evaluations for software systems, especially execution based grading that verifies against tests.
  • Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them against human raters.
  • Experience with build-system-aware test selection, such as Bazel or mapping changed files to the tests that cover them.
  • Experience building reproducible test environments and managing versioned evaluation datasets. Comfortable writing up methodology and results for engineering leadership.

Stage 4 Solutions is an equal opportunity employer. We celebrate diversity and are committed to providing employees with an inclusive environment that is free of discrimination and harassment. All employment decisions are based on the job requirements and candidates' qualifications, without regard to race, color, religion/belief, national origin, gender identity, age, disability, marital status, genetic information or other applicable legally protected characteristics.

Compensation: $72/hr. - $78.57/hr. on W2

#LI-SW1

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Software Engineer: LLM Evaluation & Automation
Remote AI Software Engineer: LLM Evaluation & Automation

Stage 4 Solutions Inc • United States

Remote
USD 99,000 - 108,000
Health benefits
401K
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
AI Evaluation Engineer
AI Evaluation Engineer

Wipro Technologies • San Diego (CA)

On-site
USD 60,000 - 149,000
Medical and dental benefits
Disability insurance
Paid time off
AI Engineer
AI Engineer

Eliassen Group • Carson City (NV)

Remote
USD 124,000 - 152,000
Medical benefits
Dental benefits
Vision benefits
+2
Senior Software Engineer, AI Training - US
Senior Software Engineer, AI Training - US

G2i Inc. • United States

On-site
USD 138,000 - 276,000
Remote worldwide
Flexible hours
Software Engineer - AI Eval & Automation
Software Engineer - AI Eval & Automation

ServiceNow • San Diego (CA)

On-site
USD 150,000 - 190,000
Magnit Global benefits
Senior Software Engineer, AI Training - LATAM
Senior Software Engineer, AI Training - LATAM

BlockchainHQ • United States

Remote
USD 138,000 - 276,000
Senior Software Engineer, AI Training
Senior Software Engineer, AI Training

G2i Inc. • United States

Remote
USD 138,000 - 276,000
Senior Software Engineer, AI Training
Senior Software Engineer, AI Training

BlockchainHQ • United States

Remote
USD 138,000 - 276,000
AI Engineer (Automation & LLM Systems)
AI Engineer (Automation & LLM Systems)

Huzzle • United States

Remote
USD 120,000 - 160,000
Fully remote
Competitive salary
Growth opportunities
+3