AI Software Engineering Evaluator — Systems & Code Quality

turing

United States

Remote

USD 124,000 - 179,000

Part time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Turing, based in San Francisco, is seeking a Software Engineering evaluator to help create datasets for training and benchmarking large language models. You will curate code examples across Python, C/C++, Rust, Go, Java, and JavaScript (ReactJS) and assess AI-generated code for accuracy, efficiency, and reliability.

You'll collaborate with researchers and engineers to improve enterprise-grade AI tooling, with flexible hours (10–40 hrs/week) for a 1-month engagement.

Qualifications

  • Several years of software engineering experience (3 years or more).
  • Strong expertise in systems programming, infrastructure, or backend development using languages like Python, C/C++, Rust, and Go.
  • Experience building and deploying scalable, production-grade software using modern languages and tools.
  • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment.
  • Excellent oral and written communication skills for clear, structured evaluation rationales.

Responsibilities

  • Curate code examples and build solutions in Python, C/C++, Rust, Go, Java, and JavaScript (including ReactJS).
  • Evaluate and refine AI-generated code with an emphasis on systems-level correctness, performance, and reliability.
  • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks.
  • Build agents that can verify the quality of systems-level and infrastructure code and identify error patterns.
  • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them.
  • Design verification mechanisms that can automatically verify a solution to a software engineering task.

Skills

Software engineering
Systems programming
Backend development
Python
C/C++
Rust
Go
Code review

Job description

Turing, based in San Francisco, is seeking a Software Engineering evaluator to help create datasets for training and benchmarking large language models. You will curate code examples across Python, C/C++, Rust, Go, Java, and JavaScript (ReactJS) and assess AI-generated code for accuracy, efficiency, and reliability.

You'll collaborate with researchers and engineers to improve enterprise-grade AI tooling, with flexible hours (10–40 hrs/week) for a 1-month engagement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Code Evaluator & Dataset Engineer
AI Code Evaluator & Dataset Engineer

turing • San Francisco (CA)

Remote
USD 83,000 - 165,000
Remote AI Software Engineer: Code Quality & Evaluation
Remote AI Software Engineer: Code Quality & Evaluation

Turing • United States

Remote
USD 55,000 - 110,000
Java AI Code Evaluator (Contract, 10–40 hrs/wk)
Java AI Code Evaluator (Contract, 10–40 hrs/wk)

Turing • New York (NY)

On-site
USD 83,000 - 165,000
Remote AI Code Evaluator (Contractor)
Remote AI Code Evaluator (Contractor)

turing • San Francisco (CA)

Remote
USD 69,000 - 138,000
Remote AI Software Engineer - Code & Data Evaluation
Remote AI Software Engineer - Code & Data Evaluation

Turing • United States

Remote
USD 83,000 - 138,000
Remote C/C++ AI Software Engineer (Contractor)
Remote C/C++ AI Software Engineer (Contractor)

Turing • United States

Remote
USD 55,000 - 96,000
Flexible hours
1 month contract
Contractor: AI Code Evaluation Engineer
Contractor: AI Code Evaluation Engineer

turing • San Francisco (CA)

Remote
USD 55,000 - 110,000
C/C++ Software Engineer for AI Data & Benchmarks
C/C++ Software Engineer for AI Data & Benchmarks

turing • San Francisco (CA)

Remote
USD 83,000 - 138,000
Remote AI Code Engineer — Benchmark & Validate Models
Remote AI Code Engineer — Benchmark & Validate Models

Turing • United States

Remote
USD 69,000 - 138,000
Senior Software Engineer (Contract) — AI Coding & Evaluation
Senior Software Engineer (Contract) — AI Coding & Evaluation

turing • United States

Remote
USD 83,000 - 152,000