Remote AI Code Engineer — Benchmark & Validate Models

Turing

United States

Remote

USD 69,000 - 138,000

Part time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Turing, based in San Francisco, is seeking a Software Engineering evaluator to create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers.

You will curate code examples, provide precise solutions, and refine AI-generated code across Python, JavaScript (including ReactJS), C/C++, Java, Rust, and Go; evaluating and refining AI-generated code for efficiency, scalability, and reliability.

Qualifications

  • 3+ years of software engineering experience.
  • Strong expertise in building full-stack applications and deploying production-grade software.
  • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment.
  • Excellent oral and written communication skills for clear, structured evaluation rationales.

Responsibilities

  • Curate code examples for model training and benchmarking.
  • Provide precise solutions and corrections in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go.
  • Evaluate and refine AI-generated code for efficiency and scalability.
  • Collaborate with researchers and cross-functional teams on AI-driven coding solutions.
  • Build agents to verify code quality and identify error patterns.
  • Hypothesize steps in the software engineering cycle and test model capabilities.

Skills

Software engineering
Full-stack development
Architecture & design
Code quality & reviews
Communication skills

Tools

Python
JavaScript/ReactJS
C/C++
Java
Go
Rust

Job description

Turing, based in San Francisco, is seeking a Software Engineering evaluator to create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers.

You will curate code examples, provide precise solutions, and refine AI-generated code across Python, JavaScript (including ReactJS), C/C++, Java, Rust, and Go; evaluating and refining AI-generated code for efficiency, scalability, and reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Software Engineer - Code & Data Evaluation
Remote AI Software Engineer - Code & Data Evaluation

Turing • United States

Remote
USD 83,000 - 138,000
Remote AI Software Engineer: Code Quality & Evaluation
Remote AI Software Engineer: Code Quality & Evaluation

Turing • United States

Remote
USD 55,000 - 110,000
Remote C/C++ AI Software Engineer (Contractor)
Remote C/C++ AI Software Engineer (Contractor)

Turing • United States

Remote
USD 55,000 - 96,000
Flexible hours
1 month contract
Remote AI Software Engineer - Data & Model Training
Remote AI Software Engineer - Data & Model Training

turing • San Francisco (CA)

Remote
USD 83,000 - 124,000
Remote AI Software Engineer – ML Data & Code Evaluation
Remote AI Software Engineer – ML Data & Code Evaluation

engineeringjobs.net, Inc. • Pittsburgh

Remote
USD 90,000 - 130,000
Java AI Code Evaluator (Contract, 10–40 hrs/wk)
Java AI Code Evaluator (Contract, 10–40 hrs/wk)

Turing • New York (NY)

On-site
USD 83,000 - 165,000
Remote AI Code Engineer (Rust & Full-Stack)
Remote AI Code Engineer (Rust & Full-Stack)

Turing • Boston (MA)

Remote
USD 120,000 - 180,000
Remote ML Infrastructure Engineer for LLM Benchmarking
Remote ML Infrastructure Engineer for LLM Benchmarking

engineeringjobs.net, Inc. • Pittsburgh

Remote
USD 120,000 - 160,000
Java Software Engineer — AI Data & Code Evaluation
Java Software Engineer — AI Data & Code Evaluation

Turing • Boston (MA)

Remote
USD 120,000 - 170,000
Senior AI Code Evaluator (Python/TypeScript)
Senior AI Code Evaluator (Python/TypeScript)

Turing • United States

Remote
USD 83,000 - 152,000