Remote Python Engineer (AI Benchmarking & Task Design)

24-MAG

New York (NY)

Remote

USD 76,000 - 117,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote

Job summary

24-MAG LLC is offering a fully remote, full-time consulting opportunity for experienced software engineers with strong Python skills. You will design realistic, multi-step coding tasks and build verifiable Python reference solutions, including environment setup and tests.

The role evaluates AI coding assistants and requires rigorous documentation and peer-review contributions. The ideal candidate has 1+ year in software or research engineering, strong Python, and fluent Git/IDE workflows.

Qualifications

  • At least 1 year of software engineering or research engineering experience.
  • Strong hands-on Python scripting, implementation, and debugging skills.
  • Fluent with Git, IDEs, repositories, and standard software development workflows.
  • Ability to configure environments, dependencies, tooling, and validation processes.

Responsibilities

  • Create realistic, multi-step software engineering challenges based on practical development workflows.
  • Design technically demanding problems requiring implementation, debugging, environment configuration, and analytical reasoning.
  • Define clear requirements, constraints, expected outputs, and acceptance criteria.
  • Ensure tasks assess genuine software engineering capability rather than superficial code generation.
  • Build complete and verifiable reference solutions in Python with setup, tests, and validation checks.
  • Document implementation decisions and expected behavior clearly.
  • Evaluate how frontier models approach complex coding and debugging tasks using AI assistants.
  • Identify errors, unsupported assumptions, inefficient approaches, and incomplete solutions.
  • Document failure patterns and provide evidence for evaluation conclusions.
  • Review tasks and reference solutions created by others; provide actionable feedback.

Job description

24-MAG LLC is offering a fully remote, full-time consulting opportunity for experienced software engineers with strong Python skills. You will design realistic, multi-step coding tasks and build verifiable Python reference solutions, including environment setup and tests.

The role evaluates AI coding assistants and requires rigorous documentation and peer-review contributions. The ideal candidate has 1+ year in software or research engineering, strong Python, and fluent Git/IDE workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Python Engineer for AI Benchmarking
Remote Python Engineer for AI Benchmarking

24-Mag Llc • New York (NY)

Remote
Remote AI Benchmark QA Engineer
Remote AI Benchmark QA Engineer

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 117,000
Remote Python Engineer for AI Lab (Part-Time, Proj-Based)
Remote Python Engineer for AI Lab (Part-Time, Proj-Based)

United States Digital Space LLC • Germany (OH)

Remote
USD 62,000 - 110,000
Remote work
Remote Part-Time AI Software Engineering Trainer
Remote Part-Time AI Software Engineering Trainer

United States Digital Space LLC • Germany (OH)

On-site
USD 62,000 - 110,000
Remote Part-Time Senior Engineering Assessment Consultant
Remote Part-Time Senior Engineering Assessment Consultant

24-Mag Llc • New York (NY)

Remote
USD 69,000 - 96,000
Remote Physics AI Benchmark Designer
Remote Physics AI Benchmark Designer

United States Digital Space LLC • Germany (OH)

On-site
USD 45,000 - 65,000
Remote CS Research Expert - Benchmark & ML Systems
Remote CS Research Expert - Benchmark & ML Systems

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 103,000
Remote work
Flexible schedule
Competitive hourly rate
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
Senior Python Engineer — AI Task Designer & Evaluator
Senior Python Engineer — AI Task Designer & Evaluator

Dorado • United States

Remote
GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000