Software Engineering Expert

Weekday 1

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Fully remote
Weekly payments

Job summary

Weekday 1 is seeking experienced Software Engineers for a fully remote, contractor engagement to build multi-step software engineering benchmarks for frontier AI models. You will implement, debug, configure environments, and reason about tasks alongside AI researchers.

The role emphasizes Python reference solutions, reproducibility, and rigorous validation checks, with approximately 35 hours per week. Weekly payments apply for this remote engagement.

Qualifications

  • Master's degree, PhD, or equivalent practical experience in CS/Software Engineering or STEM
  • Minimum 1 year of professional software engineering or research engineering experience
  • Strong Python proficiency including design, scripting, debugging, testing, and clean code
  • Solid experience with Git, IDEs, and modern software development workflows
  • Good understanding of software architecture, debugging methods, and best practices
  • Experience with AI coding assistants, prompt engineering, or AI agent workflows is preferred

Responsibilities

  • Design realistic, multi-step software engineering benchmarks reflecting typical development workflows
  • Develop comprehensive Python reference solutions that are reproducible and validated
  • Use AI coding assistants as part of the workflow while evaluating their strengths and failure modes
  • Review benchmark tasks for technical correctness, clarity, completeness, and difficulty
  • Analyze AI-generated solutions to identify errors, debugging challenges, and reasoning gaps
  • Collaborate with AI researchers and engineering teams to refine evaluation methodologies

Skills

Python
Git
Debugging
Software design
Code testing

Education

Master's degree / PhD in CS or STEM

Tools

IDEs
AI coding assistants

Job description

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Software Engineers to design and validate sophisticated engineering tasks that mirror real-world software development challenges.

In this role, you will create complex, multi-step software engineering benchmarks that require implementation, debugging, environment configuration, and technical reasoning. Working alongside AI researchers, you'll help identify where advanced AI coding systems succeed, where they fail, and how benchmark quality can be continuously improved.

This is a fully remote, full-time engagement requiring approximately 35 hours per week

Requirements

Key Responsibilities
  • Design realistic, multi-step software engineering challenges based on practical development workflows, including implementation, debugging, system configuration, testing, and documentation.
  • Develop comprehensive reference solutions in Python, ensuring each task is reproducible, verifiable, and supported by appropriate validation checks.
  • Leverage AI coding assistants as part of your development workflow while evaluating their strengths, limitations, and failure modes on complex engineering tasks.
  • Review benchmark tasks created by fellow experts, providing detailed feedback on technical correctness, clarity, completeness, and level of difficulty.
  • Analyze AI-generated solutions to identify implementation errors, debugging challenges, reasoning gaps, and opportunities for improving benchmark quality.
  • Collaborate closely with AI researchers and engineering teams to refine evaluation methodologies and maintain consistently high technical standards.
Required Qualifications
  • Master's degree, PhD, or equivalent practical experience in Computer Science, Software Engineering, or another STEM discipline involving substantial programming and technical problem solving.
  • Minimum 1 year of professional experience in software engineering, research engineering, applied research, or another technically intensive engineering role.
  • Strong hands‑on proficiency in Python, including software design, scripting, debugging, testing, and writing clean, maintainable code.
  • Solid experience using Git, integrated development environments (IDEs), and modern software development workflows.
  • Strong understanding of software architecture, debugging methodologies, and engineering best practices.
  • Experience using AI coding assistants, prompt engineering techniques, or AI agent workflows is preferred.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently.
  • Strong written communication skills for documenting technical solutions and engineering decisions.
  • Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
  • Experience developing complex software systems across backend, infrastructure, automation, or developer tooling.
  • Familiarity with large language models, AI‑assisted software development, or autonomous coding agents.
  • Background in benchmark design, software quality engineering, or technical evaluation frameworks.
  • Contributions to open‑source projects, technical publications, or advanced engineering initiatives.
Why Join
  • Help shape how next‑generation AI coding systems are evaluated against real‑world engineering challenges.
  • Collaborate with leading AI researchers developing frontier evaluation benchmarks.
  • Apply your software engineering expertise to improve AI reliability, technical reasoning, and coding performance.
  • Contribute directly to benchmark development that influences the evolution of advanced AI systems.
  • Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives.
Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details
  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

STEM Researcher - Computational Fields
STEM Researcher - Computational Fields

Weekday 1 • United States

Remote
USD 83,000 - 124,000
LLM Red Team Specialist - Failure Modes & Edge Cases
LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour

24-MAG • United States

Remote
USD 69,000 - 138,000
Remote work
Flexible hours
Contractor engagement
Software Engineering Expert (AI Training)
Software Engineering Expert (AI Training)

Weekday 1 • United States

Remote
USD 138,000 - 317,000
Remote | Senior Software Engineer — $50–$100/hour
Remote | Senior Software Engineer — $50–$100/hour

engineeringjobs.net, Inc. • New York (NY)

Remote
USD 69,000 - 138,000
Remote Software Engineer for AI-Driven Systems
Remote Software Engineer for AI-Driven Systems

YO IT Consulting • United States

On-site
USD 80,000 - 120,000
Technology AI Evaluation Expert
Technology AI Evaluation Expert

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
Remote | Senior Software Engineer — $50–$100/hour
Remote | Senior Software Engineer — $50–$100/hour

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 69,000 - 138,000
Senior Full-Stack Software Engineer
Senior Full-Stack Software Engineer

Appsierra Group • American Samoa

On-site
USD 104,000 - 208,000