Software Engineering Expert

Weekday AI

United States

Remote

USD 83,000 - 124,000

Part time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Weekday AI is seeking experienced Software Engineers to design and validate multi-step benchmarks that mirror real-world software tasks for frontier AI models. You will implement complex tasks, configure environments, and reason through solutions alongside AI researchers.

This is a fully remote, part-time engagement requiring about 35 hours per week, with weekly payments and strong emphasis on code quality, reproducibility, and rigorous validation checks.

Qualifications

  • Master's degree, PhD, or equivalent practical experience in Computer Science, Software Engineering, or another STEM discipline involving substantial programming.
  • Minimum 1 year of professional software engineering or research engineering experience.
  • Strong hands-on Python including design, scripting, debugging, testing and clean code.
  • Solid experience with Git, IDEs and modern software development workflows.
  • Experience with AI coding assistants, prompt engineering, or AI agent workflows is preferred.

Responsibilities

  • Design realistic, multi-step software engineering challenges based on practical workflows.
  • Develop reference solutions in Python that are reproducible and validated.
  • Leverage AI coding assistants while evaluating their strengths and failure modes on tasks.
  • Review tasks created by peers, providing feedback on correctness and difficulty.
  • Analyze AI-generated solutions to identify errors and improvement opportunities.
  • Collaborate with AI researchers and engineers to refine evaluation methodologies.

Skills

Python
Git
Software engineering
Debugging
Testing
Code review
Documentation

Education

Master's/PhD or equivalent

Tools

Git
IDEs

Job description

This role is for one of our clients


Compensation: $60-$90 per hour


Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Software Engineers to design and validate sophisticated engineering tasks that mirror real-world software development challenges.


In this role, you will create complex, multi-step software engineering benchmarks that require implementation, debugging, environment configuration, and technical reasoning. Working alongside AI researchers, you'll help identify where advanced AI coding systems succeed, where they fail, and how benchmark quality can be continuously improved.


This is a fully remote, full-time engagement requiring approximately 35 hours per week.


Key Responsibilities


  • Design realistic, multi-step software engineering challenges based on practical development workflows, including implementation, debugging, system configuration, testing, and documentation.

  • Develop comprehensive reference solutions in Python, ensuring each task is reproducible, verifiable, and supported by appropriate validation checks.

  • Leverage AI coding assistants as part of your development workflow while evaluating their strengths, limitations, and failure modes on complex engineering tasks.

  • Review benchmark tasks created by fellow experts, providing detailed feedback on technical correctness, clarity, completeness, and level of difficulty.

  • Analyze AI-generated solutions to identify implementation errors, debugging challenges, reasoning gaps, and opportunities for improving benchmark quality.

  • Collaborate closely with AI researchers and engineering teams to refine evaluation methodologies and maintain consistently high technical standards.


Required Qualifications


  • Master's degree, PhD, or equivalent practical experience in Computer Science, Software Engineering, or another STEM discipline involving substantial programming and technical problem solving.

  • Minimum 1 year of professional experience in software engineering, research engineering, applied research, or another technically intensive engineering role.

  • Strong hands-on proficiency in Python, including software design, scripting, debugging, testing, and writing clean, maintainable code.

  • Solid experience using Git, integrated development environments (IDEs), and modern software development workflows.

  • Strong understanding of software architecture, debugging methodologies, and engineering best practices.

  • Experience using AI coding assistants, prompt engineering techniques, or AI agent workflows is preferred.

  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.

  • Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently.

  • Strong written communication skills for documenting technical solutions and engineering decisions.

  • Ability to commit approximately 35 hours per week on a consistent basis.


Preferred Qualifications


  • Experience developing complex software systems across backend, infrastructure, automation, or developer tooling.

  • Familiarity with large language models, AI-assisted software development, or autonomous coding agents.

  • Background in benchmark design, software quality engineering, or technical evaluation frameworks.

  • Contributions to open-source projects, technical publications, or advanced engineering initiatives.


Why Join


  • Help shape how next-generation AI coding systems are evaluated against real-world engineering challenges.

  • Collaborate with leading AI researchers developing frontier evaluation benchmarks.

  • Apply your software engineering expertise to improve AI reliability, technical reasoning, and coding performance.

  • Contribute directly to benchmark development that influences the evolution of advanced AI systems.

  • Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives.


Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.


Contract & Engagement Details


  • Independent contractor engagement.

  • Fully remote with flexible working hours.

  • Expected commitment of approximately 35 hours per week.

  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.

  • Work does not require access to confidential or proprietary information from any current or former employer.

  • Payments are issued weekly based on approved work completed.

  • At this time, we are unable to support H1-B or STEM OPT candidates.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday AI • United States

Remote
USD 83,000 - 124,000
LLM Red Team Specialist - Failure Modes & Edge Cases
LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday AI • United States

Remote
USD 171,924,000 - 257,887,000
STEM Researcher - Computational Fields
STEM Researcher - Computational Fields

Weekday AI • United States

Remote
USD 83,000 - 124,000
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday AI • United States

Remote
USD 83,000 - 124,000
QA/Test Engineer
QA/Test Engineer

Weekday AI • United States

Remote
USD 83,000 - 124,000
Software Engineering Expert (AI Training)
Software Engineering Expert (AI Training)

Weekday AI • United States

Remote
USD 138,000 - 317,000
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour

24-MAG • United States

Remote
USD 69,000 - 138,000
Remote work
Flexible hours
Contractor engagement
Remote | Senior Software Engineer — $50–$100/hour
Remote | Senior Software Engineer — $50–$100/hour

engineeringjobs.net, Inc. • New York (NY)

Remote
USD 69,000 - 138,000
SWE-Bench AI Task Auditor - Freelance AI Trainer Project
SWE-Bench AI Task Auditor - Freelance AI Trainer Project

Jobgether SRL • United States

Remote
USD 68,000 - 97,000
Fully remote freelance contract
Flexible project-based work
Impact on AI training workflows
Software Engineer - Remote - Contract
Software Engineer - Remote - Contract

Xperteez Technology Pvt Ltd • United States

Remote
USD 34,000 - 83,000