Python Engineer

Explore more jobs

Kuwait City

On-site

KWD 26,000 - 32,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Explore more jobs seeks a software-testing specialist to help build evaluation tasks for AI coding agents. You will design realistic developer environments and write tests that differentiate safe from unsafe completion.

This role focuses on ensuring tests cover breadth of repos, CI workflows and codebases, with 4-5+ years in software development and hands-on agent experience.

Qualifications

  • 4-5+ years in software development.
  • Strong test design skills - functional and integration tests that separate safe from unsafe completion.
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar).
  • Familiarity with GitHub PRs and CI workflows as a user.
  • Stack breadth is welcome; exposure to databases, CI pipelines, and deploy scripts is useful.

Responsibilities

  • Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes.
  • Write tests that verify whether the agent completed the task the right way - catching corners cut, not just outputs.
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until evaluation is fair and robust.

Skills

Test design
Software development
QA automation

Tools

Claude Code
GitHub Copilot CLI
Codex

Job description

What this opportunity involves

Frontier coding agents are already good at passing tests. We measure whether they pass them the right way. We're building a dataset to evaluate the safety and conduct of AI coding agents - not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.

You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
  • Write tests that verify whether the agent completed the task the right way - catching corners cut, not just checking outputs
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is not:
  • Not cybersecurity or red-teaming - there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
  • Not writing code from scratch - the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
  • 4-5+ years in software development;
  • Strong test design skills - functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
  • Familiarity with GitHub PRs and CI workflows as a user;
  • Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful - but you don't need to be an expert in every layer;
Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation - a scenario where the unsafe or out-of-scope path is the path of least resistance - and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.

Project time expectations

For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

On this project, contributors can earn up to $75 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Evaluation Specialist
Software Engineering Evaluation Specialist

Mindrift • Kuwait

On-site
KWD 8,900 - 15,000
Paid contributions
Incentive payments
Engineer
Engineer

Explore more jobs • Kuwait City

On-site
KWD 18,000 - 25,000
Senior Python Engineer
Senior Python Engineer

Explore more jobs • Kuwait City

On-site
KWD 20,000 - 33,000
Agent Evaluation Engineer
Agent Evaluation Engineer

Explore more jobs • Kuwait City

On-site
KWD 17,000 - 21,000
Competitive Programming Expert - Freelance AI Trainer
Competitive Programming Expert - Freelance AI Trainer

Mindrift • Kuwait

On-site
KWD 17,000 - 38,000
Compliance Analyst
Compliance Analyst

Explore more jobs • Kuwait

Hybrid
KWD 18,000 - 25,000
Compliance Analyst
Compliance Analyst

Explore more jobs • Kuwait City

On-site
KWD 17,000 - 26,000
Electrical Engineer and Python Expert
Electrical Engineer and Python Expert

Explore more jobs • Kuwait City

On-site
KWD 9,400 - 16,000
Python Developer
Python Developer

Explore more jobs • Kuwait

On-site
KWD 10,000 - 17,000
Mechanical Engineer and Python Expert
Mechanical Engineer and Python Expert

Explore more jobs • Kuwait City

On-site
KWD 9,400 - 16,000