Security Software Engineer - Python & AI Evaluation

Braintrust

United States

Remote

USD 107,453,000 - 128,943,000

Part time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Leading AI Lab is seeking an experienced software engineer for security-focused coding task evaluation and prompt creation. This contracting engagement involves evaluating vulnerabilities, validating exploits, and writing clear feedback, with potential for longer-term work.

Remote candidates in the United States, Canada and 21+ other countries are eligible; expected 20 hours per week at $75–$90/hr, requiring 5+ years of software development and strong Python skills.

Qualifications

  • At least five years of professional software-development experience and strong Python skills.
  • Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
  • The ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluency in written and spoken English.

Responsibilities

  • Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
  • Create high-quality coding prompts and reference answers for benchmark-style problems.
  • Evaluate model outputs for code generation, refactoring, debugging and implementation.
  • Identify and document model failures, edge cases and reasoning gaps.
  • Compare private language models with leading external models.
  • Build or configure coding environments for evaluation and reinforcement learning.
  • Follow detailed annotation and evaluation guidelines consistently.

Skills

Python
Vulnerability research
Security patching
Technical feedback
English communication

Job description

  • Rate: $75 – $90/hr
  • Hours: 20 hours / week
  • Experience: 5 - 10 years
  • Location: United States | Canada + 21 more

Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands‑on security experience to work involving vulnerabilities, exploit verification and security patches.

This is a contracting engagement, with potential for a longer‑term engagement. Remote candidates in the selected countries are elegible.

What you will do
  • Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
  • Create high‑quality coding prompts and reference answers for benchmark‑style problems.
  • Evaluate model outputs for code generation, refactoring, debugging and implementation.
  • Identify and document model failures, edge cases and reasoning gaps.
  • Compare private language models with leading external models.
  • Build or configure coding environments for evaluation and reinforcement learning.
  • Follow detailed annotation and evaluation guidelines consistently.
What you bring
  • At least five years of professional software‑development experience and strong Python skills.
  • Hands‑on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
  • The ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluency in written and spoken English.
Helpful, not required
  • Professional code review, coding annotation, LLM/code evaluation or benchmark design.
  • Knowledge of another programming language.
  • Team leadership or mentoring experience.

Company: Leading AI Lab

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Application Security Engineer - AI Code Evaluation
Senior Application Security Engineer - AI Code Evaluation

Braintrust • United States

Remote
USD 103,000 - 124,000
Remote Security Software Engineer (Python) for AI Evaluation
Remote Security Software Engineer (Python) for AI Evaluation

Braintrust • United States

Remote
USD 107,453,000 - 128,943,000
Remote Senior App Security Engineer for AI Code Eval
Remote Senior App Security Engineer for AI Code Eval

Braintrust • United States

Remote
USD 103,000 - 124,000
GenAI Security Evaluation Engineer (Up to $150/hr)
GenAI Security Evaluation Engineer (Up to $150/hr)

Turing • United States

Remote
USD 170,000 - 243,000
Security Engineer - Fully Remote | Upto $85/hr
Security Engineer - Fully Remote | Upto $85/hr

mercor • San Francisco (CA)

Hybrid
GBP 73,000 - 104,000
Cybersecurity & Software Engineering Specialist (AI Projects)
Cybersecurity & Software Engineering Specialist (AI Projects)

Gramian Consultancy Group • United States

Remote
USD 120,000 - 160,000
Remote contractor position
Software Engineer (Python/C++) – Code Evaluation | Remote
Software Engineer (Python/C++) – Code Evaluation | Remote

Crossing Hurdles • United States

On-site
GBP 25,000 - 41,000
Principal Coding Annotator / LLM Evaluation Engineer
Principal Coding Annotator / LLM Evaluation Engineer

Braintrust • United States

Remote
USD 103,000 - 124,000
Backend Security Engineer - AI Code Review and Evaluation
Backend Security Engineer - AI Code Review and Evaluation

AuraOne • United States

Remote
USD 165,000 - 193,000
Backend Security Engineer – AI Coding Evaluation Project
Backend Security Engineer – AI Coding Evaluation Project

eDataBae • United States

On-site
USD 120,000 - 180,000