Senior Application Security Engineer - AI Code Evaluation

Braintrust

United States

Remote

USD 103,000 - 124,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Leading AI Lab is seeking a contractor to evaluate security-focused coding tasks and improve LLM performance. Remote candidates in the United States and other eligible locations may participate for a 20 hours per week commitment at $75–$90 per hour.

You will bring 5+ years of software development and strong Python skills, along with hands-on vulnerability research, exploit verification, and security patch experience. Fluency in English is required.

Qualifications

  • Minimum 5 years of professional software development experience.
  • Strong Python skills required.
  • Hands-on vulnerability research or exploit verification experience.
  • Ability to document findings with clear technical feedback.
  • Fluent in written and spoken English.

Responsibilities

  • Evaluate coding tasks involving vulnerabilities, exploit verification and security patches.
  • Create high-quality coding prompts and reference answers for benchmarks.
  • Evaluate model outputs for code generation, refactoring, debugging and implementation.
  • Identify and document model failures, edge cases and reasoning gaps.
  • Compare private language models with leading external models.
  • Build or configure coding environments for evaluation and reinforcement learning.
  • Follow detailed annotation and evaluation guidelines consistently.

Skills

5+ years software development
Python
Vulnerability research
Exploit verification
Security patches
English fluency

Tools

Python

Job description

  • Rate: $75 – $90/hr
  • Hours: 20 hours / week
  • Experience: 5 - 10 years
  • Location: United States | Canada + 21 more

Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands‑on security experience to work involving vulnerabilities, exploit verification and security patches.

This is a contracting engagement, with potential for a longer‑term engagement. Remote candidates in the selected countries are elegible.

What you will do
  • Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
  • Create high‑quality coding prompts and reference answers for benchmark‑style problems.
  • Evaluate model outputs for code generation, refactoring, debugging and implementation.
  • Identify and document model failures, edge cases and reasoning gaps.
  • Compare private language models with leading external models.
  • Build or configure coding environments for evaluation and reinforcement learning.
  • Follow detailed annotation and evaluation guidelines consistently.
What you bring
  • At least five years of professional software‑development experience and strong Python skills.
  • Hands‑on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
  • The ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluency in written and spoken English.
Helpful, not required
  • Professional code review, coding annotation, LLM/code evaluation or benchmark design.
  • Knowledge of another programming language.
  • Team leadership or mentoring experience.

Company: Leading AI Lab

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Security Software Engineer - Python & AI Evaluation
Security Software Engineer - Python & AI Evaluation

Braintrust • United States

Remote
USD 107,453,000 - 128,943,000
Remote Senior App Security Engineer for AI Code Eval
Remote Senior App Security Engineer for AI Code Eval

Braintrust • United States

Remote
USD 103,000 - 124,000
Remote Security Software Engineer (Python) for AI Evaluation
Remote Security Software Engineer (Python) for AI Evaluation

Braintrust • United States

Remote
USD 107,453,000 - 128,943,000
GenAI Security Evaluation Engineer (Up to $150/hr)
GenAI Security Evaluation Engineer (Up to $150/hr)

Turing • United States

Remote
USD 170,000 - 243,000
Principal Coding Annotator / LLM Evaluation Engineer
Principal Coding Annotator / LLM Evaluation Engineer

Braintrust • United States

Remote
USD 103,000 - 124,000
Backend Security Engineer - AI Code Review and Evaluation
Backend Security Engineer - AI Code Review and Evaluation

AuraOne • United States

Remote
USD 165,000 - 193,000
Application Security Analyst
Application Security Analyst

Alignerr Corp. • United States

On-site
MXN 551,000 - 827,000
Security Engineer - Fully Remote | Upto $85/hr
Security Engineer - Fully Remote | Upto $85/hr

mercor • San Francisco (CA)

Hybrid
GBP 73,000 - 104,000
Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based
Software & Firmware Engineering Experts — AI Code Evaluation
Software & Firmware Engineering Experts — AI Code Evaluation

AuraOne, Inc. • United States

Remote
USD 138,000 - 165,000