RLVR Engineer — Verifiable Rewards for AI Security Roles

Bugcrowd Inc.

United States

Remote

USD 176,000 - 243,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bugcrowd is seeking a Reinforcement Learning Engineer specializing in RLVR to design and scale automated verification pipelines that convert real-world vulnerabilities into deterministic reward signals. You will bridge low-level security analysis with modern LLM reasoning, building environments where AI agents learn to discover, exploit, and remediate software vulnerabilities with mathematical certainty.

You will work at the intersection of fuzzing, dynamic program analysis, system exploitation,

Qualifications

  • Understanding RL training workflows used by modern LLM systems (RLVR).
  • Proficiency developing applications in Python and low-level systems programming in C; Rust is a strong plus.
  • Experience with DevOps pipelines (GitHub Actions), reproducible builds (Docker, BuildKit, Nix).

Responsibilities

  • Design, build, and deploy high-throughput RLVR environments that evaluate LLM action sequences against deterministic execution outcomes.
  • Develop automated test harnesses, sandboxes, and verification engines that convert vulnerability research into binary pass/fail signals.
  • Integrate Bugcrowd’s Mayhem platform and vulnerability feeds into RL environment pipelines.
  • Architect safe, isolated execution environments using Docker, BuildKit, or Nix for thousands of agent-driven trajectories.
  • Collaborate with researchers at AI labs to define benchmark formats, observation spaces, and verifiable metrics.
  • Implement telemetry, ground-truth verifiers, and trajectory logging to analyze reasoning paths and prevent reward hacking.
  • Build low-level instrumentation to monitor memory, processes, and network behaviors during agent interactions.
  • Optimize infrastructure performance and environment reset latency for large-scale RL training.

Skills

RLVR
Python
C
Rust
DevOps
Linux
Docker
Build systems

Tools

Docker
BuildKit
Nix

Job description

Bugcrowd is seeking a Reinforcement Learning Engineer specializing in RLVR to design and scale automated verification pipelines that convert real-world vulnerabilities into deterministic reward signals. You will bridge low-level security analysis with modern LLM reasoning, building environments where AI agents learn to discover, exploit, and remediate software vulnerabilities with mathematical certainty.

You will work at the intersection of fuzzing, dynamic program analysis, system exploitation,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote RL Engineer for Cybersecurity AI Training
Remote RL Engineer for Cybersecurity AI Training

Bugcrowd • United States

On-site
USD 176,400 - 242,550
Reinforcement Learning Engineer
Reinforcement Learning Engineer

Bugcrowd Inc. • United States

Remote
USD 176,000 - 243,000
Remote AI Training Engineer for RL Software Tasks
Remote AI Training Engineer for RL Software Tasks

YO AI Labs • Maryland

Remote
USD 69,000 - 165,000
AI Research Engineer - RL Security & Red Teaming
AI Research Engineer - RL Security & Red Teaming

Qualis • Sunnyvale (CA)

On-site
USD 150,000 - 230,000
Research Engineer, Cybersecurity RL (Reinforcement Learning)
Research Engineer, Cybersecurity RL (Reinforcement Learning)

Anthropic • San Francisco (CA)

On-site
USD 300,000 - 405,000
Equity donation matching
Generous vacation
Flexible working hours
+1
Remote AI Research Engineer — RL & LLM Systems
Remote AI Research Engineer — RL & LLM Systems

Not specified • United States

Remote
USD 120,000 - 180,000
Remote RL Engineer - Scale & Deploy AI
Remote RL Engineer - Scale & Deploy AI

Bright Vision Technologies • Naperville (IL)

Remote
USD 96,000 - 120,000
Senior AI Software Engineer - Remote RL Environments
Senior AI Software Engineer - Remote RL Environments

YO AI Labs • Houston (TX)

Remote
USD 5,510,000 - 11,021,000
Remote Senior Software Engineer – RL Environments for AI
Remote Senior Software Engineer – RL Environments for AI

YO AI Labs • Los Angeles (CA)

Remote
USD 96,000 - 165,000
Research Scientist, Reinforcement Learning
Research Scientist, Reinforcement Learning

Pramaana Labs • Palo Alto (CA)

On-site
USD 150,000 - 230,000