AI Red Team Intern: Break & Harden Safe Models

Realm Labs

Sunnyvale (CA)

On-site

USD 34,000 - 55,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Market aligned intern compensation -,u

Job summary

Realm Labs in Sunnyvale, CA is seeking an Intern for AI Red Teaming in Fall 2026. This in-office, full-time role focuses on breaking systems and models, turning findings into reproducible attacks, and documenting why they work.

You will engage with adversarial ML methods and agentic systems, with a goal to publish a technical report and improve the evaluation suite. Ideal candidates have hands-on attack experience, ability to implement research papers, and familiarity with ML tools and security

Qualifications

  • Hands-on experience attacking or stress-testing models.
  • Read a paper and implement its attack.
  • (nice to have) Offensive security background outside ML: CTFs, vulnerability research.
  • (nice to have) Familiarity with agentic systems and their attack surface.

Responsibilities

  • You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works.
  • Expect mix of eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation; and measuring guardrails under pressure.
  • Aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.

Skills

Adversarial ML
Red Teaming
Python
Pytorch
LLM safety

Tools

AWS
GCP
Jupyter

Job description

Realm Labs in Sunnyvale, CA is seeking an Intern for AI Red Teaming in Fall 2026. This in-office, full-time role focuses on breaking systems and models, turning findings into reproducible attacks, and documenting why they work.

You will engage with adversarial ML methods and agentic systems, with a goal to publish a technical report and improve the evaluation suite. Ideal candidates have hands-on attack experience, ability to implement research papers, and familiarity with ML tools and security

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Intern: AI Red Teaming (Fall 2026)
Intern: AI Red Teaming (Fall 2026)

Realm Labs • Sunnyvale (CA)

On-site
USD 34,000 - 55,000
Market aligned intern compensation -,u
AI Red Teamer: Offensive Security for AI Systems (Remote)
AI Red Teamer: Offensive Security for AI Systems (Remote)

Handshake • United States

Remote
USD 150,000 - 210,000
AI Red Team Engineer
AI Red Team Engineer

Confidential • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Red Teaming Engineer
AI Red Teaming Engineer

DeWinter Group • Campbell (CA)

Remote
AI Red Team Specialist — Adversarial Testing (Remote)
AI Red Team Specialist — Adversarial Testing (Remote)

Mercor • San Francisco (CA)

Remote
USD 130,000 - 170,000
Remote AI Red Team Engineer
Remote AI Red Team Engineer

Mercor • New York (NY)

Remote
USD 90,000 - 160,000
AI Red Teamer: LLM Safety & Stress-Testing Generalist
AI Red Teamer: LLM Safety & Stress-Testing Generalist

Handshake • Seattle (WA)

On-site
USD 83,000 - 138,000
AI Security Research Engineer: RL Training & Red Teaming
AI Security Research Engineer: RL Training & Red Teaming

Qualis • Sunnyvale (CA)

On-site
USD 120,000 - 180,000
Remote AI Red Team - Safety & Adversarial Testing
Remote AI Red Team - Safety & Adversarial Testing

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 140,000
Staff AI Safety Engineer — Red Team & Guardrails
Staff AI Safety Engineer — Red Team & Guardrails

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Top-tier compensation
Stock options
Health & wellness
+3