Intern: AI Red Teaming (Fall 2026)

Realm Labs

Sunnyvale (CA)

On-site

USD 34,000 - 55,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Market aligned intern compensation -,u

Job summary

Realm Labs in Sunnyvale, CA is seeking an Intern for AI Red Teaming in Fall 2026. This in-office, full-time role focuses on breaking systems and models, turning findings into reproducible attacks, and documenting why they work.

You will engage with adversarial ML methods and agentic systems, with a goal to publish a technical report and improve the evaluation suite. Ideal candidates have hands-on attack experience, ability to implement research papers, and familiarity with ML tools and security

Qualifications

  • Hands-on experience attacking or stress-testing models.
  • Read a paper and implement its attack.
  • (nice to have) Offensive security background outside ML: CTFs, vulnerability research.
  • (nice to have) Familiarity with agentic systems and their attack surface.

Responsibilities

  • You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works.
  • Expect mix of eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation; and measuring guardrails under pressure.
  • Aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.

Skills

Adversarial ML
Red Teaming
Python
Pytorch
LLM safety

Tools

AWS
GCP
Jupyter

Job description

Intern: AI Red Teaming (Fall 2026)

Sunnyvale, CA

AI/ML

In office

Full-time

Role Overview
  • You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works.
  • Expect some mix of: eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation rather than hand-crafting one-off prompts; and measuring whether guardrails hold under pressure. Where RealmLabs' interpretability work gives you access to a model's internals, use it.
  • We aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.
Expected Background: Adversarial ML and Red Teaming
  • Hands-on experience attacking or stress-testing models, from any direction: jailbreaks, prompt injection, adversarial examples, data poisoning, model extraction, or evaluating safety and moderation systems.
  • Able to read a paper and implement its attack.
  • (nice to have) Offensive security background outside ML: CTFs, vulnerability research, penetration testing.
  • (nice to have) Familiarity with agentic systems and their attack surface - tool calls, retrieval, memory, multi-agent orchestration.
Expected Background: ML
  • Machine learning tools: pytorch, huggingface, transformers, datasets.
  • Applied deep learning and LLM experience.
    • Training and evaluating deep models.
    • (nice to have) finetuning LLMs, multi-modal LLMs.
  • (nice to have) Familiarity with ML[NLP,LLM,Vision] interpretability methods, sparse autoencoders, linear probes - as a way of locating failure modes, not as an end in itself.
Expected Background: Software Engineering
  • Development environments and tools:
    • unix, git, basic clouds usage on AWS and/or GCP
    • jupyter
  • Programming:
    • python
    • (nice to have) "programming languages well-roundedness"\
    • experience in statically-typed and functional languages
Compensation & Benefits
  • Market aligned compensation for interns in the bay area.
Requirements
  • Must be authorized to work in the USA or must be able to obtain CPT (Curricular Practical Training) approval from host university.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Red Team Intern: Break & Harden Safe Models
AI Red Team Intern: Break & Harden Safe Models

Realm Labs • Sunnyvale (CA)

On-site
USD 34,000 - 55,000
Market aligned intern compensation -,u
ML Engineer Intern | Summer 2026
ML Engineer Intern | Summer 2026

Crustdata (YC F24) • San Francisco (CA)

On-site
Competitive stipend
Housing stipend for relocation
Direct mentorship from founders
+1
AI Red Teamer (LLM Generalist)
AI Red Teamer (LLM Generalist)

Handshake • Seattle (WA)

On-site
USD 83,000 - 138,000
Member of Technical Staff (intern)
Member of Technical Staff (intern)

Adaptive ML • New York (NY)

On-site
Paid internship
Mentorship
Exposure to real-world AI systems
AI Red Team Engineer
AI Red Team Engineer

Confidential • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer Internship, Agent Systems
Software Engineer Internship, Agent Systems

Armadin • Palo Alto (CA)

On-site
USD 88,000 - 118,000
Health, dental, and vision coverage
Equity ownership
In-office meals
+4
AI Red Teaming Engineer
AI Red Teaming Engineer

DeWinter Group • Campbell (CA)

Remote
Research Intern: Interpretability & Reliability (Summer 2027)
Research Intern: Interpretability & Reliability (Summer 2027)

CTGT • San Francisco (CA)

On-site
USD 52,000 - 68,000
Strategic Projects Lead, Red Team
Strategic Projects Lead, Red Team

Front Door Defense • New York (NY)

On-site
USD 152,000 - 190,000
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+2
Head of AI Red Teaming
Head of AI Red Teaming

Trajectory Labs, PBC • Berkeley (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Equity
Health coverage
401(k)
+1