MTS - Research (Cybersecurity)

Collinear AI

Sunnyvale (CA)

On-site

USD 120,000 - 180,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Collinear AI in Sunnyvale is building cybersecurity environments, tasks, and evaluations to help frontier AI models improve at real work. We are growing our security team and seeking hands-on security experts who can turn practical problems into runnable environments, with verifiable fixes and scalable testing.

You will own work from initial scenario through runnable environment, task instructions, reference solution, and verifier.

Qualifications

  • Hands-on security expertise across cybersecurity domains.
  • Ability to read unfamiliar code, reproduce issues, and implement fixes.
  • Comfort with Linux, Git, containers, debugging, and automated testing.
  • Clear technical writing and careful judgment on evaluation realism and trustworthiness.
  • Curiosity about applying cybersecurity to AI; prior LLM experience not required.

Responsibilities

  • Build cybersecurity environments to host realistic tasks and assessments.
  • Turn security work into clearly defined agent tasks and objectives.
  • Develop reference solutions and verifiers to validate fixes.
  • Test the tests and guard against shortcuts and hidden solutions.
  • Run and analyze agents, reviewing tool calls and changes for gaps.
  • Contribute to benchmarks and research with researchers and engineers.

Skills

Practical cybersecurity
Linux
Git
Containers
Debugging
Automated testing
Technical writing
Curiosity about AI/LLMs

Job description

Collinear AI builds environments, tasks, and evaluations that help frontier AI models improve at real work. We are growing our cybersecurity team and looking for someone who can turn practical security problems into environments where AI agents can investigate, act, and learn.

We recently released CWE-bench, a defensive cybersecurity benchmark with 100 held‑out audit‑and‑patch tasks across 54 weakness types. Agents must find and fix vulnerabilities in real codebases, with checks that confirm the vulnerability is resolved and existing functionality still works. You will help build what comes next: richer cybersecurity environments, realistic tasks, and reliable ways to measure whether agents succeed.

You will own work from the initial security scenario through the runnable environment, task instructions, reference solution, and verifier. We are looking for hands‑on security knowledge, strong programming skills, and curiosity about how AI agents fail.

Deep cybersecurity expertise is enough to get started—you do not need prior AI or machine‑learning experience. If you know how to investigate vulnerabilities, reason about security failures, and verify that a fix works, we want to hear from you. We will teach you our AI tooling and evaluation workflows.

What you will do
  • Build cybersecurity environments. Create isolated, reproducible environments with real repositories, applications, services, logs, and access controls. Make them straightforward to launch, reset, and evaluate at scale.
  • Turn security work into tasks. Design scenarios around vulnerability discovery and remediation, application and API security, authentication and authorization, incident investigation, and system hardening. Define what the agent knows, which tools it can use, and what it must accomplish.
  • Develop reference solutions and verifiers. Reproduce the underlying issue, implement a valid solution, and write checks that distinguish a real fix from a superficial workaround. Verify that attacks fail after remediation while legitimate behavior continues to work.
  • Test the tests. Challenge graders with incomplete fixes, disabled features, hard‑coded answers, and other shortcuts. Keep hidden solutions and test data out of the agent's environment, and separate genuine model failures from broken infrastructure.
  • Run and analyze agents. Evaluate frontier models, inspect their tool calls and code changes, and explain where their security reasoning or execution breaks down. Use those findings to improve task coverage and difficulty.
  • Contribute to benchmarks and research. Work with researchers and engineers to turn strong environments into evaluation suites and training data. Review other contributors' tasks and help document results for future releases.
Who we are looking for
  • Practical depth in at least one area of cybersecurity, such as application security, vulnerability research, penetration testing, systems security, cloud security, or incident response. Evidence can come from internships, research, open‑source work, bug bounties, CTFs, or independent projects.
  • The ability to read unfamiliar code, reproduce a security issue, understand its root cause, and implement or assess a fix.
  • Comfort with Linux, Git, containers, debugging, and automated testing.
  • Clear technical writing and careful judgment about what an evaluation does and does not demonstrate. You can explain why a task is realistic and why its grading is trustworthy.
  • Curiosity about applying your cybersecurity expertise to AI, and a willingness to learn how to evaluate tool‑using agents. Prior experience with LLMs is not expected.

NOTE: Recent graduates and early‑career engineers or researchers are encouraged to apply. A PhD, professional certification, or previous role at an AI lab is not required. We value demonstrated ability and the quality of your work.

Nice to have
  • Disclosed vulnerabilities, accepted security patches, strong CTF results, or useful security tools and write‑ups.
  • Experience with fuzzing, static or dynamic analysis, reverse engineering, or building security labs.
  • Familiarity with CWE and OWASP classifications and how they relate to concrete software failures.
  • Experience building agent evaluations, adversarial tests, prompt‑injection defenses, or reinforcement‑learning environments.
Examples of what you might build
  • A repository audit where an agent must discover and repair an authorization flaw without being told where it is.
  • A small service environment where an agent investigates suspicious activity from logs and configuration, then applies and verifies a remediation.
  • An LLM application where an agent must repair a prompt‑injection or tool‑permission weakness while preserving legitimate functionality.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MTS - Research (Cybersecurity)
MTS - Research (Cybersecurity)

Collinear AI, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Evals Engineer, Offensive Cyber
Evals Engineer, Offensive Cyber

Zealot Labs • New York (NY)

On-site
USD 140,000 - 200,000
Applied AI Engineer (Autonomous Defense)
Applied AI Engineer (Autonomous Defense)

Horizon3 • United States

Remote
USD 140,000 - 210,000
Growth opportunities
Innovation-driven culture
Flexible work environment
+1
Frontier AI Offensive Security Specialist
Frontier AI Offensive Security Specialist

Gexel Telecom International Inc. • Santa Clara (CA)

Hybrid
USD 140,000 - 200,000
Specialized training
Mentorship
Global cybersecurity organization
AI SOC Engineer
AI SOC Engineer

ByLabs • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI/ML Engineer - Cybersecurity
AI/ML Engineer - Cybersecurity

Scan Ninja Inc. • Houston (TX)

On-site
USD 100,000 - 130,000
Frontier AI Offensive Security Specialist
Frontier AI Offensive Security Specialist

Hitachi Cyber • United States

On-site
USD 140,000 - 190,000
Frontier AI focus
Specialized AI security training
Mentorship within global cyber team
+1
Research Engineer, Evaluations
Research Engineer, Evaluations

General Analysis • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 230,000
Cybersecurity for AI Safety (Contract)
Cybersecurity for AI Safety (Contract)

Next Frontier Capital • United States

On-site
USD 103,000 - 165,000
Cybersecurity for AI Safety (Contract)
Cybersecurity for AI Safety (Contract)

Empathy • Northern (KY)

On-site
USD 82,656 - 165,312