MTS - Research (Cybersecurity)

Collinear AI, Inc.

Sunnyvale (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Collinear AI, Inc. is growing its cybersecurity team and seeks a hands-on engineer to design executable security environments, real repositories, and verifiable tasks for AI agents. You will own the project from scenario to runnable setup and verifier.

You bring deep cybersecurity know-how and strong Python and systems programming skills. No AI/ML background required; we will teach our tooling and evaluation workflows. Recent grads are encouraged to apply.

Qualifications

  • Hands-on cybersecurity knowledge with strong programming skills.
  • Ability to read unfamiliar code, reproduce issues, and implement fixes.
  • Experience with Python and at least one systems language (C/C++, Go, Java, JS/TS, or Rust).
  • Familiarity with Linux, Git, containers, debugging, and automated testing.
  • Recent graduates or early-career engineers encouraged to apply; PhD not required.

Responsibilities

  • Build cybersecurity environments. Create isolated, reproducible environments with real repositories, applications, services, logs, and access controls.
  • Turn security work into tasks. Design scenarios around vulnerability discovery and remediation, application and API security, authentication and authorization, incident investigation, and system hardening.
  • Develop reference solutions and verifiers. Reproduce the underlying issue, implement a valid solution, and write checks that distinguish a real fix from a superficial workaround.
  • Test the tests. Challenge graders with incomplete fixes, disabled features, hard-coded answers, and other shortcuts.
  • Run and analyze agents. Evaluate frontier models, inspect their tool calls and code changes, and explain where their security reasoning or execution breaks down.
  • Contribute to benchmarks and research. Work with researchers and engineers to turn strong environments into evaluation suites and training data.

Skills

Python
C/C++
Go
Java
JavaScript/TypeScript
Rust
Linux
Git
Containers
Debugging
Automated testing
Code analysis

Job description

Collinear AI builds environments, tasks, and evaluations that help frontier AI models improve at real work. We are growing our cybersecurity team and looking for someone who can turn practical security problems into environments where AI agents can investigate, act, and learn.

We recently released CWE-bench, a defensive cybersecurity benchmark with 100 held-out audit-and-patch tasks across 54 weakness types. Agents must find and fix vulnerabilities in real codebases, with checks that confirm the vulnerability is resolved and existing functionality still works. You will help build what comes next: richer cybersecurity environments, realistic tasks, and reliable ways to measure whether agents succeed.

You will own work from the initial security scenario through the runnable environment, task instructions, reference solution, and verifier. We are looking for hands-on security knowledge, strong programming skills, and curiosity about how AI agents fail.

Deep cybersecurity expertise is enough to get started- you do not need prior AI or machine-learning experience. If you know how to investigate vulnerabilities, reason about security failures, and verify that a fix works, we want to hear from you. We will teach you our AI tooling and evaluation workflows.

What you will do
  • Build cybersecurity environments. Create isolated, reproducible environments with real repositories, applications, services, logs, and access controls. Make them straightforward to launch, reset, and evaluate at scale.

  • Turn security work into tasks. Design scenarios around vulnerability discovery and remediation, application and API security, authentication and authorization, incident investigation, and system hardening. Define what the agent knows, which tools it can use, and what it must accomplish.

  • Develop reference solutions and verifiers. Reproduce the underlying issue, implement a valid solution, and write checks that distinguish a real fix from a superficial workaround. Verify that attacks fail after remediation while legitimate behavior continues to work.

  • Test the tests. Challenge graders with incomplete fixes, disabled features, hard-coded answers, and other shortcuts. Keep hidden solutions and test data out of the agent's environment, and separate genuine model failures from broken infrastructure.

  • Run and analyze agents. Evaluate frontier models, inspect their tool calls and code changes, and explain where their security reasoning or execution breaks down. Use those findings to improve task coverage and difficulty.

  • Contribute to benchmarks and research. Work with researchers and engineers to turn strong environments into evaluation suites and training data. Review other contributors' tasks and help document results for future releases.

Who we are looking for
  • Practical depth in at least one area of cybersecurity, such as application security, vulnerability research, penetration testing, systems security, cloud security, or incident response. Evidence can come from internships, research, open-source work, bug bounties, CTFs, or independent projects.

  • The ability to read unfamiliar code, reproduce a security issue, understand its root cause, and implement or assess a fix.

  • Strong programming skills in Python and at least one language used in the systems you investigate, such as C/C++, Go, Java, JavaScript/TypeScript, or Rust.

  • Comfort with Linux, Git, containers, debugging, and automated testing.

  • Clear technical writing and careful judgment about what an evaluation does and does not demonstrate. You can explain why a task is realistic and why its grading is trustworthy.

  • Curiosity about applying your cybersecurity expertise to AI, and a willingness to learn how to evaluate tool-using agents. Prior experience with LLMs is not expected.

Recent graduates and early-career engineers or researchers are encouraged to apply. A PhD, professional certification, or previous role at an AI lab is not required. We value demonstrated ability and the quality of your work.

Nice to have
  • Disclosed vulnerabilities, accepted security patches, strong CTF results, or useful security tools and write-ups.

  • Experience with fuzzing, static or dynamic analysis, reverse engineering, or building security labs.

  • Familiarity with CWE and OWASP classifications and how they relate to concrete software failures.

  • Experience building agent evaluations, adversarial tests, prompt-injection defenses, or reinforcement-learning environments.

Examples of what you might build
  • A repository audit where an agent must discover and repair an authorization flaw without being told where it is.

  • A small service environment where an agent investigates suspicious activity from logs and configuration, then applies and verifies a remediation.

  • An LLM application where an agent must repair a prompt-injection or tool-permission weakness while preserving legitimate functionality.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MTS - Research (Cybersecurity)
MTS - Research (Cybersecurity)

Collinear AI • Sunnyvale (CA)

On-site
USD 120,000 - 180,000
Applied AI Engineer (Autonomous Defense)
Applied AI Engineer (Autonomous Defense)

Horizon3 • United States

Remote
USD 140,000 - 210,000
Growth opportunities
Innovation-driven culture
Flexible work environment
+1
Evals Engineer, Offensive Cyber
Evals Engineer, Offensive Cyber

Zealot Labs • New York (NY)

On-site
USD 140,000 - 200,000
Research Engineer, Evaluations
Research Engineer, Evaluations

General Analysis • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 230,000
Frontier AI Offensive Security Specialist
Frontier AI Offensive Security Specialist

Gexel Telecom International Inc. • Santa Clara (CA)

Hybrid
USD 140,000 - 200,000
Specialized training
Mentorship
Global cybersecurity organization
Frontier AI Offensive Security Specialist
Frontier AI Offensive Security Specialist

Hitachi Cyber • United States

On-site
USD 140,000 - 190,000
Frontier AI focus
Specialized AI security training
Mentorship within global cyber team
+1
AI Red Teamer (Cybersecurity)
AI Red Teamer (Cybersecurity)

Handshake • United States

Remote
USD 140,000 - 210,000
Health Insurance
401k match
Parental leave
+1
AI Engineer - SF
AI Engineer - SF

Foundation Capital • San Francisco (CA)

On-site
USD 150,000 - 210,000
Infrastructure Engineer Lead
Infrastructure Engineer Lead

Aegis AI Security • New York (NY)

On-site
USD 180,000 - 280,000
Member of Technical Staff, AI Engineer San Francisco, CA
Member of Technical Staff, AI Engineer San Francisco, CA

Parameter • San Francisco (CA)

Hybrid
USD 180,000 - 240,000