AI Researcher, Agent Systems

Greylock Partners

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Greylock Partners is partnering with an early-stage AI company to explore AI agents in real company workflows and to build the systems needed to test their effectiveness.

You will work directly with the founders across research and engineering, focusing on agent environments, deployment, evaluation, and multimodal understanding. Model post-training is a possible future direction.

Qualifications

  • Experience with agent evaluation, reinforcement-learning environments, simulation, or synthetic workflows.
  • Experience with multimodal or video understanding to turn long recordings into useful representations.
  • Research engineering that connects experiments to reliable, usable products.

Responsibilities

  • Develop research approaches for understanding work from video, accessibility data, and other workflow context.
  • Build or improve simulated environments that let agents practice and be evaluated before live deployment.
  • Design evaluations that measure whether agents complete useful work reliably and know when to involve a person.
  • Research how multiple agents can coordinate, operate safely and improve through deployment feedback.
  • Build the software and experimental infrastructure needed to move promising ideas into working systems.
  • Help define which research problems matter most as the company learns from customers.

Skills

Agent systems
Reinforcement learning
Multi-agent coordination
Research engineering

Job description

Greylock is partnering with an early-stage AI company focused on a gap between model capability and adoption. AI can already change how work gets done, but bringing agents into a real organization remains slow, manual, and difficult to scale beyond individual pilots.

The company is building a repeatable path from an initial AI-agent deployment to broader organizational impact. Its goal is to make those agents correct, trustworthy, robust, and scalable—and to help companies build on each successful deployment.

Summary

This is a research role for someone who wants to work on AI agents in real company workflows and build the systems needed to test whether they work.

Depending on your strengths, you could help create simulated company environments where agents can practice before deployment, develop ways to evaluate and improve multiple agents operating inside an organization, or turn raw video into representations that capture tasks, decisions, and handoffs.

You'll work directly with the founders across research and engineering, with close feedback from real deployments. The company is not training models today; its current research is centered on agent environments, deployment, evaluation, and multimodal understanding. Model post‑training is a possible future direction.

What You'll Own
  • Develop research approaches for understanding work from video, accessibility data, and other workflow context
  • Build or improve simulated environments that let agents practice and be evaluated before live deployment
  • Design evaluations that measure whether agents complete useful work reliably and know when to involve a person
  • Research how multiple agents can coordinate, operate safely and improve through deployment feedback
  • Build the software and experimental infrastructure needed to move promising ideas into working systems
  • Help define which research problems matter most as the company learns from customers
What We're Looking For

You may bring depth in one or more of these areas:

  • Computer-use agents, tool use, planning, or multi-agent systems
  • Agent evaluation, reinforcement-learning environments, simulation, or synthetic workflows
  • Multimodal or video understanding, including ways to turn long recordings into useful representations
  • Research engineering that connects experiments to reliable, usable products

You should be able to formulate an open-ended problem, build and test a solution, and learn from where it fails in practice. Research publications or frontier-lab experience are valuable, but neither is a substitute for strong technical judgment and hands‑on execution.

People who have built digital twins or sim‑to‑real systems in robotics or physical AI may also find the company‑simulation problem compelling. The key question is whether you want to apply those skills to software‑based work inside organizations.

About Us

Greylock is an early‑stage investor in companies including Airbnb, LinkedIn, Dropbox, Workday, Cloudera, Facebook, Instagram, Roblox, Coinbase, and Palo Alto Networks. Learn more at greylock.com.

How We Work

Greylock's Core Talent team provides free candidate referrals and introductions to our active investments. This posting is for direct employment with one of our portfolio companies. We review every application and reach out directly when we believe there's a strong potential fit.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Engineer — AI Agent Security
Founding Engineer — AI Agent Security

Greylock Partners • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior Data Infrastructure Engineer
Senior Data Infrastructure Engineer

Greylock Partners • San Francisco (CA)

On-site
USD 180,000 - 260,000
Applied AI Engineer
Applied AI Engineer

Judgment Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Lead Product Manager (AI Dev Tooling)
Lead Product Manager (AI Dev Tooling)

Greylock Partners • San Francisco (CA)

On-site
USD 120,000 - 160,000
Applied AI Engineer
Applied AI Engineer

Day One Partners • New York (NY)

On-site
USD 120,000 - 210,000
Sr. Machine Learning Engineer, Physical AI
Sr. Machine Learning Engineer, Physical AI

Greylock Partners • New York (NY)

On-site
USD 140,000 - 230,000
GTM Lead, Artifact Collection
GTM Lead, Artifact Collection

Greylock Partners • New York (NY)

On-site
USD 110,000 - 170,000
Candidate referrals
Introductions to investments
Staff Research Engineer, Multi-Agent Scaling
Staff Research Engineer, Multi-Agent Scaling

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Founding Senior Applied AI Engineer
Founding Senior Applied AI Engineer

Goaly AI • Palo Alto (CA)

Hybrid
USD 180,000 - 240,000
Hybrid in Palo Alto
Visa sponsorship
Meals and snacks
Founding Research Engineer, Applied AI
Founding Research Engineer, Applied AI

Agentio Inc. • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
Flexible PTO
Health Coverage
Dental & Vision
+6