Head of Research

talp

San Francisco (CA)

On-site

USD 180,000 - 320,000

Full time

13 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Talp, a simulation infrastructure company, is seeking an experienced leader to guide its scientific direction and turn research into working systems. You will oversee modeling, evaluation, and the infrastructure that supports experimentation and deployment.

This is a hands-on leadership role: design experiments, write code, inspect data, and build a research team in collaboration with engineering. Strong background in applied AI is essential.

Qualifications

  • Strong track record in applied AI or ML research with production impact.
  • Hands-on experience with language models and modern training techniques.
  • Solid Python skills and familiarity with PyTorch or JAX.
  • Ability to design and run rigorous experiments and benchmarks.
  • Experience building research tooling and ML infrastructure.
  • Leadership experience in research teams or programs.

Responsibilities

  • Set the research direction and align experiments with product goals.
  • Develop and improve models across data prep, adaptation, training, and fine-tuning.
  • Define benchmarks to measure behavioral accuracy and usefulness.
  • Build evaluation systems with traceable datasets and experiments.
  • Collaborate with engineering on data pipelines and model serving.
  • Grow and mentor a research team and establish rigorous standards.

Skills

Applied Research
Deep Models
Python
Statistical Judgment
Team Leadership
Research-to-Production
Autonomy

Tools

PyTorch
JAX

Job description

About Talp

Talp is simulation infrastructure for enterprise decisions. We build a company's customer base from its own behavioral data so decisions can be tested before they are made, and we already work with large enterprises across three regions.

We are the wind tunnel for decisions. What we install stays and sharpens every time it runs, which is why we call it infrastructure rather than a tool. Testing decisions this way will soon be as ordinary as version control.

We have raised $2.2 million to date, most recently at a $25 million valuation, from Formus Capital, Sunshine Lake Ventures, Aito Capital, the a16z Scout Fund, and a group of GPs and founders.

The Role

Everyone in this field renders the simulation. Nobody renders the proof. Confidence scores, calibration methods, benchmarks measured against human evidence... None of it has been published by anyone in this category, and the company that publishes it first will define how everybody else gets measured.

That is the center of this job.

You will lead our scientific direction and turn research into working systems. Your scope spans modeling, evaluation, and the infrastructure that supports experimentation and deployment.

This is a hands-on leadership role. You will design experiments, write code, inspect data, and work closely with engineering while building a research team. You should be comfortable asking difficult scientific questions and taking full responsibility for how they are answered.

Responsibilities

Set the research direction: Identify the most valuable questions, prioritize experiments, and connect scientific progress directly to product improvements.

Develop and improve models: Work across data preparation, model adaptation, training, and fine-tuning. Reproduce relevant research, challenge its assumptions, and test approaches against strong alternatives.

Define how quality is measured: Design benchmarks and experiments that assess behavioral accuracy, consistency, and usefulness. Compare simulations against human evidence while accounting for uncertainty, population variance, and data limitations.

Build reliable evaluation systems: Develop repeatable workflows for comparing models, investigating failures, and detecting regressions. Keep datasets, experiments, and results strictly traceable, protecting evaluation data from leaking into development.

Combine human and automated judgment: Create clear scoring criteria, review processes, and automated evaluators, then verify that those evaluators consistently agree with qualified human reviewers.

Make research practical to run: Partner with engineering on data pipelines, experiment tracking, training workflows, and model serving to maximize iteration speed, reliability, and compute efficiency.

Carry improvements into production: Follow promising results through implementation and deployment, verifying that their advantages hold outside the original experiment.

Build the research team and culture: Recruit and mentor researchers and engineers, establish clear standards for rigor, and communicate findings through technical writing, publications, and external collaborations.

Who You Are

Applied Research Record: A strong track record in applied AI or machine learning research, with evidence of taking ideas from experiments into working production systems.

Deep Model Experience: Hands‑on experience with language models, modern training/fine‑tuning techniques, and rigorous model evaluation.

Core Engineering Skills: Strong Python skills and deep familiarity with frameworks such as PyTorch or JAX.

Sound Statistical Judgment: You reason intuitively about sampling, uncertainty, and bias, and can tell immediately whether an apparent improvement is statistically meaningful.

Systems & Tooling Ownership: Experience building or owning research tooling, evaluation pipelines, or ML infrastructure that other engineers rely on.

Team Leadership: Experience leading researchers or substantial research programs, and the judgment to build a team whose strengths complement your own.

High Agency: The ability to drive technical work autonomously and make clear decisions even when the empirical evidence is incomplete.

Intellectual Honesty: Clear communication and the willingness to discard an approach the moment the data challenges it.

Typically five or more years of relevant AI or ML research experience, including at least one year leading research teams or substantial research programs. Relevant academic research counts toward this; we also welcome candidates with exceptional demonstrated impact on a shorter timeline.

Valuable Additional Experience

Behavioral science, computational social science, causal inference, or experimental design

Synthetic data generation, human feedback loops, annotation systems, or complex behavioral datasets

Multi‑agent systems and simulations involving interacting populations

Distributed training, GPU optimization, or efficient model serving

You do not need equal depth in every area. We look for deep research instincts, strong engineering ability, and the judgment to build a team with complementary strengths. A PhD is welcome but not required.

Our Process

We prioritize thoughtful conversations and clear examples of past work. Our hiring process is designed to help both sides align on mutual fit, working style, and expectations.

Reapplication Policy: To ensure a fair and thorough evaluation for all applicants, Talp observes a 90‑day waiting period before reconsidering candidates for the same role.

Talp is an equal opportunity workplace. We welcome applicants of every background and identity. If you need support or an accommodation at any point in the process, let us know and we will arrange it.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Research Scientist
Staff Research Scientist

Turing • United States

On-site
USD 250,000 - 400,000
Head of Applied AI Research & Systems
Head of Applied AI Research & Systems

talp • San Francisco (CA)

On-site
USD 180,000 - 320,000
Senior Research Scientist, STEM
Senior Research Scientist, STEM

Turing • United States

On-site
USD 250,000 - 350,000
Research Scientist, STEM
Research Scientist, STEM

Turing • United States

On-site
USD 150,000 - 300,000
Research - Member of Technical Staff
Research - Member of Technical Staff

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health & Wellness
Equity
Flexible time off
Senior Research Scientist – Frontier AI Benchmarking
Senior Research Scientist – Frontier AI Benchmarking

Foundation Capital • United States

On-site
USD 250,000 - 350,000
Member of Technical Staff — Research
Member of Technical Staff — Research

Collective Intuition, Inc. • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 250,000
Equity
Research Engineer
Research Engineer

Tessera Labs • New York (NY)

On-site
USD 200,000 - 300,000
Frontier AI Research Scientist - Benchmarks & Data
Frontier AI Research Scientist - Benchmarks & Data

Foundation Capital • United States

On-site
USD 250,000 - 400,000
Research Infrastructure - Member of Technical Staff
Research Infrastructure - Member of Technical Staff

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Equity grants
Health, dental, vision
Flexible time off