Founding AI Research Engineer

Socket.dev

San Francisco (CA)

On-site

USD 150,000 - 250,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
YC-backed
Founding team

Job summary

Klavis AI is hiring a founding engineer to build, test, debug, and ship software that powers frontier AI post-training data projects. You’ll work with the founders to create tasks, environments, and data pipelines using Python, TypeScript, Docker, and advanced AI tooling.

You should have extensive experience with LLMs, coding agents, and AI-assisted engineering workflows, and you’ll own core product areas from design to production within a founding team.

Qualifications

  • 3+ years of professional software engineering experience.
  • Extremely strong with LLMs, coding agents, and AI-assisted engineering workflows.
  • Proficient across Python, TypeScript, shell, Docker, Git, APIs, databases, and modern dev tools.
  • Strong taste for realistic long-horizon coding tasks, test design, agent workflows, evals, and edge cases.
  • Care deeply about correctness, verification, reproducibility, and data quality.
  • Move fast without accepting sloppy work.
  • Desire ownership and intensity of founding-stage work.

Responsibilities

  • Build high-quality long-horizon coding datasets for frontier AI post-training.
  • Create coding tasks, hidden tests, gold solutions, rubrics, dockerized environments, and agent trajectories.
  • Design realistic software engineering workflows across terminals, repos, APIs, databases, and developer tools.
  • Use LLMs and coding agents to accelerate engineering and data production.
  • Build infrastructure for data generation, environment orchestration, verification, evaluation, and QA.
  • Work with human experts to define valuable tasks for frontier post-training.
  • Turn customer needs into reliable, scalable data products.

Skills

LLMs & coding agents
Python
TypeScript
Shell
Docker
Git
APIs & databases
AI-assisted engineering
Product ownership

Tools

MCP servers
Eval systems
LLM orchestration tools
Claude Code / Codex

Job description

About Klavis AI

Klavis AI is building high-quality agentic and coding data for frontier AI post-training.

Frontier models are increasingly bottlenecked not just by compute, but by the quality of the coding environments, trajectories, rubrics, rewards, and verification data used to train them. We build that data layer: long-horizon coding tasks, terminal-based software engineering environments, hidden-test verification, expert rubrics, gold trajectories, dockerized environments, and agentic tool-use workflows ready for RL and SFT. We already work with multiple frontier AI labs on production coding and agentic data.

The Role

We’re hiring a founding engineer who is genuinely exceptional at using LLMs and coding agents to build, test, debug, and ship real software.

You should be the kind of engineer who can make Claude Code, Codex, MCP tools, shell environments, custom evals, Docker, and agent pipelines feel like an extension of your hands. We are not looking for someone who has only tried basic ChatGPT prompts or simple API wrappers. We are looking for someone who already uses AI agents to build, debug, refactor, test, and ship faster than traditional engineering teams.

You’ll work directly with the founders to build the systems and datasets that help frontier labs train better coding and tool-use agents. In your first 30 days, you’ll onboard into our internal systems, ship improvements to them, and personally produce high-quality tasks end-to-end. By the end of the first month, you should understand what makes a task valuable for frontier post-training. In 90 days, you’ll own a core product or infrastructure area from design to production. You’ll help define what “excellent” agentic and coding data means, build systems that scale production, and directly influence how frontier labs train their next-generation agents.

What you’ll work on

  • Build high-quality long-horizon coding datasets for frontier AI post-training
  • Create coding tasks, hidden tests, gold solutions, rubrics, dockerized environments, and agent trajectories
  • Design realistic software engineering workflows across terminals, repos, APIs, databases, and developer tools
  • Use LLMs and coding agents aggressively to accelerate engineering and data production
  • Build infrastructure for data generation, environment orchestration, verification, evaluation, and QA
  • Work with human experts to define difficult, realistic, and verifiable coding tasks
  • Design agentic tool-use workflows across SaaS apps, APIs, MCP servers, and external tools where needed
  • Turn ambiguous customer needs into reliable, scalable data products

What we're looking for

  • Have 3+ years of professional software engineering experience
  • Are extremely strong with LLMs, coding agents, and AI-assisted engineering workflows
  • Can build across Python, TypeScript, shell, Docker, Git, APIs, databases, and modern dev tools
  • Have strong taste for realistic long-horizon coding tasks, test design, agent workflows, evals, and edge cases
  • Care deeply about correctness, verification, reproducibility, and data quality
  • Move fast without accepting sloppy work
  • Want the ownership, ambiguity, and intensity of joining at the founding stage

Strong signals

  • You have created long-horizon coding benchmarks, hidden tests, task environments, coding-agent evals, or data pipelines
  • You have built custom coding-agent workflows, MCP servers, eval systems, or LLM orchestration tools
  • You use Claude Code, Codex, Cursor, or similar tools daily and deeply understand their failure modes
  • You have strong open-source, infra, systems, ML engineering, devtools, or competitive programming experience
  • You can show examples of agents helping you ship real software, not just demos

Why join

  • Work on a core bottleneck for frontier AI: post-training data quality
  • Build products already used by frontier AI labs
  • Join a YC-backed company at the founding stage
  • Work directly with technical founders with deep agentic AI, infra, and ML systems experience
  • Own important engineering and product decisions from day one
  • Help define how future AI coding and tool-use agents are trained

Compensation

$150K - $250K + Equity 0.50% - 1.00%

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding AI Research Engineer: Agentic Coding Systems
Founding AI Research Engineer: Agentic Coding Systems

Socket.dev • San Francisco (CA)

On-site
USD 150,000 - 250,000
Equity
YC-backed
Founding team
Member of Technical Staff
Member of Technical Staff

Collinear AI • Sunnyvale (CA)

On-site
USD 170,000 - 210,000
Competitive salary and equity packages
Health insurance
Backend Software Engineer - India
Backend Software Engineer - India

Reacher • United States

Remote
USD 30,000 - 60,000
High autonomy
Cutting-edge AI tools
Strong engineering culture
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
Founding Member of Technical Staff
Founding Member of Technical Staff

ROI-AI • San Francisco (CA)

On-site
USD 180,000 - 220,000
Founding AI Engineer
Founding AI Engineer

Bonfirevc • New York (NY)

On-site
USD 120,000 - 160,000
100% employer-paid health insurance premiums
Flexible PTO
Seed-stage equity grant
+2
Principal Applied AI Engineer
Principal Applied AI Engineer

ByLabs • Seattle (WA)

On-site
USD 150,000 - 190,000
ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Mercor • New York (NY)

On-site
USD 220,000 - 551,000
AI Engineer
AI Engineer

Valsoft Corporation • Northern (KY)

Hybrid
USD 120,000 - 180,000