Research Engineer, LangSmith Engine

Neura Market

New York (NY)

On-site

USD 180,000 - 260,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LangChain's LangSmith Engine team is seeking an experienced research engineer to improve the Engine agent's capabilities and efficiency. You will study real agent failures, build benchmarks, run experiments across models, prompting, context, tools, and orchestration to drive measurable improvements in production.

This role requires hands-on experience with LLMs and AI agents, strong software engineering, and the ability to translate research ideas into production impact, cost-aware and scalable

Qualifications

  • 4+ years of experience in ML/AI research or closely related field.
  • Hands-on experience working with LLMs and AI agents, analyzing model behavior and improving production performance.
  • Strong ability to design benchmarks, evaluations, and experiments for AI/ML systems.
  • Proven track record of turning research ideas into measurable production impact.
  • Strong software engineering skills with end-to-end delivery.
  • Excellent judgment and ability to work through ambiguity.

Responsibilities

  • Build and maintain benchmarks and evaluations that measure the quality and efficiency of Engine agents on real-world tasks.
  • Design and run experiments to improve agent performance across models, prompting, context, tools, orchestration, and agent strategies.
  • Explore and implement post-training and fine-tuning techniques when they can meaningfully improve agent capabilities, quality, or cost.
  • Turn successful experiments into production improvements, working closely with engineers and researchers to measure impact and prevent regressions.
  • Help define the ML roadmap and technical direction for improving Engine agents, and mentor other engineers through strong technical leadership.

Skills

ML/AI research
LLMs & AI agents
Benchmarks & experiments
Software engineering
Research judgment

Education

Master's or Ph.D. in relevant field

Job description

About Us

At LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. We began as widely adopted open-source tools and have grown to also offer a platform for building, evaluating, deploying, and operating agents at scale.

With 125M raised at Series B from IVP, Sequoia, Benchmark, CapitalG, and Sapphire Ventures, we’re at a stage where we’re continuing to develop new products, growth is accelerating, and all team members have meaningful impact on what we build and how we work together. LangChain is a place where your contributions can shape how this technology shows up in the real world.

Today, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. We have 100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of the Fortune 500 overall), including teams at Klarna, Clay, Coinbase, Workday, Lyft, Cloudflare, Harvey, Rippling, Vanta, LinkedIn, Monday.com, Nvidia, and Bridgewater.

About the Team:

The LangSmith Engine team is building a proactive agent engineer that analyzes production traces, identifies important failures, recommends and writes fixes, and helps prevent those issues from coming back. We’re building agents that can understand complex software systems and continuously improve the quality of other AI agents.

About the Role:

We’re looking for an experienced research engineer to help make the Engine agent more capable and more efficient.

You’ll study real agent failures, build benchmarks that capture what matters, run experiments to improve performance, and turn successful ideas into production. This may include prompting and agent-harness improvements, model selection, fine-tuning and post-training custom models. The focus is on measurable improvements to the overall agent.

This role also requires a understanding of production engineering and system-level tradeoffs. Engine is a production system, so improving an agent is not just about maximizing benchmark performance—it also means understanding the impact on cost, latency, reliability, and scalability. You’ll work in the same team with production engineers to design, test, and ship improvements that work reliably in real-world environments.

Location: SF and NYC

What You’ll Do:
  • Build and maintain benchmarks and evaluations that measure the quality and efficiency of Engine agents on real-world tasks.
  • Design and run experiments to improve agent performance across models, prompting, context, tools, orchestration, and agent strategies.
  • Explore and implement post-training and fine-tuning techniques when they can meaningfully improve agent capabilities, quality, or cost.
  • Turn successful experiments into production improvements, working closely with engineers and researchers to measure impact and prevent regressions.
  • Help define the ML roadmap and technical direction for improving Engine agents, and mentor other engineers through strong technical leadership.
What You’ll Bring:
  • 4+ years of experience in ML/AI research, or a closely related field.
  • Master’s or Ph.D. in a relevant scientific field.
  • Hands-on experience working with LLMs and AI agents, including analyzing model behavior and improving real-world performance
  • Strong experience designing benchmarks, evaluations, and experiments for AI/ML systems; you know how to tell whether a change actually made an agent better.
  • Strong software engineering skills, with a track record of taking ideas from research prototype to measurable production impact.
  • You have maximum agency and strong research judgment: you can identify high-impact problems, work through ambiguity, move quickly, and communicate your findings clearly.
Nice to Have:
  • Ph.D. in Machine Learning, Computer Science or Physics.
  • Hands on experience with LLM-as-a-judge, automated graders, synthetic data generation, or human evaluation.
  • Hands on experience with reinforcement learning, preference optimization, SFT, RLHF/RLAIF, or other post-training techniques for LLMs.
  • Experience optimizing LLM agents for cost, latency, or task efficiency on productions
  • Experience with model serving, inference optimization, distributed systems, or GPU infrastructure.
Compensation & Benefits

We offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Benefits include things like medical, dental, and vision coverage, flexible vacation, a 401(k) plan, and life insurance. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.

Compensation Philosophy:

We offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.

Benefits

Benefits include medical, dental, and vision coverage, flexible vacation, a 401(k) plan, meals on in-office days in the US and more.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, LangSmith Engine
Research Engineer, LangSmith Engine

AI Chopping Block • New York (NY)

On-site
USD 170,000 - 230,000
Research Engineer, LangSmith Engine
Research Engineer, LangSmith Engine

LangChain • New York (NY)

On-site
USD 180,000 - 260,000
Equity
Medical, dental, vision coverage
401(k)
Research Engineer, LangSmith Engine
Research Engineer, LangSmith Engine

LangChain, Inc. • New York (NY)

On-site
USD 150,000 - 230,000
Competitive compensation
Equity
Medical, dental, vision coverage
+1
Production AI Agent Engineer — Pre‑Sales & Deployment
Production AI Agent Engineer — Pre‑Sales & Deployment

LangChain • San Francisco (CA)

On-site
USD 165,000 - 315,000
Deployed Engineer (Bay Area)
Deployed Engineer (Bay Area)

LangChain • San Francisco (CA)

On-site
USD 165,000 - 315,000
Medical, dental, and vision coverage
401(k) plan
Meals on in-office days in the US
Deployed Engineer (Early Career-NYC)
Deployed Engineer (Early Career-NYC)

LangChain • New York (NY)

On-site
USD 155,000 - 165,000
Medical benefits
401(k) plan
Office meals
+1
Deployed Engineer (Early Career-NYC)
Deployed Engineer (Early Career-NYC)

Neura Market • New York (NY)

On-site
USD 155,000 - 165,000
Deployed Engineer (Early Career-NYC)
Deployed Engineer (Early Career-NYC)

LangChain, Inc. • New York (NY)

On-site
USD 155,000 - 165,000
Medical coverage
Flexible vacation
401(k) plan
+1
Deployed Engineer (Early Career- SF/NY)
Deployed Engineer (Early Career- SF/NY)

LangChain, Inc. • San Francisco (CA)

On-site
USD 155,000 - 165,000
Medical benefits
401(k) plan
Meals on in-office days
Principal Software Engineer, AI Observability & Evals Platform
Principal Software Engineer, AI Observability & Evals Platform

LangChain • Cambridge (MA)

On-site
USD 230,000 - 270,000
Medical, dental, and vision coverage
401(k) plan with company match
Meals on in-office days (US)