Member of Technical Staff, Applied Research

LlamaIndex, Inc.

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

13 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

LlamaIndex, Inc. is seeking an AI Research Engineer to join the document understanding team in San Francisco. You will work between research and engineering to develop vision-language models for document processing, build data pipelines, and benchmark systems used by customers.

You will prototype quickly, evaluate rigorously, and help translate promising approaches into production-ready solutions that improve accuracy, latency, and cost in real-world workflows.

Qualifications

  • 3–7 years of ML engineering or research experience.
  • Strong Python skills and PyTorch experience.
  • Experience with CV, vision-language, NLP, or OCR.
  • Ability to build experiments and iterate quickly.
  • Production-quality code and good engineering judgment.

Responsibilities

  • Develop and train vision-language models for document processing.
  • Build data pipelines for data curation, labeling, and benchmarks.
  • Evaluate models and perform post-training or fine-tuning.
  • Improve accuracy, latency, and cost in production.
  • Collaborate to move research prototypes into production.
  • Work with customers to translate requirements into benchmarks.

Skills

Python
PyTorch
Vision-language
NLP
Experimentation
Engineering judgment
Production code

Tools

vLLM
Pydantic
uv
ruff
mypy
Claude Code
Cursor

Job description

Join us and help shape the future of AI by defining the narrative around document understanding.

About The Role

We are looking for an AI Research Engineer to join our document understanding team. This role is ideal for someone who sits between applied research and strong engineering. You will work on vision-language models, document processing, data curation, synthetic data generation, benchmarking, training, fine-tuning, and post-training. The goal is simple: make our document AI systems more accurate, faster, and more cost-effective in production. You should be excited by frontier AI work, but equally motivated by practical product impact. This is not a pure research role where ideas stay in papers. You will be expected to prototype quickly, evaluate rigorously, and help turn promising approaches into production systems used by customers.

What You’ll Do
  • Develop and train vision-language models for document processing and document understanding.
  • Build data pipelines for data curation, synthetic data generation, labeling, and benchmark creation.
  • Evaluate base models and perform post-training or fine-tuning to hit specific performance targets.
  • Improve model accuracy, latency, and cost-effectiveness across real-world document workflows.
  • Design and maintain benchmarks to measure extraction quality, layout understanding, OCR performance, reasoning accuracy, and end-to-end system reliability.
  • Work with messy real-world documents, including PDFs, scanned documents, tables, charts, forms, and multi-page enterprise documents.
  • Collaborate with engineering to move successful research prototypes into production.
  • Work directly with customers when needed to translate product requirements into benchmarks, experiments, and model improvements.
  • Stay close to the latest research in vision-language models, document AI, post-training, synthetic data, and agentic systems.
  • Use modern AI coding workflows and tools to move quickly.
What We’re Looking For
  • 3-7 years of experience in machine learning engineering, applied research, or research engineering.
  • Strong ML foundation, including hands-on experience benchmarking and training models.
  • Strong Python skills and comfort with modern ML tooling, especially PyTorch.
  • Experience with computer vision, vision-language models, NLP, document AI, OCR, extraction, or agentic AI systems.
  • Ability to build experiments, evaluate results, and iterate quickly toward measurable performance improvements.
  • Strong engineering judgment and ability to write clean, production-quality code.
  • Comfort working in a fast-paced startup environment with high ownership and limited structure.
  • Adaptable, scrappy, and self-directed - someone who can figure things out without waiting to be told.
  • Strong technical writing and communication skills.
Nice to Have
  • Prior startup experience, especially at an early-stage or high-growth AI company.
  • Experience as a founder or early startup engineer.
  • Experience building or improving document processing systems.
  • Experience with synthetic data generation, post-training, fine-tuning, or benchmark design.
  • Familiarity with tools such as vLLM, Pydantic, uv, ruff, mypy, Claude Code, Cursor, or similar modern AI engineering workflows.
  • Experience with open-source AI infrastructure or developer tools.
Who You’ll Work With

You will work closely with the CTO and the document understanding team, partnering across research, engineering, product, and customer-facing teams.

Why Join LlamaIndex
  • Work on a core AI infrastructure problem: making complex documents understandable and actionable for AI systems.
  • Build production systems at the frontier of vision-language models and document AI.
  • Join a fast-growing startup with strong open-source adoption and commercial traction.
  • Work directly with technical founders and a highly ambitious engineering team.
  • Have real ownership over model quality, product capability, and technical direction.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Compensation Range: $180K - $250K

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Applied Research
Member of Technical Staff, Applied Research

LlamaIndex Inc. • San Francisco (CA)

On-site
USD 140,000 - 230,000
Applied AI Research Engineer - Document Understanding
Applied AI Research Engineer - Document Understanding

LlamaIndex, Inc. • San Francisco (CA)

On-site
USD 180,000 - 250,000
Applied AI Research Engineer - Document Understanding
Applied AI Research Engineer - Document Understanding

LlamaIndex Inc. • San Francisco (CA)

On-site
USD 140,000 - 230,000
03 Member of Technical Staff, Backend San Francisco
03 Member of Technical Staff, Backend San Francisco

LlamaIndex, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive base salary and equity
Medical/dental/vision coverage for you
Unlimited paid time off
+1
Developer Evangelist, Growth
Developer Evangelist, Growth

LlamaIndex • San Francisco (CA)

Hybrid
USD 130,000 - 250,000
AI Content Engineer
AI Content Engineer

LlamaIndex • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

LlamaIndex, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 275,000
Competitive base salary and equity
Comprehensive medical/dental/vision
Unlimited paid time off
+1
Developer Evangelist, Growth
Developer Evangelist, Growth

LlamaIndex Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 130,000 - 190,000
Senior Backend Engineer
Senior Backend Engineer

SupportFinity™ • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Competitive compensation
Equity compensation
Daily catered lunch
Developer Evangelist, Technical Content
Developer Evangelist, Technical Content

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000