Member of Technical Staff, Applied Research

LlamaIndex Inc.

San Francisco (CA)

On-site

USD 140,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

LlamaIndex is hiring an AI Research Engineer to join the document understanding team in San Francisco. You will work at the intersection of applied research and engineering, prototyping quickly and turning promising approaches into production systems used by customers.

You’ll develop vision‑language models, build data pipelines, and evaluate post‑training/fine‑tuning to improve accuracy, latency, and cost. Collaborative work with engineering and customers is expected.

Qualifications

  • 3–7 years of experience in machine learning engineering, applied research, or research engineering.
  • Strong ML foundation, including hands‑on experience benchmarking and training models.
  • Strong Python skills and comfort with modern ML tooling, especially PyTorch.
  • Experience with computer vision, vision‑language models, NLP, document AI, OCR, extraction, or agentic AI systems.
  • Ability to build experiments, evaluate results, and iterate quickly toward measurable performance improvements.
  • Strong engineering judgment and ability to write clean, production‑quality code.
  • Comfort working in a fast‑paced startup environment with high ownership and limited structure.
  • Adaptable, scrappy, and self‑directed — someone who can figure things out without waiting to be told.
  • Strong technical writing and communication skills.

Responsibilities

  • Develop and train vision-language models for document processing and document understanding.
  • Build data pipelines for data curation, synthetic data generation, labeling, and benchmark creation.
  • Evaluate base models and perform post-training or fine-tuning to hit specific performance targets.
  • Improve model accuracy, latency, and cost-effectiveness across real-world document workflows.
  • Design and maintain benchmarks to measure extraction quality, layout understanding, OCR performance, reasoning accuracy, and end-to-end system reliability.
  • Work with messy real-world documents, including PDFs, scanned documents, tables, charts, forms, and multi-page enterprise documents.
  • Collaborate with engineering to move successful research prototypes into production.
  • Work directly with customers when needed to translate product requirements into benchmarks, experiments, and model improvements.
  • Stay close to the latest research in vision-language models, document AI, post-training, synthetic data, and agentic systems.
  • Use modern AI coding workflows and tools to move quickly.

Skills

Python
PyTorch
ML engineering
Vision-language

Tools

vLLM
Pydantic
uv
ruff
mypy
Claude Code
Cursor

Job description

Join us and help shape the future of AI by defining the narrative around document understanding.

About the Role

We are looking for an AI Research Engineer to join our document understanding team.

This role is ideal for someone who sits between applied research and strong engineering. You will work on vision-language models, document processing, data curation, synthetic data generation, benchmarking, training, fine-tuning, and post-training. The goal is simple: make our document AI systems more accurate, faster, and more cost-effective in production.

You should be excited by frontier AI work, but equally motivated by practical product impact. This is not a pure research role where ideas stay in papers. You will be expected to prototype quickly, evaluate rigorously, and help turn promising approaches into production systems used by customers.

What You’ll Do
  • Develop and train vision-language models for document processing and document understanding.

  • Build data pipelines for data curation, synthetic data generation, labeling, and benchmark creation.

  • Evaluate base models and perform post-training or fine-tuning to hit specific performance targets.

  • Improve model accuracy, latency, and cost-effectiveness across real-world document workflows.

  • Design and maintain benchmarks to measure extraction quality, layout understanding, OCR performance, reasoning accuracy, and end-to-end system reliability.

  • Work with messy real-world documents, including PDFs, scanned documents, tables, charts, forms, and multi-page enterprise documents.

  • Collaborate with engineering to move successful research prototypes into production.

  • Work directly with customers when needed to translate product requirements into benchmarks, experiments, and model improvements.

  • Stay close to the latest research in vision-language models, document AI, post-training, synthetic data, and agentic systems.

  • Use modern AI coding workflows and tools to move quickly.

What We’re Looking For
  • 3–7 years of experience in machine learning engineering, applied research, or research engineering.

  • Strong ML foundation, including hands‑on experience benchmarking and training models.

  • Strong Python skills and comfort with modern ML tooling, especially PyTorch.

  • Experience with computer vision, vision‑language models, NLP, document AI, OCR, extraction, or agentic AI systems.

  • Ability to build experiments, evaluate results, and iterate quickly toward measurable performance improvements.

  • Strong engineering judgment and ability to write clean, production‑quality code.

  • Comfort working in a fast‑paced startup environment with high ownership and limited structure.

  • Adaptable, scrappy, and self‑directed — someone who can figure things out without waiting to be told.

  • Strong technical writing and communication skills.

Nice to Have
  • Prior startup experience, especially at an early‑stage or high‑growth AI company.

  • Experience as a founder or early startup engineer.

  • Experience building or improving document processing systems.

  • Experience with synthetic data generation, post‑training, fine‑tuning, or benchmark design.

  • Familiarity with tools such as vLLM, Pydantic, uv, ruff, mypy, Claude Code, Cursor, or similar modern AI engineering workflows.

  • Experience with open‑source AI infrastructure or developer tools.

Who You’ll Work With

You will work closely with the CTO and the document understanding team, partnering across research, engineering, product, and customer‑facing teams.

Why Join LlamaIndex
  • Work on a core AI infrastructure problem: making complex documents understandable and actionable for AI systems.

  • Build production systems at the frontier of vision‑language models and document AI.

  • Join a fast‑growing startup with strong open‑source adoption and commercial traction.

  • Work directly with technical founders and a highly ambitious engineering team.

  • Have real ownership over model quality, product capability, and technical direction.

Pursuant to the SanFrancisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

LlamaIndex does not accept unsolicited agency resumes. Please do not forward resumes to our jobs alias, employees, or any other organization location. LlamaIndex is not responsible for any fees related to unsolicited resumes.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Applied Research
Member of Technical Staff, Applied Research

LlamaIndex, Inc. • San Francisco (CA)

On-site
USD 180,000 - 250,000
Applied AI Research Engineer - Document Understanding
Applied AI Research Engineer - Document Understanding

LlamaIndex, Inc. • San Francisco (CA)

On-site
USD 180,000 - 250,000
Applied AI Research Engineer - Document Understanding
Applied AI Research Engineer - Document Understanding

LlamaIndex Inc. • San Francisco (CA)

On-site
USD 140,000 - 230,000
Developer Evangelist, Growth
Developer Evangelist, Growth

LlamaIndex Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 130,000 - 190,000
Developer Evangelist, Technical Content
Developer Evangelist, Technical Content

LlamaIndex Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
03 Member of Technical Staff, Backend San Francisco
03 Member of Technical Staff, Backend San Francisco

LlamaIndex, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive base salary and equity
Medical/dental/vision coverage for you
Unlimited paid time off
+1
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

LlamaIndex, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 275,000
Competitive base salary and equity
Comprehensive medical/dental/vision
Unlimited paid time off
+1
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

LlamaIndex • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive base salary and equity compensation
Comprehensive benefits coverage
Unlimited paid time off
+1
Developer Evangelist, Technical Content
Developer Evangelist, Technical Content

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Backend Engineer
Senior Backend Engineer

SupportFinity™ • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Competitive compensation
Equity compensation
Daily catered lunch