Member of Technical Staff - Post Training, Applied (Vision)

Liquid AI

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive salary with equity
100% health premiums for employees and dependents
401(k) matching
Unlimited PTO

Job summary

A cutting-edge AI firm in San Francisco seeks a VLM Post-Training Owner to lead enterprise engagements and enhance vision-language models. The role combines project ownership with technical execution, ensuring quality data generation and customer satisfaction in AI solutions. Ideal candidates have hands-on experience in multimodal AI, strong communication skills, and the ability to manage complex projects. Competitive salary and comprehensive health benefits included.

Qualifications

  • Hands-on experience with data generation and evaluation for VLM or multimodal post-training.
  • Experience training or fine-tuning vision-language models.
  • Strong intuition for visual data quality and annotation design.
  • Familiarity with vision encoders, image‑text architectures, and how visual representations interact with language model backbones.

Responsibilities

  • Act as the technical owner for enterprise VLM post-training engagements.
  • Translate customer requirements into multimodal post-training specifications.
  • Design and execute visual data generation and quality assessment processes.
  • Run supervised fine-tuning, preference alignment, and reinforcement learning workflows for vision-language models.
  • Design task-specific evaluations for visual understanding, grounding, OCR, document parsing, and other multimodal capabilities. Interpret results and feed learnings back into core post-training pipelines.

Skills

Ownership of projects
End-to-end thinking
Pragmatism
Clear communication

Job description

About Liquid AI

Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there.

The Opportunity

This is a rare chance to sit at the intersection of frontier vision-language models and real-world deployment. You'll own applied post-training work for VLMs end-to-end for some of the world's largest enterprises, while still contributing directly to Liquid's core multimodal model development.

Unlike most roles that force a trade-off between customer impact and foundational work, this role gives you both: deep ownership over how vision-language models are adapted, evaluated, and shipped, and a direct line into the evolution of Liquid's multimodal post-training stack.

If you care about visual understanding, data quality, evaluation, and making VLMs actually work in production, this is a chance to shape how applied multimodal AI is done at a foundation model company.

What We're Looking For

We need someone who:

  • Takes ownership: Owns VLM post-training projects end-to-end, from customer requirements through delivery and evaluation.

  • Thinks end-to-end: Can reason across visual data curation, training, alignment, and evaluation as a single system.

  • Is pragmatic: Optimizes for model quality and customer outcomes over publications or theory.

  • Communicates clearly: Can translate between customer needs and internal technical teams, and push back when needed.

The Work
  • Act as the technical owner for enterprise customer VLM post-training engagements.

  • Translate customer requirements into concrete multimodal post-training specifications and workflows.

  • Design and execute visual data generation, filtering, and quality assessment processes, including image-text pair curation, annotation pipelines, and synthetic data generation for visual tasks.

  • Run supervised fine-tuning, preference alignment, and reinforcement learning workflows for vision-language models.

  • Design task-specific evaluations for visual understanding, grounding, OCR, document parsing, and other multimodal capabilities. Interpret results and feed learnings back into core post-training pipelines.

Desired Experience

Must-have:

  • Hands‑on experience with data generation and evaluation for VLM or multimodal post‑training.

  • Experience training or fine‑tuning vision‑language models using SFT, preference alignment, and/or RL.

  • Strong intuition for visual data quality, annotation design, and multimodal evaluation.

  • Familiarity with vision encoders, image‑text architectures, and how visual representations interact with language model backbones.

Nice‑to‑have:

  • Experience with visual grounding, document understanding, OCR, or video understanding tasks.

  • Experience contributing to shared or general‑purpose multimodal post‑training infrastructure.

  • Prior exposure to customer‑facing or applied ML delivery environments.

  • Familiarity with alignment or RL techniques beyond basic supervised fine‑tuning in the multimodal setting.

What Success Looks Like (Year One)
  • Independently owns and delivers enterprise VLM post‑training projects with minimal oversight.

  • Is trusted by customers as the technical owner, demonstrating strong judgment and delivery quality on multimodal workloads.

  • Has made durable contributions to Liquid's general‑purpose multimodal post‑training pipelines by feeding applied learnings back into baseline model development.

What We Offer
  • Real ML work: You will fine‑tune vision‑language models, generate multimodal data, and ship solutions, not configure API calls. Your work feeds directly back into our core model development.

  • Compensation: Competitive base salary with equity in a unicorn‑stage company.

  • Health: We pay 100% of medical, dental, and vision premiums for employees and dependents.

  • Financial: 401(k) matching up to 4% of base pay.

  • Time Off: Unlimited PTO plus company‑wide Refill Days throughout the year.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Post Training, Applied (Text)
Member of Technical Staff - Post Training, Applied (Text)

Liquid AI • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive base salary with equity
100% paid health, dental, and vision premiums
401(k) matching up to 4%
+1
Member of Technical Staff - Post Training, Applied (Audio)
Member of Technical Staff - Post Training, Applied (Audio)

Liquid AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive base salary
Equity in a unicorn-stage company
100% paid health premiums
+2
Research Engineer, Visual Knowledge Work
Research Engineer, Visual Knowledge Work

Anthropic • New York (NY)

Hybrid
USD 350,000 - 850,000
Generous vacation
Parental leave
Flexible working hours
ML Researcher - Posttraining
ML Researcher - Posttraining

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 230,000
Competitive salary
Equity package
Health & dental insurance
+5
ML Researcher - Posttraining
ML Researcher - Posttraining

Krea • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 260,000
Health & dental insurance
Flexible PTO
401k with company match
+3
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Clearview AI • United States

On-site
USD 180,000 - 250,000
Medical plans
Dental plans
Vision plans
+1
Research Engineer, Visual Knowledge Work
Research Engineer, Visual Knowledge Work

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Vision-Language Models (VLMs)
Vision-Language Models (VLMs)

TalentOla • Waukesha (WI)

On-site
USD 120,000 - 150,000
Solutions Architect
Solutions Architect

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive base salary
Equity in a unicorn-stage company
100% medical, dental, and vision premiums
+2
Research Engineer, Multimodal Data
Research Engineer, Multimodal Data

Eventual • San Francisco (CA)

On-site
USD 120,000 - 150,000
Catered lunches and dinners
Commuter benefit
Health, vision, and dental coverage
+2