Applied AI Engineer

Soulside AI

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Remote-friendly culture
Founders access
Professional development budget
Conference attendance
Equity

Job summary

Soulside AI on-site in San Francisco is seeking an Applied AI Engineer to own the model layer that makes our documentation trustworthy. You will build post-training pipelines, fine-tune and deploy open-source models, and establish rigorous evaluation sets to measure clinical note quality and medical necessity.

Working at the intersection of applied ML and product, you will collaborate with clinical experts, optimize the end-to-end ML stack, and drive measurable improvements in our clinical

Qualifications

  • 3+ years in applied ML / AI engineering, or a Master's degree in a related field.
  • Hands-on experience taking LLM-based systems into production.
  • Experience with post-training / fine-tuning open-source models using SFT, LoRA/PEFT, or preference-based methods.
  • Experience serving or fine-tuning models on managed platforms such as Fireworks AI, Baseten, or Together AI.
  • Strong Python and familiarity with the modern ML tooling ecosystem (PyTorch, Hugging Face, etc.).
  • Solid grounding in prompt engineering and structured-output validation.

Responsibilities

  • Build post-training pipelines on open-source models—supervised fine-tuning, preference optimization (DPO/RLHF), LoRA/adapters, and distillation—for domain-specific clinical tasks.
  • Fine-tune, deploy, and serve models across managed inference and fine-tuning platforms such as Fireworks AI, Baseten, and Together AI, and make pragmatic build-vs-buy calls on where each workload should run.
  • Design and maintain rigorous evaluation sets for high-stakes tasks like clinical reasoning and AI note generation—defining metrics, curating gold-standard data, and building automated and human-in-the-loop eval harnesses.
  • Turn eval results into a fast, trustworthy iteration loop: catch regressions before they ship, and quantify the impact of every model or prompt change.
  • Optimize the full LLM pipeline—prompting, retrieval, structured output validation, latency, and cost.
  • Partner with clinical experts to translate documentation and compliance requirements into model behavior and evaluation criteria.
  • Monitor models in production for quality, drift, and failure modes, and close the loop back into training data and evals.

Skills

Python
Applied ML
LLM deployment
Prompt engineering
Hugging Face

Education

Master's degree in a related field

Tools

PyTorch
Hugging Face
LoRA/PEFT
SFT
Baseten
Fireworks AI
Together AI

Job description

Soulside AI · US On-Site · Reports to the CTO

About Soulside

Soulside AI is the specialist AI platform for behavioral health documentation and compliance. We generate audit-ready clinical documentation across individual and group sessions, virtual and in-person care, admissions, and treatment planning—and we embed real-time chart audits and payer-aligned compliance checks into everyday workflows. The result is immediate and measurable: higher-quality charts, stronger medical necessity, and hours given back to clinicians every week.

We're backed by Counterpart Ventures, GreyMatter Capital, and One Mind, and we're a UCSF Rosenman Institute and One Mind Accelerator company. We've reached strong product-market fit and are scaling fast.

The Role

We're looking for an Applied AI Engineer to own the model layer that makes Soulside's documentation trustworthy. In behavioral health, a note isn't just text—it has to be clinically sound, defensible for medical necessity, and safe. Your job is to build the post-training pipelines and evaluation systems that get our models there, and keep them there as we scale.

This is a hands‑on role for someone who lives at the intersection of applied ML and product. You'll fine-tune and adapt open-source models, stand up the infrastructure to serve them, and build the rigorous evaluation sets that tell us—objectively—whether a change made the product better or worse.

Why This Role Matters

Accuracy Isn't Optional: In behavioral health, a wrong or unsupported note has real clinical and financial consequences. The pipelines and evals you build are what let us ship model changes with confidence.

Own the Model Layer: You'll define how we post-train, evaluate, and deploy models end‑to‑end—not inherit someone else's stack.

Direct Clinical Impact: Every improvement in clinical reasoning or note quality directly reduces documentation burden and strengthens the charts clinicians and payers rely on.

What You'll Do

Build post‑training pipelines on open-source models—supervised fine‑tuning, preference optimization (DPO/RLHF), LoRA/adapters, and distillation—for domain‑specific clinical tasks.

Fine‑tune, deploy, and serve models across managed inference and fine‑tuning platforms such as Fireworks AI, Baseten, and Together AI, and make pragmatic build‑vs‑buy calls on where each workload should run.

Design and maintain rigorous evaluation sets for high‑stakes tasks like clinical reasoning and AI note generation—defining metrics, curating gold‑standard data, and building automated and human‑in‑the‑loop eval harnesses.

Turn eval results into a fast, trustworthy iteration loop: catch regressions before they ship, and quantify the impact of every model or prompt change.

Optimize the full LLM pipeline—prompting, retrieval, structured output validation, latency, and cost.

Partner with clinical experts to translate documentation and compliance requirements into model behavior and evaluation criteria.

Monitor models in production for quality, drift, and failure modes, and close the loop back into training data and evals.

What We're Looking For

3+ years in applied ML / AI engineering, or a Master's degree in a related field, with hands‑on experience taking LLM-based systems into production.

Practical experience with post‑training / fine‑tuning open‑source models (e.g., Llama, Qwen, Mistral) using SFT, LoRA/PEFT, or preference‑based methods.

Experience serving or fine‑tuning models on managed platforms such as Fireworks AI, Baseten, or Together AI (or comparable inference/training infra).

Demonstrated ability to build evaluation frameworks for LLM tasks—you think in terms of measurable quality, not vibes.

Strong Python and familiarity with the modern ML tooling ecosystem (PyTorch, Hugging Face, etc.).

Solid grounding in prompt engineering and structured‑output validation.

Ability to thrive in a fast‑paced, remote startup and communicate clearly with technical and clinical teammates.

We're willing to sponsor visas, including H‑1B and O‑1, for the right candidate.

Bonus Points

Experience with healthcare, clinical NLP, or other high‑stakes / regulated domains.

Familiarity with HIPAA and handling sensitive clinical data.

RAG systems, retrieval quality tuning, or long‑context document workflows.

Experience with LLM observability, monitoring, and drift detection in production.

Data pipeline and labeling workflow experience for curating high‑quality training and eval sets.

Open‑source contributions in the ML/LLM ecosystem.

What We Offer

Salary range of $150,000–$200,000, plus equity with significant upside potential as a founding team member

Comprehensive health, dental, and vision insurance

Flexible, remote‑first culture

Direct access to founders and influence on technical direction

Professional development budget and conference attendance

The chance to build AI that measurably improves mental health care at scale

548 Market St, PMB 48695

San Francisco, California 94104-5401 US

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Engineer
Applied AI Engineer

Sarah Smith Fund • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Remote-first culture
Direct access to founders
Professional development budget
Director of Machine Learning (Healthcare AI)
Director of Machine Learning (Healthcare AI)

Nxt Level • United States

Hybrid
USD 150,000 - 200,000
Competitive salary
Meaningful equity
Direct line to CEO
+1
Applied AI Engineer
Applied AI Engineer

Norbert Health • New York (NY)

On-site
USD 100,000 - 150,000
Equity participation
Competitive salary
High autonomy and technical ownership
Applied AI Engineer — Clinical ML for Behavioral Health
Applied AI Engineer — Clinical ML for Behavioral Health

Sarah Smith Fund • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Remote-first culture
Direct access to founders
Professional development budget
Applied AI Engineer — Remote Behavioral Health ML
Applied AI Engineer — Remote Behavioral Health ML

Soulside AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Health insurance
Dental insurance
Vision insurance
+5
Senior AI Engineer
Senior AI Engineer

Citizen Health • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Competitive salary + equity package
Comprehensive health, dental, and vision insurance
Unlimited paid time off and generous parental leave
+1
Founding AI Engineer
Founding AI Engineer

Nolla Health • New York (NY)

On-site
USD 160,000 - 220,000
Health Insurance
Flexible unlimited vacation
Meal stipends
+1
Senior AI Engineer
Senior AI Engineer

Weave • San Francisco (CA)

Hybrid
USD 175,000 - 215,000
Comprehensive health, dental and vision insurance
Generous PTO and parental leave
Career development opportunities
+1
Senior AI Engineer
Senior AI Engineer

Citizen Health • San Francisco (CA)

On-site
USD 150,000 - 300,000
Competitive salary + equity package
Comprehensive health, dental, and vision insurance
Unlimited paid time off
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Ambience Healthcare • San Francisco (CA)

Hybrid
USD 225,000 - 300,000
Medical, dental, vision
401(k) with company match
Remote-friendly culture
+2