Get more replies from employers
Send a job-specific resume in minutes.
Soulside AI, a specialist AI platform for behavioral health documentation and compliance, seeks an Applied AI Engineer to own the model layer that makes our documentation trustworthy. You will build post-training pipelines and evaluation systems to keep models accurate and defensible at scale.
This is a hands-on role at the intersection of applied ML and product. You’ll fine-tune and adapt open-source models, stand up infrastructure to serve them, and develop rigorous eval sets to measure impact
Soulside AI · US On-Site · Reports to the CTO
Soulside AI is the specialist AI platform for behavioral health documentation and compliance. We generate audit-ready clinical documentation across individual and group sessions, virtual and in-person care, admissions, and treatment planning-and we embed real-time chart audits and payer-aligned compliance checks into everyday workflows. The result is immediate and measurable: higher-quality charts, stronger medical necessity, and hours given back to clinicians every week.
We're backed by Counterpart Ventures, GreyMatter Capital, and One Mind, and we're a UCSF Rosenman Institute and One Mind Accelerator company. We've reached strong product-market fit and are scaling fast.
We're looking for an Applied AI Engineer to own the model layer that makes Soulside's documentation trustworthy. In behavioral health, a note isn't just text—it has to be clinically sound, defensible for medical necessity, and safe. Your job is to build the post-training pipelines and evaluation systems that get our models there, and keep them there as we scale.
This is a hands-on role for someone who lives at the intersection of applied ML and product. You'll fine-tune and adapt open-source models, stand up the infrastructure to serve them, and build the rigorous evaluation sets that tell us—objectively—whether a change made the product better or worse.
Accuracy Isn't Optional: In behavioral health, a wrong or unsupported note has real clinical and financial consequences. The pipelines and evals you build are what let us ship model changes with confidence.
Own the Model Layer: You'll define how we post-train, evaluate, and deploy models end-to-end—not inherit someone else's stack.
Direct Clinical Impact: Every improvement in clinical reasoning or note quality directly reduces documentation burden and strengthens the charts clinicians and payers rely on.
Build post-training pipelines on open-source models-supervised fine-tuning, preference optimization (DPO/RLHF), LoRA/adapters, and distillation-for domain-specific clinical tasks.
Fine-tune, deploy, and serve models across managed inference and fine-tuning platforms such as Fireworks AI, Baseten, and Together AI, and make pragmatic build-vs-buy calls on where each workload should run.
Design and maintain rigorous evaluation sets for high-stakes tasks like clinical reasoning and AI note generation-defining metrics, curating gold-standard data, and building automated and human-in-the-loop eval harnesses.
Turn eval results into a fast, trustworthy iteration loop: catch regressions before they ship, and quantify the impact of every model or prompt change.
Optimize the full LLM pipeline-prompting, retrieval, structured-output validation, latency, and cost.
Partner with clinical experts to translate documentation and compliance requirements into model behavior and evaluation criteria.
Monitor models in production for quality, drift, and failure modes, and close the loop back into training data and evals.
3+ years in applied ML / AI engineering, or a Master's degree in a related field, with hands-on experience taking LLM-based systems into production.
Practical experience with post-training / fine-tuning open-source models (e.g., Llama, Qwen, Mistral) using SFT, LoRA/PEFT, or preference-based methods.
Experience serving or fine-tuning models on managed platforms such as Fireworks AI, Baseten, or Together AI (or comparable inference/training infra).
Demonstrated ability to build evaluation frameworks for LLM tasks—you think in terms of measurable quality, not vibes.
Strong Python and familiarity with the modern ML tooling ecosystem (PyTorch, Hugging Face, etc.).
Solid grounding in prompt engineering and structured-output validation.
Ability to thrive in a fast-paced, remote startup and communicate clearly with technical and clinical teammates.
We're willing to sponsor visas, including H-1B and O-1, for the right candidate.
Experience with healthcare, clinical NLP, or other high-stakes / regulated domains.
Familiarity with HIPAA and handling sensitive clinical data.
RAG systems, retrieval quality tuning, or long-context document workflows.
Experience with LLM observability, monitoring, and drift detection in production.
Data pipeline and labeling workflow experience for curating high-quality training and eval sets.
Open-source contributions in the ML/LLM ecosystem.
Salary range of $150,000-$200,000, plus equity with significant upside potential as a founding team member
Comprehensive health, dental, and vision insurance
Flexible, remote-first culture
Direct access to founders and influence on technical direction
Professional development budget and conference attendance
The chance to build AI that measurably improves mental health care at scale