ML/AI Engineer

Kelly

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kelly is seeking a Staff AI/ML Engineer who is product-minded, hands-on, and excited to ship production ML/LLM systems that turn messy, real-world text and behavioral data into reliable experiences for patients.

You will develop LLM applications with retrieval and tool use, convert unstructured text into structured signals, design production pipelines, and collaborate with product, engineering, data, and clinical experts to ensure safe, measurable ML outcomes.

Qualifications

  • 8+ years building and shipping production ML systems with measurable impact.
  • Strong Python skills and ML/LLM library experience (Hugging Face, LangChain).
  • Designs production-grade training/inference/eval pipelines and monitors models in live systems.
  • Solid grounding in NLP, deep learning, statistics and evaluation.
  • MLOps practices: versioning, reproducibility, CI/CD, model registry, data management.
  • Experience with Spark/Databricks or equivalent for large-scale data.
  • Excellent product sense to translate business needs into ML outcomes.

Responsibilities

  • Develop LLM applications with retrieval and tool use for trusted health experiences.
  • Convert unstructured text into structured signals (topics, entities, intent, sentiment).
  • Build data pipelines for training, inference, evaluation and analytics.
  • Design evaluation systems: offline metrics, golden datasets, online A/B testing alignment.
  • Implement guardrails to reduce harm and ensure proper citations/attribution.
  • Set up monitoring for latency, cost, drift, and quality metrics.
  • Partner with Product, Engineering, Data, and clinical experts to validate outputs.
  • Lead architecture and technical direction for applied AI; mentor engineers.

Skills

Production ML
Python
MLOps
NLP/DL fundamentals
Data processing (Spark/Databricks)
Product stakeholder instincts
ML pipelines design

Tools

Redshift
Looker/Tableau
Terraform
Airflow
GitHub Actions
AWS Bedrock
Databricks

Job description

  • We’re looking for a Staff AI/ML Engineer who is product-minded, hands‑on, and excited to ship production ML/LLM systems that turn messy, real-world text and behavioral data into reliable experiences for patients.
RESPONSIBILITIES
  • Develop LLM applications with retrieval and tool use (e.g., RAG, orchestration/workflows, structured extraction) to deliver trustworthy consumer health experiences.
  • Convert unstructured text (posts, comments, messages, search queries) into structured signals
  • (topics, entities, intent, sentiment, safety flags) using a mix of classical NLP and modern LLMs.
  • Create and maintain data pipelines for training, inference, evaluation, and analytics (batch and/or
  • streaming as needed).
  • Design evaluation systems that measure quality and safety: offline metrics, golden datasets, human
  • review workflows, and online A/B testing alignment.
  • Implement production guardrails to reduce harm and misinformation risk (policy constraints, refusal behavior, citations/attribution when appropriate, red‑teaming, monitoring, and incident response).
  • Set up monitoring for model + system health (latency, cost, drift, regressions, quality metrics).
  • Partner closely with Product, Engineering, Data, and clinical/subject‑matter experts to validate
  • outputs and define what “correct” means for sensitive health‑adjacent use cases.
  • (Staff scope) Lead architecture and technical direction for applied AI across the organization; mentor engineers; establish best practices and reusable platforms.
Examples of problems you might work on :
  • Personalized recommendations for communities, posts, resources, or next‑best actions
  • Safer content understanding: detection of misleading/high‑risk health claims, escalation workflows
  • Search and discovery improvements using embeddings, hybrid retrieval, and ranking
  • Summarization and structuring of long threads into navigable insights (with safety constraints)
  • Member intent understanding from behavioral + text signals
MUST‑HAVE QUALIFICATIONS
  • 8+ years building and shipping production ML systems (or equivalent experience with demonstrable impact).
  • Strong Python skills and experience with ML/LLM libraries and tooling (e.g., Hugging Face ecosystem, LangChain/LangGraph or equivalent).
  • Proven ability to design production‑grade pipelines (training/inference/eval), and operate models in real systems (monitoring, rollbacks, incident handling).
  • Solid grounding in ML fundamentals (NLP, deep learning, statistical reasoning, evaluation).
  • Experience with MLOps best practices: versioning, reproducibility, CI/CD, model registry patterns,
  • feature/data management, and infrastructure collaboration.
  • Experience working with large‑scale data using Spark/Databricks or equivalent distributed
  • processing.
  • Strong product and stakeholder instincts: you can translate ambiguous business needs into
  • measurable ML outcomes.
NICE‑TO‑HAVE QUALIFICATIONS
  • Experience building RAG and retrieval systems: vector databases, hybrid search, ranking,
  • recommendation, query understanding.
  • Experience in healthcare or regulated environments, including privacy‑by‑design, auditability, and
  • safety reviews (HIPAA/PHI familiarity a plus).
  • Experience with streaming/clickstream data, experimentation platforms, and causal/measurement
  • thinking.
  • Ability to prototype end‑to‑end experiences (e.g., Streamlit, Gradio, lightweight frontends).
SOME TOOLS WE USE
  • Redshift and BI tools (Looker/Tableau) for analytics
  • Terraform for infrastructure‑as‑code; Airflow for orchestration; GitHub Actions for CI/CD
  • AWS (including Bedrock) and a mix of private and open‑weight models (including fine‑tunes where appropriate)
  • Experimentation tooling (A/B testing) and internal UX analytics tools
  • AI‑assisted coding tools (e.g., Cursor, Copilot, Claude/OpenAI tooling)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Syndesus, Inc. • Austin (TX)

On-site
USD 100,000 - 140,000
100% employer-paid health, vision, and dental insurance
Retirement plans (401(k))
Disability insurance
+1
Senior AI Engineer
Senior AI Engineer

7AI • Boston (MA)

On-site
USD 140,000 - 200,000
Staff AI Software Engineer
Staff AI Software Engineer

Harnham • San Francisco (CA)

On-site
USD 150,000 - 200,000
AI/ML Engineer
AI/ML Engineer

RiskForce • Northern (KY)

Hybrid
USD 120,000 - 155,000
Senior MLOps Engineer
Senior MLOps Engineer

C the Signs • United States

Hybrid
USD 120,000 - 160,000
Competitive salary and benefits
Flexible working arrangements
Continuous learning opportunities
Senior LLM Systems Architect — AI for Security & Automation
Senior LLM Systems Architect — AI for Security & Automation

7AI • Boston (MA)

On-site
USD 140,000 - 200,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
Staff AI/ML Engineer - Production Health LLM Systems
Staff AI/ML Engineer - Production Health LLM Systems

Kelly • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000