Machine Learning Engineer

ESB Technologies

United States

On-site

USD 150,000 - 210,000

Full time

7 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

ESB Technologies seeks an experienced ML/LLM systems engineer to build production-grade pipelines that turn client interview data into structured process maps and SOPs. You will design retrieval and evaluation mechanisms across Claude, ChatGPT, Gemini and Slack, and implement autonomy scoring to optimize workflow automation.

You will work with founding engineers to ship end-to-end in production, balancing cost, latency and quality while ensuring privacy and tenant isolation across client

Qualifications

  • 5+ years building ML or LLM systems that run in production with real users.
  • Strong Python and sound engineering habits: testing, versioning, observability.
  • Hands-on LLM experience: context design, structured output, tool use, retrieval, agents.
  • A record of building the evaluation before trusting the model change.
  • Experience with unstructured input: transcripts, audio, documents, screenshots.
  • Working statistics: design a measurement, size a sample, and interpret results.
  • Comfort with ambiguity and with direct contact with customers and founders.

Responsibilities

  • Own the pipeline from interview transcripts, screen recordings and uploaded files to structured process maps and draft SOPs.
  • Build the evaluation harness: golden sets, regression tests, and scoring for extraction accuracy and SOP quality.
  • Design retrieval over client workflows so the assistant answers consistently across Claude, ChatGPT, Gemini and Slack.
  • Build the scoring behind the autonomy ladder and determine when to auto-publish.
  • Detect deviations from SOPs, classify, and route retries and human review.
  • Improve time and cost estimates per workflow with clear labeling as estimates.
  • Choose models per task based on cost, latency and quality; maintain portability across providers.
  • Ensure tenant isolation and privacy in every pipeline.

Skills

Python
LLM systems in production
Testing & observability
Customer interaction
Ambiguity tolerance

Tools

Claude
ChatGPT
Gemini
Slack

Job description

We find the manual work slowing teams down, rebuild it as automations inside tools clients already own, and run them. Our platform has three parts: a client workspace where operations are mapped and priced; a delivery system that runs projects and the builder network; and an intelligence layer across both. We have delivered 311 projects for more than 100 companies.

The role

You make the AI inside our intelligence layer and AI interviewer measurably reliable. The interviewer talks with client teams by chat, voice and screen share, then turns what it captures into process maps and draft SOPs. You own the quality of that pipeline: extraction, retrieval, evaluation, and the signals that decide when a workflow can run with less human review. This is applied work on foundation models from Anthropic, OpenAI and Google, not model research.

What you will do
  • Own the pipeline from interview transcript, screen recording and uploaded file to structured process map and draft SOP.
  • Build the evaluation harness: golden sets, regression tests, and scoring for extraction accuracy, SOP quality and assistant answers.
  • Design retrieval over each client's approved workflows so the assistant answers consistently in Claude, ChatGPT, Gemini and Slack.
  • Build the scoring behind the autonomy ladder: the run-quality signals that move a workflow from full review to spot-checks to auto-publish.
  • Detect when a run deviates from its SOP, classify the deviation, and route retries and human review.
  • Improve the time and money estimate attached to each workflow, and keep estimates labelled as estimates.
  • Choose models per task on cost, latency and quality, and keep the system portable across providers.
  • Hold tenant isolation and privacy in every pipeline. Client interviews and SOPs are private by default.
  • Ship end to end with the founding engineers and own what you ship in production.
What you bring
  • 5+ years building ML or LLM systems that run in production with real users.
  • Strong Python and sound engineering habits: testing, versioning, observability.
  • Hands-on LLM experience: context design, structured output, tool use, retrieval, agents.
  • A record of building the evaluation before trusting the model change.
  • Experience with unstructured input: transcripts, audio, documents, screenshots.
  • Working statistics. You can design a measurement, size a sample, and say what a result does not show.
  • Comfort with ambiguity and with direct contact with customers and founders.
Nice to have
  • Speech and multimodal pipelines for voice interviews and screen capture.
  • Process mining, workflow modelling or operations research.
  • Assistant integrations for Claude, ChatGPT, Gemini or Slack.
  • Classical ML for anomaly detection, classification or forecasting.
  • Time in a services or forward-deployed team.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Scientist, NLP & LLM Fine-Tuning
Senior Data Scientist, NLP & LLM Fine-Tuning

Xenoss • New York (NY)

Remote
USD 160,000 - 210,000
Staff Software Engineer, Applied AI
Staff Software Engineer, Applied AI

Emergence Capital Partners • San Francisco (CA)

On-site
USD 160,000 - 230,000
AI/ML Engineer
AI/ML Engineer

RiskForce • Northern (KY)

On-site
USD 120,000 - 155,000
Senior AI Engineer (Workflow & Systems)
Senior AI Engineer (Workflow & Systems)

TubeScience • Los Angeles (CA)

On-site
USD 110,000 - 160,000
Founding Engineer (Stackpoint Portfolio Company)
Founding Engineer (Stackpoint Portfolio Company)

Stackpoint Ventures • United States

Remote
USD 150,000 - 210,000
AI Engineer
AI Engineer

Ready Health • California (MO)

On-site
USD 120,000 - 180,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Malleable • Northern (KY)

Remote
USD 150,000 - 180,000
Automation AI Engineer
Automation AI Engineer

GigaBrands • United States

On-site
MXN 600,000 - 1,200,000
Competitive salary
High-impact role
Scale AI systems
Full Stack Automation Engineer
Full Stack Automation Engineer

GigaBrands • United States

Remote
USD 120,000 - 180,000
Remote work
PTO after probation
AI Engineer, Multimodal LLMs
AI Engineer, Multimodal LLMs

Eloquent AI • San Francisco (CA)

On-site
USD 150,000 - 210,000