Member of Technical Staff [AI/ML Engineer]

Burnt Group

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 275,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

BURNT is seeking an AI/ML Engineer focused on MLOps for our on-site San Francisco office. You will own the model layer from feature engineering to deployment in a fast-growing supply chain context.

You will work on time series forecasting, ontology/knowledge graphs, and in-house LLM fine-tuning, building production-grade systems and collaborating closely with product and engineering teams to ship value.

Qualifications

  • Python in production environments.
  • Experience deploying ML models to production.
  • Experience with AWS SageMaker and Spark.

Responsibilities

  • Own MLOps end to end: model creation, deployment, iteration, monitoring, and support.
  • Build time series forecasting models for production traffic.
  • Fine-tune LLMs with LoRA/PEFT on our data, and build evals.
  • Design ontology and knowledge graph for agents.
  • Engineer data pipelines at scale with Spark.
  • Architect production systems end to end and ship full stack work when needed.

Skills

Python
ML engineering
MLOps
Time series forecasting
Knowledge graph
Data pipelines

Tools

AWS SageMaker
Apache Spark
S3/Glue/Step Functions
LoRA/PEFT
Neo4j
TypeScript/React

Job description

BURNT
Member of Technical Staff
AI/ML Engineer (MLOps-Focused)

Location On-site, San Francisco · Experience 5–7 years · Compensation $150,000–$275,000 + equity

About Burnt

Burnt isn't building software on top of ERPs. We don't believe ERPs will exist in the long run. They were built for a world where humans key in data and software stores it. That world is ending.

We're building the Operating Brain for the global supply chain. A living, evolving brain that will run half of these businesses on autopilot.

The brain runs on models. Every order a distributor has ever placed, every substitution a buyer has ever made, every vendor that has ever shorted a delivery. That history is sitting in ERPs doing nothing. We turn it into forecasts, into an ontology our agents can reason over, and into models tuned on our own data instead of someone else's API.

We're starting in food, a $1T+ industry that feeds the country and has been ignored by modern software. Not because it's small. Because it's operationally complex and unforgiving. Product moves in cases, pounds, and pallets at the same time. Shelf life is measured in days. Demand is intermittent and seasonal and breaks every time a customer runs a promo. This is a hard modeling problem, which is exactly why nobody has solved it.

Culturally, we are extremely competitive. We run through walls for customers. We build elbows up. We're here to build how supply chain companies will run for the next 20 years.

The role

You are our first dedicated ML hire. The model layer is yours to build.

Today our agents run on general-purpose APIs and rules. That gets us to production. It doesn't get us to a system that gets sharper every month, and it doesn't get us to margins that work at scale. Both of those are your job.

Three problems, in priority order.

Demand forecasting. Distributors buy on gut and get punished for it. Overbuy on a perishable and you write it off. Underbuy and you short a customer who leaves. You'll own forecasting end to end: pulling order history out of ERPs, engineering features from it, choosing the model, backtesting it against real order books rather than a random split, deploying it, and watching it drift. The hard parts here are intermittent demand on the long tail, cold-start on new SKUs, and separating a real trend from a customer who ran a promo last March.

Ontology and knowledge graph. Agents can't reason about a business they can't represent. A single product exists as five different vendor SKUs, three pack sizes, and two units of measure, and every distributor names it differently. You'll design the ontology and the graph that resolves that, and it has to hold up when we onboard a customer whose data is worse than the last one's.

Bringing LLMs in-house. We fine-tune with LoRA and PEFT on our own data to beat the general-purpose APIs on our tasks and cut token spend at the same time. You'll own the datasets, the training runs, and the evals that decide whether a fine-tune actually ships.

This is not a research seat. Models that live in a notebook are worth nothing to us. Everything you build goes into production, carries traffic, retrains itself, and gets measurably better.

It is also not a siloed one. You need to be able to architect a system, not just a model, and defend the decisions you made. We are a small team shipping fast, so there will be weeks where the highest-leverage thing you can do is pick up full stack work and ship a feature with the product engineers. The people who do well here are the ones who reach for that instead of waiting for it to be someone else's problem.

The data

You’ll work with order and transaction history pulled from customer ERPs across thousands of distributors, 500k+ SKUs, and years of history, plus the unstructured side: emails, PDFs, and messages that carry the orders these systems never captured. It is real operational data, which means it is messy, inconsistent between customers, and full of the kind of edge cases that only show up in production. That is the job.

What you'll do
  • Own MLOps end to end: model creation, deployment, iteration, monitoring, and support. No handoffs.
  • Build and maintain time series forecasting models that serve production traffic and hold up under backtest against real order books.
  • Fine-tune LLMs with LoRA and PEFT on our own data, and build the evals that decide what ships.
  • Design and maintain the ontology and knowledge graph our agents reason over.
  • Engineer data pipelines at scale with Spark, over messy multi-tenant supply chain data.
  • Build the versioning, drift detection, and retraining pipelines that keep models honest after launch.
  • Run the AWS ML stack: SageMaker at the core, with S3, Glue, and Step Functions around it.
  • Architect the systems your models live inside, not just the models, and own those design decisions.
  • Step into full stack work when the team needs it, from the API that serves a prediction to the interface a buyer actually uses.
Mandatory tech stack

You should be fluent in most of this and able to ramp fast on the rest:

  • Core: Python at expert level. This is the whole job, not a nice-to-have.
  • Data: Apache Spark and the surrounding data engineering ecosystem.
  • Platform: AWS SageMaker as the core platform, plus S3, Glue, Step Functions and the rest.
  • MLOps: MLflow, Kubeflow, or equivalent.
  • Fine-tuning: LoRA and PEFT libraries such as HuggingFace PEFT or TRL.
  • Forecasting: Prophet, NeuralForecast, statsmodels, or similar.
  • Knowledge graph: Neo4j, RDF, OWL, SPARQL, or similar.
  • Application layer: enough TypeScript and React to be useful in our codebase. You don't need to have shipped a frontend last quarter, but you do need to be willing to.
What we expect you've done
  • Owned MLOps end to end, from model creation through deployment, iteration, monitoring, and support.
  • Fine-tuned LLMs with LoRA or PEFT on real datasets, not toy ones.
  • Built and maintained time series forecasting models serving production traffic.
  • Worked inside systems backed by ontologies and knowledge graphs.
  • Operated across the AWS ecosystem beyond SageMaker.
  • Engineered data pipelines at scale with Spark.
  • Built model versioning, drift detection, and retraining pipelines that ran without you watching them.
  • Architected production systems end to end and can walk through the tradeoffs you chose and what you would do differently now.
  • Worked outside the model layer when it was needed, shipping application code alongside product engineers.
Round 1 filter — the non-negotiables

If you do not clear every line below, this is not the right role. We screen on these first, no exceptions:

  • Python at an expert level. Demonstrable, not claimed.
  • ML models, not agents, deployed and maintained in production.
  • AWS SageMaker hands-on.
  • A time series forecasting model in production. Hard filter, no exceptions.
  • LLM fine-tuning with LoRA or PEFT. You've done it, not read about it.
  • Ontology or knowledge-graph-backed systems in a real product context.
  • Can articulate system design decisions you personally architected.
  • Willing and able to pick up full stack work when the team needs it. No "that's not my job."
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Lead (AI Data Labeling)
Machine Learning Lead (AI Data Labeling)

NewtonX • United States

On-site
USD 140,000 - 190,000
Medical, dental, and vision insurance
401K match
Paid vacation and holidays
+3
Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Senior AI/ML Engineer
Senior AI/ML Engineer

UMATR • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding Software Engineer - AI & Operations (Full-Stack)
Founding Software Engineer - AI & Operations (Full-Stack)

United Concrete Inc • Wallingford (CT), Northern (KY)

Hybrid
USD 120,000 - 190,000
Paid Vacation
Paid Holidays and Sick Time
Company 401K 10% Match
+1
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
Senior AI-First Software Engineer (End-to-End SaaS)
Senior AI-First Software Engineer (End-to-End SaaS)

Turing Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

Socket.dev • Shelton (CT)

On-site
USD 140,000 - 190,000
Senior Full Stack Engineer
Senior Full Stack Engineer

Bucket Robotics • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision insurance
401(k)
Equity
+2
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Cassi Home • United States

On-site
USD 180,000 - 260,000