About the role
Running CI on real hardware produces a kind of data most ML teams never get to touch: power traces, sensor streams, build artifacts, console logs, test results, all tied to specific physical machines doing specific physical work. There's serious signal in there — predicting flaky boards, classifying failure modes, scoring job risk, and eventually closed-loop optimization of how the fleet itself is used.
We're hiring a founding ML / data engineer to build the pipeline that turns that data into models, and the models into product features. You'll own this end-to-end — not "hand off a notebook and hope." Data ingestion, labeling, training infrastructure, evaluation, deployment, monitoring. You'll set the technical direction for ML at Primitive and shape what good looks like — from the first models in production to ML as a core part of the product.
This is the first dedicated ML hire. It's a build-from-zero role on top of a rich, real-world dataset.
What you'll do
- Design and build the data platform end-to-end: extend instrumentation where signals are missing today, then ingest from Postgres, SeaweedFS / S3, and streaming telemetry into a clean, versioned analytical layer
- Build the labeling workflow that lets us (and eventually customers) label hardware events without it becoming a permanent side project
- Design and operate a reproducible training stack on AWS — distributed where it needs to be, with experiment tracking, dataset versioning, and a real eval harness
- Ship inference for product features: low-latency serving where it matters, batch scoring where it fits
- Operate models in production: drift monitoring, regression gates, the dashboards that tell us when a model is silently rotting
- Partner with the full-stack and hardware teams to integrate predictions cleanly into the product surface
- Set the bar for ML rigor at Primitive: eval-first development, reproducibility, honest reporting of model quality
- Mentor new engineers on data and ML patterns as the team grows; raise the bar on data contracts, eval design, and reviewability
About you
- 5+ years in ML / data engineering, with at least one production ML system you took from raw data to served predictions
- Strong Python; comfortable with modern ML tooling (PyTorch, Hugging Face, Ray, or equivalents — we're not religious)
- Real opinions about data versioning, feature stores, and experiment tracking — you've used DVC, LakeFS, MLflow, or Weights & Biases and know what each is good and bad at
- Production data pipeline experience with a real data warehouse — schema design, contracts, ownership
- AWS chops: S3, EKS-hosted training (or SageMaker), IAM that doesn't terrify the security team
- Built or operated a human-in-the-loop labeling workflow
- Comfortable setting architectural direction for ML/data at a small company, balancing vision with pragmatism
Bonus
- Time-series, sensor, or signal-processing ML — we have a lot of it
- LLM fine-tuning, retrieval, or agent eval experience (there's product surface here too)
- Background in hardware, EE, or anything physical — helps a lot when the data is from real machines
- ClickHouse, dbt, Airflow / Dagster / Prefect at production scale
- Contributed to open-source ML or data tooling
Stack
Python, PyTorch, AWS (S3, EKS, possibly SageMaker), PostgreSQL, SeaweedFS, Grafana / Mimir. Pipeline orchestration is open (Airflow / Dagster / Prefect).
How we work
- Offices in New York, NY and San Francisco, CA — flexible in-office attendance, no fixed days per week
- We like working together in person: regular team meetups across both offices
- Small team, high ownership — most engineers ship to production in their first week
- Light on-call for serving infrastructure once models are in production
Benefits
- Health, dental, and vision for you and your dependents
- Substantial equity, with early exercise and an extended post-termination exercise window
- Unlimited paid vacation, plus local holidays
- Equinox membership
Compensation
Base salary: $180,000 – $250,000, adjusted based on location, level, and experience. Total compensation includes substantial equity and the benefits above.
A note on applying
If you don't tick every box on the list above, apply anyway. The bullets describe the engineer we'd be thrilled to hire; what we actually need is someone who can do the work and learn the rest. We especially encourage applications from people who don't see themselves represented in tech today.