Senior ML Infra Engineer - Low-Latency Scale

Patreon

New York, San Francisco (NY, CA)

Hybrid

USD 180,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare
401k with matching
Paid time off

Job summary

Patreon seeks a Senior Machine Learning Engineer, Infrastructure, to design and scale high-throughput ML infrastructure with low-latency live inference capabilities. The role is hybrid, with in-office presence in San Francisco or New York three days a week.

You will own the feature store lifecycle from ingestion to production serving, ensuring data consistency between online and offline features, and will work with product, data engineering, and trust & safety to deliver robust infrastructure.

Qualifications

  • Deep experience building, deploying, and maintaining production-grade ML infrastructure at scale, with low-latency live inference pipelines and feature store architectures.
  • Strong background in distributed systems and backend engineering, writing robust, maintainable Python code.
  • Systematic debugging of high-throughput systems and performance bottlenecks.
  • Experience building '0 to 1' infrastructure foundations that stand the test of time.
  • Clear documentation for system architectures and infrastructure strategies.
  • Growth mindset and improving developer velocity through code reviews.

Responsibilities

  • Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support relevance systems.
  • Own the end-to-end feature store lifecycle—from ingestion and transformation to production serving, ensuring high availability and consistency between online and offline features.
  • Design and implement observability, monitoring, and validation frameworks to detect performance gaps, latency spikes, and production drift.
  • Collaborate with product, data engineering, and trust and safety to translate product requirements into robust, scalable infrastructure solutions.
  • Automate model deployment and reliability testing to improve developer velocity and ensure system stability.
  • Debug complex relevance systems when monitoring identifies performance bottlenecks or reliability issues.

Skills

ML infra
Low-latency
Python
Distributed systems
Back-end engineering
Documentation
Collaboration

Tools

Kubernetes
Terraform
Airflow

Job description

Patreon seeks a Senior Machine Learning Engineer, Infrastructure, to design and scale high-throughput ML infrastructure with low-latency live inference capabilities. The role is hybrid, with in-office presence in San Francisco or New York three days a week.

You will own the feature store lifecycle from ingestion to production serving, ensuring data consistency between online and offline features, and will work with product, data engineering, and trust & safety to deliver robust infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: Scale Real-Time ML Infra & Data
Senior ML Engineer: Scale Real-Time ML Infra & Data

Ultra • New York (NY)

Hybrid
USD 180,000 - 240,000
Cloud-Scale Backend Engineer for ML Inference
Cloud-Scale Backend Engineer for ML Inference

Praxis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 250,000
Staff ML Infra Engineer: Global, Low-Latency ML Systems
Staff ML Infra Engineer: Global, Low-Latency ML Systems

Uber • San Francisco (CA)

On-site
USD 232,000 - 258,000
Bonus program
Equity award
401(k) plan
+1
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
ML Platform Engineer — Scalable AI Infra
ML Platform Engineer — Scalable AI Infra

Whatnot • Seattle (WA)

Hybrid
USD 245,000 - 345,000
Health Insurance
Work From Home Support
Care benefits
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
ML Infra Tech Lead: Scalable Training & Inference
ML Infra Tech Lead: Scalable Training & Inference

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO