Senior Machine Learning Engineer, Infrastructure

United States Digital Space LLC

New York (NY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking a Senior Machine Learning Engineer, Infrastructure to build and scale production-grade ML infrastructure for high-throughput relevance systems. The role emphasizes low-latency live inference, feature store management, and robust observability.

You will collaborate with cross-functional teams across Product, Data Engineering, and Trust & Safety to deploy reliable infrastructure and improve developer velocity.

Qualifications

  • Deep experience building, deploying, and maintaining production-grade ML infrastructure at scale.
  • Strong background in distributed systems and backend engineering.
  • Proficient in Python with reliable, maintainable code.

Responsibilities

  • Architect, scale, and maintain high-throughput, low-latency live inference infrastructure.
  • Own end-to-end feature store lifecycle—from ingestion to production serving.
  • Design observability, monitoring, and validation frameworks.
  • Collaborate with product, data engineering, and trust & safety.
  • Automate model deployment and reliability testing.
  • Debug complex relevance systems to fix bottlenecks.

Skills

ML infrastructure
Distributed systems
Python
Performance debugging

Job description

the company is a media and community platform where over 300,000 creators give their biggest fans access to exclusive work and experiences. We offer creators a variety of ways to engage with their fans and build a lasting business including: paid memberships, free memberships, community chats, live video, and selling to fans directly with one-time purchases.

Ultimately our goal is simple: fund the creative class. And we're leaders in that space, with:

  • $10 billion+ generated by creators since the company's inception
  • 100 million+ free memberships for fans who may not be ready to pay just yet, and
  • 25 million+ paid memberships on the company today.

We're continuing to invest heavily in building the best creator platform with the best team in the creator economy and are looking for a Senior Machine Learning Engineer, Infrastructure to support our mission.

This role is based in San Francisco or New York as an in-office 3 days per week on a hybrid work model.

About the Team

You’ll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how content surfaces across the company. The team is responsible for search, feed ranking, and creator-fan matching. You’ll work closely with a small, collaborative group of MLEs on shared infrastructure, code reviews, and roadmap alignment, while partnering cross-functionally with Product, Data Engineering, and Trust & Safety to deliver measurable impact across the platform.

About the Role
  • Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support our relevance systems.
  • Own the end-to-end feature store lifecycle—from ingestion and transformation to production serving, ensuring high availability and consistency between online and offline features.
  • Design and implement observability, monitoring, and validation frameworks to detect performance gaps, latency spikes, and production drift.
  • Collaborate with cross-functional partners, such as product, data engineering, and trust and safety, to translate product requirements into robust, scalable infrastructure solutions.
  • Automate model deployment and reliability testing to improve developer velocity and ensure system stability.
  • Debug complex relevance systems when monitoring identifies performance bottlenecks or reliability issues.
About You
  • You have deep experience building, deploying, and maintaining production-grade ML infrastructure at scale, specifically with low-latency live inference pipelines and feature store architectures.
  • You have a strong background in distributed systems and backend engineering, with the ability to write robust, maintainable code in Python.
  • You have a systematic approach to debugging complex, high-throughput systems and performance bottlenecks.
  • You are energized by building '0 to 1' infrastructure systems that stand the test of time and provide a reliable foundation for the team.
  • You possess strong communication skills and are effective at creating clear documentation for system architectures and infrastructure strategies.
  • You have a growth mindset, a keen eye for detail in code reviews, and a passion for empowering your teammates by improving developer velocity.

We hire talented and passionate people from different backgrounds because workplace diversity and inclusion is critical to our ability to serve creators worldwide.

About the company

the company powers creators to do what they love and get paid by the people who love what they do. Our team is passionate about making this mission and our core values come to life every day in our work. Through this work, our Patronauts:

  • Put Creators First | They’re the reason we’re here. When creators win, we win.
  • Build with Craft | We sign our name to every deliverable, just like the creators we serve.
  • Make it Happen | We don’t quit. We learn and deliver.
  • Win Together | We grow as individuals. We win as a team.

the company is proud to be an equal opportunity employer. We provide employment opportunities without regard to age, race, color, ancestry, national origin, religion, disability, sex, gender identity or expression, sexual orientation, veteran status or any other protected class. If you need a reasonable accommodation during the interview process, please let us know via email at accommodations@the company.

Benefits

the company offers a competitive benefits package including and not limited to salary, equity plans, healthcare, flexible time off, company holidays and recharge days, commuter benefits, lifestyle stipends, learning and development stipends, patronage, parental leave, and 401k plan with matching.

Work Model

the company operates under a hybrid work model, where employees based in office locations are expected to come into the office two days per week, excluding sick time and paid leave. The goal of this policy is to be intentional about the in-person time we spend together to strengthen the feeling of community at the company. Candidates hired into remote-eligible roles are not expected to meet the same requirements.

Pay

At the company, we believe in fair and transparent pay. In compliance with New York and California pay transparency laws, we are sharing the expected salary range for this role.

The posted salary range is dependent on the location and the level. This range may encompass multiple levels within the role’s job family. The final offer will be based on candidate’s experience, skills, competencies, and geographic location, aligning with the appropriate job level within the company’s leveling framework. For remote employees located outside CA and NY, salary may vary based on location and local market conditions.

the company reserves the right to modify or update compensation and benefits at any time

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer, Infrastructure
Senior Machine Learning Engineer, Infrastructure

Patreon • New York (NY), San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Healthcare
401k with matching
Paid time off
ML Engineer
ML Engineer

CreatorIQ • Los Angeles (CA)

Hybrid
USD 132,000 - 165,000
Hybrid work model
Work from home stipend
15 days vacation
+4
Staff Machine Learning Engineer, Consumer
Staff Machine Learning Engineer, Consumer

Reddit, Inc. • San Francisco (CA)

On-site
USD 230,000 - 322,000
Healthcare Benefits
401k with Employer Match
Generous Parental Leave
+1
Senior Machine Learning Infrastructure Engineer, Embedding Platform New Remote - United States
Senior Machine Learning Infrastructure Engineer, Embedding Platform New Remote - United States

Reddit, Inc. • Northern (KY)

Hybrid
USD 191,000 - 267,000
Healthcare benefits
401k with employer match
Vacation & parental leave
+1
Software Engineer, ML Infra (Junior & New Grad)
Software Engineer, ML Infra (Junior & New Grad)

NewsBreak • Mountain View (CA)

On-site
USD 125,000 - 175,000
Staff Machine Learning Engineer, AI R&D
Staff Machine Learning Engineer, AI R&D

United States Digital Space LLC • Bellevue (CA)

Hybrid
USD 180,000 - 260,000
Health insurance
Equity ownership
401(k) matching
+2
Senior Machine Learning Systems Engineer
Senior Machine Learning Systems Engineer

Reddit, Inc. • San Francisco (CA)

On-site
USD 216,000 - 304,000
Equity (RSUs)
401(k) with employer match
Generous vacation time
+2
Senior Machine Learning Manager, Video Ranking
Senior Machine Learning Manager, Video Ranking

AI Chopping Block • Northern (KY)

Hybrid
USD 266,000 - 372,000
Healthcare benefits
401k
Global benefits
+5
Senior Staff Machine Learning Platform Engineer
Senior Staff Machine Learning Platform Engineer

United States Digital Space LLC • United States

Hybrid
USD 295,000 - 406,000
Software Engineer, ML Platform
Software Engineer, ML Platform

United States Digital Space LLC • Los Angeles (CA)

Hybrid
USD 190,000 - 240,000
Equity (RSUs)