Software Engineer, ML Infra & Distributed Systems (Staff & Principal)

Tubi Tv

San Francisco (CA)

On-site

USD 130,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading streaming service is seeking a Staff Software Engineer to enhance ML infrastructure. The role involves designing scalable systems, mentoring engineers, and collaborating with cross-functional teams. Candidates should have over 8 years of experience in building distributed systems, strong skills in AWS, and knowledge of microservices. Additionally, familiarity with machine learning infrastructure is a plus. This position is based in San Francisco and allows for hybrid work.

Qualifications

  • 8+ years of experience designing and building scalable distributed systems.
  • Strong experience with AWS or an equivalent cloud platform.
  • Experience with online microservices at scale with low latency.

Responsibilities

  • Design and build scalable distributed systems using Scala.
  • Collaborate with ML engineers to optimize ML Inference stack.
  • Mentor other engineers on system design and incident management.

Skills

Scalable, distributed systems design
Machine learning
AWS or equivalent cloud platforms
Microservices
Containerization (Docker, Kubernetes)

Tools

SQL (Postgres)
NoSQL (Cassandra)
Message brokers (Kafka)
Caches (Redis)

Job description

Staff Software Engineer, ML Infra & Distributed Systems

About the Role:

As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and may contribute to open-source projects. This role grants architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams.

Responsibilities:

  • Design and build scalable, high throughput, and low latency distributed systems using Scala
  • Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration
  • Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art.
  • Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary
  • Mentor other engineers on system design, incident management, interviewing, leveraging LLMs for work, etc.
  • Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi.

Your Background:

  • 8+ years of experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM based language is a plus.
  • Strong experience with AWS or an equivalent cloud platform
  • Experience building online microservices at scale with low latency serving
  • Experience with both SQL (e.g. Postgres) and NoSQL databases (e.g. Cassandra), message brokers (e.g. Kafka), and caches (e.g. Redis)
  • Experience with containerization technologies, such as Docker or Kubernetes
  • Led the response and resolution efforts for multiple major, large-scale incidents

Bonus:

  • Familiarity with machine learning infrastructure like inference engines (e.g. torchserve, Triton, vLLM), vector stores (e.g. LanceDB, FAISS), feature stores (e.g. Feast)
  • Understanding of ML model training pipelines and model internals. Experience with Recommender Systems, Search, Autocomplete and Ads ML is a plus
  • Previous experience with Akka, Erlang, Elixir or Go
  • Proficient in data-driven analysis of complex A/B testing results
About Tubi:

Tubi is a free streaming service that entertains over 100 million monthly active users. Tubi offers a large collection of Hollywood movies and TV shows, thousands of creator-led stories and hundreds of Tubi Originals. Headquartered in San Francisco and founded in 2014, Tubi is part of Tubi Media Group, a division of Fox Corporation.

#LI-Hybrid

We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, gender identity, disability, protected veteran status, or any other characteristic protected by law. We will consider for employment qualified applicants with criminal histories consistent with applicable law.

Disclosures and additional program details are available upon request. This job description is intended to describe the general nature and level of work performed by employees assigned to this position and is not intended to be an exhaustive list of all duties and responsibilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, ML Infra & Distributed Systems (Staff & Principal)
Software Engineer, ML Infra & Distributed Systems (Staff & Principal)

Tubitv • San Francisco (CA)

On-site
USD 227,000 - 417,000
Flexible Time Off
Generous Parental Leave
Monthly Wellness Reimbursement
Software Engineer, ML Infra & Distributed Systems (Staff & Principal)
Software Engineer, ML Infra & Distributed Systems (Staff & Principal)

Tubi • United States

On-site
USD 227,000 - 417,000
Flexible Time Off
Parental Leave Program
Wellness reimbursement
Software Engineer, ML Infra & Distributed Systems (Staff & Principal)
Software Engineer, ML Infra & Distributed Systems (Staff & Principal)

Tubi • New York (NY)

On-site
USD 227,000 - 417,000
Flexible Time Off Policy
Generous Parental Leave Program
Monthly wellness reimbursement
Director, ML Engineering & Infrastructure
Director, ML Engineering & Infrastructure

Tubi • New York (NY)

Hybrid
USD 292,000 - 418,000
Flexible Time Off Policy
Generous Parental Leave Program
Monthly Wellness Reimbursement
Machine Learning Engineer (Staff & Principal)
Machine Learning Engineer (Staff & Principal)

Tubitv • San Francisco (CA)

Hybrid
USD 239,000 - 342,000
Flexible Time Off
Parental Leave Program
Monthly wellness reimbursement
Staff Software Engineer, Internal Tools
Staff Software Engineer, Internal Tools

Tubi • New York (NY)

On-site
USD 227,000 - 325,000
Flexible Time off Policy
Generous Parental Leave Program
Monthly wellness reimbursement
Staff Software Engineer (Backend)
Staff Software Engineer (Backend)

Tubitv • San Francisco (CA)

Hybrid
USD 227,000 - 325,000
Flexible time-off policy
Generous parental leave program
Monthly wellness reimbursement
Staff Software Engineer (Backend)
Staff Software Engineer (Backend)

Tubi • San Francisco (CA)

On-site
USD 227,000 - 325,000
Medical, dental, and vision insurance
401(k) plan
Paid time off
+3
Machine Learning Engineer (Staff & Principal)
Machine Learning Engineer (Staff & Principal)

Tubi Tv • Los Angeles (CA), New York (NY), San Francisco (CA)

Hybrid
USD 239,000 - 342,000
Annual discretionary bonus
Long-term incentive plan
Medical/dental/vision insurance
+2
Senior Software Engineer, Backend
Senior Software Engineer, Backend

Tubi • San Francisco (CA)

Hybrid
USD 186,000 - 267,000
Medical/dental/vision insurance
401(k) plan
Paid time off