Senior Machine Learning Engineer (DevOps/SRE)

Roku

Austin (TX)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Roku is looking for a Senior Software Engineer specializing in MLOps/DevOps. You'll play a critical role in scaling the Machine Learning infrastructure for the Advertising Performance team. Ideal candidates will have a strong background in cloud infrastructure and DevOps practices, and will collaborate closely with ML teams to enhance end-to-end ML lifecycle processes.

Key responsibilities include leading the design of scalable cloud infrastructure and ensuring CI/CD systems support reliable production releases. This position is based in Austin, Texas.

Qualifications

  • 8+ years of experience in DevOps, SRE, or ML infrastructure.
  • 4+ years supporting large-scale ML or AI systems.
  • Strong infrastructure-as-code experience with Terraform.

Responsibilities

  • Support and scale Machine Learning infrastructure.
  • Partner with ML Scientists and Engineers to streamline ML lifecycle.
  • Lead the design and operation of scalable cloud infrastructure for ML workloads.

Skills

Python
Scala
Java
Kubernetes
GCP (GKE)
AWS (EKS)
NoSQL
Apache Spark
Apache Flink
Apache Airflow
Kafka
Jenkins
GitLab Runner
Terraform
Prometheus
Grafana
Datadog
MLflow
Chronon
ML systems

Education

BS or MS in Computer Science, Engineering, or related field

Job description

Requirements
  • BS or MS in Computer Science, Engineering, or a related quantitative field
  • 8+ years of experience in DevOps, SRE, or ML infrastructure, including 4+ years supporting large-scale ML or AI systems
  • Strong programming skills in Python and/or Scala or Java for platform automation and tooling
  • Deep experience with Kubernetes and container orchestration on GCP (GKE) and/or AWS (EKS)
  • Expertise with NoSQL or low-latency data stores such as Aerospike or similar technologies
  • Hands‑on experience with data and orchestration technologies such as Apache Spark, Apache Flink, Apache Airflow, and Kafka
  • Experience building and maintaining CI/CD systems using tools such as Jenkins or GitLab Runner
  • Familiarity with feature engineering platforms such as Chronon and model lifecycle tools such as MLflow
  • Strong infrastructure‑as‑code experience with Terraform or similar tooling
  • Experience with observability platforms such as Prometheus, Grafana, and Datadog
  • Excellent communication and cross‑functional collaboration skills
  • Experience in the Advertising domain is a plus
What the job involves
  • We are seeking a talented and experienced Senior Software Engineer, MLOps/DevOps to join the Advertising Performance team and play a critical role in supporting and scaling our Machine Learning infrastructure
  • The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling — with a passion for building platforms that accelerate ML experimentation and deployment at internet scale
  • You will partner closely with ML Scientists and Engineers to streamline the end‑to‑end ML lifecycle across training, evaluation, deployment, and monitoring — on top of a modern, cloud‑native stack running on GCP and AWS using Kubernetes, Apache Airflow, Spark, Ray, MLflow, Chronon, etc
  • Lead the design and operation of scalable, production‑grade cloud infrastructure for ML workloads across AWS and GCP, including GPU/TPU‑based training and inference environments
  • Architect and improve CI/CD systems for ML models and platform services to enable fast, reliable, and safe production releases
  • Own and evolve low‑latency infrastructure for real‑time model inference, including KV store and vector databases
  • Define and enforce observability standards for ML systems, including model performance monitoring, drift detection, capacity planning, and pipeline health metrics
  • Participate in on‑call rotation, leading incident response and root‑cause analysis for critical ML training and serving infrastructure
  • Partner with data scientists and ML engineers to improve platform usability, accelerate model iteration, and implement strong MLOps and SRE best practices
  • Champion operational excellence across ML infrastructure through automation, resilience engineering, disaster recovery planning, and continuous improvement
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Operations Engineer
ML Operations Engineer

NextGen Healthcare • Georgia

On-site
USD 80,000 - 120,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Senior MLOps Engineer
Senior MLOps Engineer

AppRecode, Inc. • Town of Middletown (NY)

On-site
USD 120,000 - 160,000
Senior Machine Learning Systems Engineer, Ads ML Experience Platform
Senior Machine Learning Systems Engineer, Ads ML Experience Platform

EngineersOfAI • United States

Hybrid
USD 100,000 - 130,000
Lead Machine Learning Engineer – News
Lead Machine Learning Engineer – News

Jobtailor • California (MO)

On-site
USD 180,000 - 240,000
Machine Learning Architect
Machine Learning Architect

Tiger Analytics Inc. • New Jersey

On-site
USD 120,000 - 160,000