Senior MLOps / ML Platform Engineer

SmartRecruiters, Inc.

United States

Remote

USD 140,000 - 190,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Sigma Software is seeking a Senior MLOps Engineer to help build a production-grade ML platform for a large-scale AdTech ecosystem. You will work on infrastructure powering predictive decision-making systems processing hundreds of millions of auction requests daily.

You will contribute to scalable ML orchestration, model lifecycle automation, observability, and real-time optimization workflows, collaborating with DevOps and SRE to drive CI/CD and infrastructure automation.

Qualifications

  • 5+ years of MLOps, ML platform engineering, or production ML infrastructure experience.
  • Strong Python skills and automation focus.
  • Hands-on experience with Kubernetes and Docker.
  • Experience building CI/CD pipelines for ML workloads.
  • Familiarity with ML observability and drift monitoring.
  • Experience with multi-tenant ML systems and isolated environments.

Responsibilities

  • Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedules.
  • Implement retries, backfills, and idempotent execution mechanisms.
  • Design and support model registry workflows including versioning, lineage, evaluation gates, and promotion processes.
  • Develop isolated per-advertiser model environments with namespace and configuration separation.
  • Build scalable refresh pipelines and publishing workflows for serving infrastructure.
  • Implement shadow mode and champion/challenger deployment strategies.
  • Develop monitoring and alerting for ML metrics including drift and calibration decay.
  • Ensure reproducibility of ML workflows with containerized environments and data snapshots.
  • Collaborate with DevOps and SRE on CI/CD and infra automation.
  • Prepare operational documentation and platform handover materials.

Skills

Strong Python skills
Excellent collaboration and comms
Problem solving and scalability focus

Tools

Kubernetes
Docker
Terraform
MLflow
Kubeflow
Airflow
Argo Workflows
Vertex Pipelines
Vertex AI

Job description

We are looking for a Senior MLOps Engineer to join Sigma Software and help build a production-grade ML platform for a large-scale AdTech ecosystem. You will work on infrastructure powering predictive decision-making systems that process hundreds of millions of auction requests daily.

As part of a dedicated engineering team, you will contribute to scalable ML orchestration, model lifecycle automation, observability, and real-time optimization workflows. This role is ideal for engineers with strong production experience who enjoy solving complex platform and operational challenges.

We at Sigma Software offer the opportunity to work on cutting-edge ML infrastructure projects, collaborate with experienced engineers, and influence architecture decisions in a long-term strategic engagement.

CUSTOMER

Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a high-load ad exchange platform handling hundreds of millions of auction requests every day and is investing in advanced predictive decision-making capabilities to improve advertiser outcomes and real-time optimization processes.

PROJECT

Sigma Software is building a predictive modeling and optimization platform integrated with a live ad exchange environment. The solution enables real-time supply scoring and filtering, audience look-alike generation, contextual performance estimation, and multi-objective optimization under operational constraints.

The project combines large-scale ML infrastructure, automated model lifecycle management, multi-tenant architecture, and advanced observability practices. The team focuses on delivering reliable, reproducible, and scalable ML systems ready for long-term Customer ownership.

Job Description
  • Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedules
  • Implement retries, backfills, and idempotent execution mechanisms
  • Design and support model registry workflows including versioning, lineage, evaluation gates, and promotion processes
  • Develop isolated per-advertiser model environments with namespace and configuration separation
  • Build scalable refresh pipelines and publishing workflows for serving infrastructure
  • Implement shadow mode and champion/challenger deployment strategies
  • Develop monitoring and alerting for ML-specific metrics including feature drift, prediction drift, train/serve skew, and calibration decay
  • Ensure reproducibility of ML workflows using containerized environments, pinned dependencies, and data snapshots
  • Monitor training and scoring costs across tenants
  • Collaborate with DevOps and SRE engineers on CI/CD and infrastructure automation
  • Prepare operational documentation and platform handover materials
Qualifications
  • 5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems
  • Strong Python skills and experience building platform-level tooling and automation
  • Hands-on experience with Kubernetes and Docker
  • Experience building CI/CD pipelines for ML workloads
  • Hands-on production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platforms
  • Experience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or Kubeflow
  • Strong understanding of ML observability including drift detection, train/serve skew monitoring, and incident response
  • Experience designing or supporting multi-tenant ML systems and isolated model environments
  • Experience working with cloud platforms, preferably GCP
  • Experience with infrastructure-as-code tools such as Terraform
  • Experience with Linux environments
  • Understanding of the ML lifecycle and productionization processes
  • Upper-Intermediate English level or higher

WILL BE A PLUS

  • Experience with feature stores and feature consistency management
  • Experience with large-scale batch scoring systems operating under freshness SLAs
  • Familiarity with experiment tracking platforms and evaluation gates
  • Experience with on-premises Kubernetes or bare-metal Linux infrastructure
  • Knowledge of DVC, lakeFS, or other data versioning tools
  • Experience with Bigtable, Redis, Aerospike, or similar low-latency serving databases
  • GPU scheduling and training cost optimization experience
  • Familiarity with SOC 2, ISO 27001, or GDPR-related compliance requirements
Additional Information

PERSONAL PROFILE

  • Strong ownership mindset and focus on operational reliability
  • Ability to work independently in complex distributed systems environments
  • Strong collaboration and communication skills
  • Analytical thinking with attention to scalability and maintainability
  • Comfortable working in fast-paced product-oriented environments

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior MLOps Platform Engineer: Real-Time ML at Scale
Senior MLOps Platform Engineer: Real-Time ML at Scale

Sigma Software • United States

Remote
USD 140,000 - 190,000
ML Ops Senior Engineer
ML Ops Senior Engineer

Compunnel, Inc. • California (MO)

On-site
USD 120,000 - 160,000
Lead ML Platform Engineer
Lead ML Platform Engineer

Harnham • New York (NY)

Hybrid
USD 150,000 - 190,000
Competitive base salary and annual be
Equity participation through RSUs
Opportunity to work on cutting-edge AI
+2
Senior Machine Learning Engineer (DevOps/SRE)
Senior Machine Learning Engineer (DevOps/SRE)

Roku • Austin (TX)

On-site
USD 120,000 - 150,000
Senior Machine Learning Systems Engineer, Ads ML Experience Platform
Senior Machine Learning Systems Engineer, Ads ML Experience Platform

EngineersOfAI • Northern (KY)

Hybrid
USD 140,000 - 210,000
ML Ops Architect
ML Ops Architect

Tiger Analytics • Dallas (TX)

On-site
USD 120,000 - 150,000
Career development opportunities
Collaborative work environment
Challenging projects
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Machine Learning Engineer Chicago, IL
Senior Machine Learning Engineer Chicago, IL

Attain • Chicago (IL), Northern (KY)

Hybrid
USD 170,000 - 240,000
Senior MLOps Engineer
Senior MLOps Engineer

Hard Rock Hotel & Casino Ottawa • United States

On-site
USD 120,000 - 180,000
Health benefits
Employee wellness programs
Career growth opportunities
Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000