Staff, MLOps Engineer

Sequen AI

United States

Remote

USD 220,000 - 280,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Sequen AI in the United States is seeking a Staff MLOps Engineer to build and operate the critical systems powering its AI models in production, with a focus on ultra-low latency and high-throughput pipelines across multiple cloud environments.

You will collaborate with ML researchers and backend engineers to turn experimental breakthroughs into scalable serving topologies, implement robust CI/CD, and monitor system reliability with real-time telemetry.

Qualifications

  • 4–8+ years in MLOps, ML engineering, or distributed infra.
  • Experience deploying ultra-low-latency ML models under real-time load.
  • Proficiency with Python and PyTorch in production.
  • Comfort with AWS, GCP, or Azure and container orchestration.
  • Experience with CI/CD, model registries, and telemetry.

Responsibilities

  • Build, scale, and operate critical systems for model deployment and real-time telemetry.
  • Move models from experimentation to production, balancing latency and throughput.
  • Implement CI/CD, canary releases, and zero-downtime rollbacks.
  • Architect real-time pipelines to monitor model performance and data drift.
  • Develop evaluation loops to validate live inference accuracy.

Skills

MLOps experience
Low-latency serving
Python
PyTorch
Distributed systems
Cloud platforms (AWS/GCP/Azure)
CI/CD tooling

Tools

Docker
Kubernetes
MLflow
Git
Terraform

Job description

Staff MLOps Engineer — Machine Learning Platform

Location: New York, NY / San Francisco, CA / Remote (US)

Comp Range: $220,000 – $280,000 Base + Performance Bonus + Meaningful Equity

ABOUT US

Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with the infrastructure to run them in production—at sub-10ms latency and enterprise scale. The world’s largest retailers, marketplaces, and travel platforms use Sequen to rank, recommend, and personalize, with an autonomous research engine that compounds model performance into revenue and margin lift measured in hundreds of millions of dollars per customer.

We are a small, highly technical, early-stage team focused on turning recent advances in AI into production-grade systems that operate under unforgiving real-world constraints. The problems we work on are deeply open-ended, where minor optimizations in algorithmic multi-stage retrieval and model routing translate directly into millions of dollars in client revenue.

ABOUT THE ROLE

We are looking for an MLOps Engineer to build, scale, and operate the critical systems that power Sequen’s AI models in production.

This is a foundational, purely infrastructure-focused role sitting at the intersection of machine learning, backend distributed systems, and platform performance. You will not be client-facing; instead, your primary customer will be our internal ML research scientists. Your mission is to make model serving, evaluation, and scaling completely seamless, reliable, and highly optimized in high-throughput production environments.

KEY RESPONSIBILITIES
  • Build ML infrastructure: Design, operate, and maintain robust systems for low-latency model deployment, distributed inference pipelines, and automated real-time telemetry.

  • Scale ranking systems: Move models cleanly from experimentation to production, optimizing the critical trade-offs between execution latency, GPU/CPU throughput, and cloud infrastructure costs.

  • Implement model CI/CD: Build reliable infrastructure for automated model versioning, canary releases, hot‑swappable container rollouts, and zero‑downtime rollbacks.

  • Drive system observability: Architect and monitor real-time pipelines to track model performance, data distribution drift, and system reliability anomalies.

  • Develop evaluation loops: Engineer robust evaluation pipelines and feedback loops to continuously validate live inference accuracy and prevent training‑serving skew.

  • Optimize platform bottlenecks: Proactively isolate and eliminate performance bottlenecks across our serving layers, improving core tooling, model warm‑up times, and researcher velocity.

  • Collaborate with research: Partner closely with our internal ML researchers and backend engineers to translate experimental model breakthroughs into resilient, production‑grade serving topologies.

ABOUT YOU
  • Proven track record: Bring 4–8+ years of practical experience in MLOps, Machine Learning Engineering, or distributed platform/infrastructure engineering.

  • Low‑latency serving expertise: Demonstrate hands‑on experience deploying and serving ultra‑low‑latency machine learning models under heavy, real‑time concurrent workloads.

  • Core ML framework mastery: Maintain deep, production‑grade proficiency with Python and PyTorch.

  • Cloud & container fluency: Operate comfortably across major cloud platforms (AWS, GCP, or Azure) utilizing modern containerization and orchestration tooling (Docker, Kubernetes).

  • Pipeline engineering depth: Show experience designing robust, scalable data pipelines, model registries (e.g., MLflow), and automated CI/CD infrastructures.

  • Systems core maturity: Bring a solid, first‑principles understanding of the complete machine learning lifecycle, asynchronous event‑driven patterns, and distributed systems.

Strong Candidates May Also Bring
  • Rust systems proficiency: Bring production experience or active, hands‑on familiarity with Rust for low‑overhead systems engineering.

  • Generative AI experience: Exposure to serving and optimizing large language models (LLMs) or large‑scale generative model architectures (vLLM, Triton).

  • Modern MLOps tooling: Familiarity with enterprise‑grade feature stores, advanced experiment tracking, and systematic model evaluation frameworks.

  • Startup velocity: Prior experience building and scaling software infrastructure from scratch in fast‑moving, early‑stage, or hypergrowth startups.

WHAT WE VALUE
  • Rigorous systems & scientific thinking: You balance algorithmic complexity with microsecond runtime latency constraints. You choose the right mathematical model for our scaling constraints, prioritizing real‑world stability and performance over theoretical vanity.

  • Uncompromising ownership: You treat production stability and platform efficiency as a personal reflection of code quality, taking pride in building robust, automated pipelines that require zero manual intervention.

  • Pragmatic speed: You possess the startup velocity to design, deploy, and validate robust infra prototypes quickly without accumulating debilitating technical debt.

  • Empowering collaboration: You act as a technical multiplier for our research scientists, building clean developer interfaces, automated workflows, and robust diagnostic tools that help the team move infinitely faster.

WHAT WE OFFER
  • High‑impact influence: A foundational, high‑autonomy role directly shaping the core deployment and serving topology of a category‑defining AI infrastructure company.

  • Pioneering systems: The unique opportunity to build and scale category‑defining, low‑latency ML platforms backed by proven, highly quantified customer revenue results.

  • Top‑tier reward: Highly competitive base salary, uncapped performance metrics, and meaningful early‑employee equity.

  • Premium benefits: Full premium medical/dental/vision coverage, unlimited paid time off, and a highly collaborative, world‑class engineering culture.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Staff MLOps Engineer — Low-Latency ML Platform
Remote Staff MLOps Engineer — Low-Latency ML Platform

Sequen AI • United States

Remote
USD 220,000 - 280,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites
Sequen - Staff Software Engineer – Infrastructure
Sequen - Staff Software Engineer – Infrastructure

fabric • New York (NY)

On-site
USD 250,000 - 350,000
Health insurance
Unlimited time off
Equity
MLOps Engineer
MLOps Engineer

Atomic Machines • Emeryville (CA)

On-site
USD 200,000 - 250,000
Senior ML OPs Engineer
Senior ML OPs Engineer

Glocomms • California (MO)

On-site
USD 198,000 - 230,000
Meal stipends for remote work days
Generous paid time off
Comprehensive health coverage
+1
Lead ML Platform Engineer
Lead ML Platform Engineer

Harnham • New York (NY)

On-site
USD 150,000 - 190,000
Competitive base salary and annual be
Equity participation through RSUs
Opportunity to work on cutting-edge AI
+2
Technical Architect - ML
Technical Architect - ML

Quantiphi • United States

Remote
USD 180,000 - 260,000
MLOps Engineer
MLOps Engineer

Atomic-Machines • Emeryville (CA)

On-site
USD 200,000 - 250,000
Equity
Benefits
Senior ML Ops Engineer | $165K-$175K + Hybrid + Equity | AI Powered Outage Intelligence SaaS Startup
Senior ML Ops Engineer | $165K-$175K + Hybrid + Equity | AI Powered Outage Intelligence SaaS Startup

SmartRecruiters, Inc. • King of Prussia (PA)

Hybrid
USD 165,000 - 175,000
Equity participation
Hybrid work model
Medical, dental, vision benefits
+2
Senior Machine Learning Engineer Chicago, IL
Senior Machine Learning Engineer Chicago, IL

Attain • Chicago (IL), Northern (KY)

Hybrid
USD 170,000 - 240,000