Production ML Systems Engineer - Scale & Observability

Careervitablr

United States

Remote

USD 140,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Careervitablr seeks a seasoned ML Systems Engineer to design, build, and operate production ML pipelines at scale. You will own model serving infra, CI/CD for ML, and feature pipelines, collaborating with ML scientists to productionize research into reliable services.

You will lead design reviews, mentor engineers, and drive reliability, performance, and cost optimizations across cloud-based ML platforms.

Qualifications

  • 6 to 9 years of software engineering experience, with 4+ years specifically building and operating production ML systems.
  • Strong software engineering fundamentals: Java, Scala, or Python at a production quality bar, plus working knowledge of the other two.
  • Deep experience with cloud infrastructure (AWS preferred) EC2, EKS/Kubernetes, S3, Lambda, IAM and infrastructure-as-code (Terraform or similar).
  • Event-Driven Systems: Experience building and operating streaming/event-driven pipelines (Kafka) and distributed batch processing (Spark).
  • Hands-on experience with model serving frameworks (TorchServe, TensorFlow Serving, Triton) and CI/CD for ML (MLflow, SageMaker Pipelines, Airflow).
  • Security & Compliance: Experience handling PII securely; data governance, access controls (IAM), encryption at rest/in transit.

Responsibilities

  • Design, build, and operate production ML pipelines: feature engineering, training pipelines, model serving, and monitoring at client scale (millions of requests/day).
  • Build and maintain low-latency, high-availability model serving infrastructure for real-time and batch inference.
  • Own CI/CD for ML: automated retraining, model versioning, shadow deployments, canary releases, rollback strategies.
  • Build robust feature pipelines (batch and streaming) using Spark and Kafka, and maintain feature stores for training/serving consistency.
  • Instrument models and pipelines with observability: drift detection, data quality checks, latency/throughput SLAs.
  • Collaborate with ML specialists to productionize research prototypes into tested, scalable services.
  • Optimize training and inference cost/performance (GPU utilization, batching, quantization).
  • Participate in platform-level decisions: build vs. buy for MLOps tooling; lead reliability improvements.

Skills

Java
Scala
Python
Testing
Data quality
Security
Technical leadership
Mentoring
On-call

Tools

Kafka
Spark
Terraform
Docker
MLflow
SageMaker Pipelines
Airflow
Jenkins/GitHub Actions

Job description

Careervitablr seeks a seasoned ML Systems Engineer to design, build, and operate production ML pipelines at scale. You will own model serving infra, CI/CD for ML, and feature pipelines, collaborating with ML scientists to productionize research into reliable services.

You will lead design reviews, mentor engineers, and drive reliability, performance, and cost optimizations across cloud-based ML platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production ML Engineer: Scale & Monitor Real-World Models
Production ML Engineer: Scale & Monitor Real-World Models

Ario Ventures Ltd • United States

Remote
USD 120,000 - 160,000
Senior ML Platform Engineer - Production & MLOps
Senior ML Platform Engineer - Production & MLOps

Attain • Chicago (IL), Northern (KY)

Hybrid
USD 170,000 - 240,000
Senior ML Platform & Infra Engineer - Scale AI Pipelines
Senior ML Platform & Infra Engineer - Scale AI Pipelines

Monograph • United States

Hybrid
USD 160,000 - 240,000
Competitive base pay
Equity (RSUs)
Benefits
Production ML Engineer: Build Scalable ML Systems
Production ML Engineer: Build Scalable ML Systems

Blue Signal Search • United States

On-site
USD 120,000 - 180,000
Health insurance
Dental insurance
Life insurance
+1
Lead ML Engineer: Scale Production AI Systems
Lead ML Engineer: Scale Production AI Systems

Salt Digital Recruitment • United States

On-site
USD 180,000 - 260,000
ML Infrastructure Engineer: Scale Models in Production
ML Infrastructure Engineer: Scale Models in Production

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
Production ML Engineer — Build & Scale AI in NYC
Production ML Engineer — Build & Scale AI in NYC

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Senior ML Engineer: Scale Production ML Pipelines
Senior ML Engineer: Scale Production ML Pipelines

6sense • Northern (KY)

Hybrid
USD 140,000 - 210,000
Health coverage
Paid parental leave
Stock options
+1
Senior ML Platform & MLOps Engineer
Senior ML Platform & MLOps Engineer

Preply • United States

Remote
USD 180,000 - 270,000
Senior ML Systems Engineer: Production Pipelines
Senior ML Systems Engineer: Production Pipelines

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 260,000