Consultant Machine Learning Engineer

Careervitablr

United States

Remote

USD 140,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Careervitablr seeks a seasoned ML Systems Engineer to design, build, and operate production ML pipelines at scale. You will own model serving infra, CI/CD for ML, and feature pipelines, collaborating with ML scientists to productionize research into reliable services.

You will lead design reviews, mentor engineers, and drive reliability, performance, and cost optimizations across cloud-based ML platforms.

Qualifications

  • 6 to 9 years of software engineering experience, with 4+ years specifically building and operating production ML systems.
  • Strong software engineering fundamentals: Java, Scala, or Python at a production quality bar, plus working knowledge of the other two.
  • Deep experience with cloud infrastructure (AWS preferred) EC2, EKS/Kubernetes, S3, Lambda, IAM and infrastructure-as-code (Terraform or similar).
  • Event-Driven Systems: Experience building and operating streaming/event-driven pipelines (Kafka) and distributed batch processing (Spark).
  • Hands-on experience with model serving frameworks (TorchServe, TensorFlow Serving, Triton) and CI/CD for ML (MLflow, SageMaker Pipelines, Airflow).
  • Security & Compliance: Experience handling PII securely; data governance, access controls (IAM), encryption at rest/in transit.

Responsibilities

  • Design, build, and operate production ML pipelines: feature engineering, training pipelines, model serving, and monitoring at client scale (millions of requests/day).
  • Build and maintain low-latency, high-availability model serving infrastructure for real-time and batch inference.
  • Own CI/CD for ML: automated retraining, model versioning, shadow deployments, canary releases, rollback strategies.
  • Build robust feature pipelines (batch and streaming) using Spark and Kafka, and maintain feature stores for training/serving consistency.
  • Instrument models and pipelines with observability: drift detection, data quality checks, latency/throughput SLAs.
  • Collaborate with ML specialists to productionize research prototypes into tested, scalable services.
  • Optimize training and inference cost/performance (GPU utilization, batching, quantization).
  • Participate in platform-level decisions: build vs. buy for MLOps tooling; lead reliability improvements.

Skills

Java
Scala
Python
Testing
Data quality
Security
Technical leadership
Mentoring
On-call

Tools

Kafka
Spark
Terraform
Docker
MLflow
SageMaker Pipelines
Airflow
Jenkins/GitHub Actions

Job description

You own the path to production and the production system itself turning modeling work into reliable, scalable, observable services that run at client's traffic volumes.

You partner closely with ML Specialists/Data Scientists, but the accountability for latency, uptime, cost, retraining pipelines, and CI/CD for models is yours.

This role suits someone who is as comfortable in distributed systems and MLOps as in ML itself think software engineer who

specializes in ML systems rather than data scientist who can code.

Design, build, and operate production ML pipelines: feature engineering, training pipelines, model serving, and monitoring at client's scale (millions of requests/day).

Build and maintain low-latency, high-availability model serving infrastructure (real-time inference for search /personalization; batch scoring for pricing/forecasting).

Own CI/CD for ML: automated retraining, model versioning, shadow deployments, canary releases, and rollback strategies.

Build robust feature pipelines (batch and streaming) using Spark and Kafka, and maintain feature stores for training /serving consistency.

Instrument models and pipelines with observability: drift detection, data quality checks, latency/throughput SLAs, and alerting.

Collaborate with ML Specialists to productionize research prototypes translating notebook code into tested, maintainable, scalable services.

Optimize training and inference cost/performance (GPU utilization, batching, caching, model compression /quantization where relevant).

Contribute to platform-level decisions: build vs. buy for MLOps tooling, standards for reproducibility, and shared infrastructure across ML teams.

Participate in on-call rotation for production ML services; drive postmortems and reliability improvements.

Mentor mid-level ML/software engineers and lead design reviews for ML platform and serving infrastructure components.

Must-Have Qualifications

6 to 9 years of software engineering experience, with 4+ years specifically building and operating production ML systems.

Strong software engineering fundamentals: Java, Scala, or Python at a production quality bar (testing, code review, design patterns), plus working knowledge of the other two.

Deep experience with cloud infrastructure (AWS preferred) EC2, EKS/Kubernetes, S3, Lambda, IAM and infrastructure-as-code (Terraform or similar).

Event-Driven Systems: Experience building and operating streaming/event-driven pipelines (Kafka) and distributed batch processing (Spark), including microservices that consume and produce events at scale.

Hands-on experience with model serving frameworks (e.g., TorchServe, TensorFlow Serving, Triton, or custom microservices) and CI/CD for ML (MLflow, SageMaker Pipelines, Airflow, or similar).

Testing & Data Quality: Experience with data validation and testing practices for ML pipelines unit/integration tests for data and models, and contract tests between training and serving to prevent train/serve skew.

Security & Compliance: Experience handling PII and other sensitive data securely within ML pipelines; familiarity with data governance, access controls (IAM policies), and encryption at rest/in transit.

Solid understanding of ML fundamentals (enough to have real technical conversations with data scientists) even if you don't build models from scratch day-to-day.

Technical Leadership: Track record of mentoring engineers and leading design/architecture reviews for distributed or ML-platform systems.

Track record of owning production systems: SLAs, on-call, incident response, capacity planning.

Nice-to-Have

Experience with feature stores (Feast, Tecton, or internal equivalents) and real-time feature computation.

Familiarity with AWS SageMaker, Databricks, or Kubeflow for orchestration.

Experience with GPU infrastructure and inference optimization (quantization, distillation, batching strategies).

Background in search/ranking, personalization, pricing, or fraud systems at marketplace scale.

Experience with A/B testing infrastructure for ML models (shadow traffic, interleaving, canary analysis).

EXPERTISE AND QUALIFICATIONS
Languages

Java/Scala/Python (production services), SQL

Orchestration/Serving

Kubernetes (EKS), Airflow, internal EGAP platform services, TorchServe/TensorFlow Serving or gRPC-based custom serving

Data Platform

Kafka, Spark, S3 data lake, Presto/Trino, Databricks

MLOps

MLflow, SageMaker Pipelines, Terraform, Docker, CI/CD via Jenkins/GitHub Actions

Cloud

AWS (EC2, EKS, Lambda, S3, IAM)

Observability

CloudWatch/Datadog-style metrics, custom drift/data-quality monitors

Collaboration

Jira/Confluence (Atlassian), Git-based workflows

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
ML Ops Engineer
ML Ops Engineer

Bana Solutions • Chantilly (VA)

On-site
USD 150,000 - 210,000
Machine Learning Engineer II, Fulfillment
Machine Learning Engineer II, Fulfillment

Jobtailor • New York (NY)

On-site
USD 120,000 - 170,000
MLOps Engineer
MLOps Engineer

Evlo AI • Seattle (WA)

On-site
USD 130,000 - 190,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Confidential
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
ML Ops Senior Engineer
ML Ops Senior Engineer

Compunnel, Inc. • California (MO)

On-site
USD 120,000 - 160,000
Staff Machine Learning Systems & Reliability Engineer (Moveworks)
Staff Machine Learning Systems & Reliability Engineer (Moveworks)

ServiceNow • Mountain View (CA)

On-site
USD 250,000 - 320,000
Generous family leave
Annual learning stipend
Flexible PTO
+2
MLOps Engineer
MLOps Engineer

Codinix Consulting Services • California (MO)

On-site
USD 120,000 - 150,000