MLOps Architect: Scalable Training & Pipelines

Arrayo

Massachusetts

On-site

USD 130,000 - 185,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Arrayo is seeking an MLops Engineer to lead the scaling of machine learning training pipelines and ensure robust end-to-end workflows. The role focuses on Flyte, GPU-optimized Kubernetes, Docker, and distributed training frameworks like Ray to optimize ML infrastructure.

You will orchestrate workflows, scale multi-node GPU training, and collaborate with data scientists to productize experiments while ensuring reproducibility and cost efficiency.

Qualifications

  • Experience with GPU scheduling on Kubernetes and cluster autoscaling.
  • Hands-on with Flyte or similar tools (Airflow, Prefect).
  • Deep knowledge of distributed ML training (e.g., PyTorch DDP, Ray, Horovod).

Responsibilities

  • Develop and maintain ML workflows using Flyte to manage training, testing, and deployment.
  • Scale large-scale ML training systems on GPU-backed Kubernetes clusters with auto-scaling and tuning.
  • Implement distributed model training pipelines using Ray for parallelization and efficiency.
  • Design, build, and optimize Docker images for ML workloads with reproducibility and security.
  • Debug and optimize GPU utilization, memory, and compute bottlenecks during training and inference.
  • Integrate monitoring for ML jobs, track resource consumption, and enforce cost-efficient resource usage.
  • Collaborate with data scientists and ML engineers to productize and scale ML experiments.

Skills

Workflow orchestration
Distributed computing
GPU optimization
ML training
Collaboration with data scientists

Tools

Kubernetes
Flyte
Docker
Ray
PyTorch DDP
Horovod
CUDA
NCCL
GitHub Actions
ArgoCD
Prometheus
Grafana

Job description

Arrayo is seeking an MLops Engineer to lead the scaling of machine learning training pipelines and ensure robust end-to-end workflows. The role focuses on Flyte, GPU-optimized Kubernetes, Docker, and distributed training frameworks like Ray to optimize ML infrastructure.

You will orchestrate workflows, scale multi-node GPU training, and collaborate with data scientists to productize experiments while ensuring reproducibility and cost efficiency.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLops Engineer
MLops Engineer

Arrayo • Massachusetts

On-site
USD 130,000 - 185,000
MLOps Engineer: Scale AI Pipelines & Platforms
MLOps Engineer: Scale AI Pipelines & Platforms

Kensho Technologies • New York (NY)

On-site
USD 120,000 - 150,000
Medical, Dental, and Vision insurance
Unlimited Paid Time Off
26 weeks paid Parental Leave
+7
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Confidential
MLOps Engineer — Scalable AI Infra & Deployment, Equity
MLOps Engineer — Scalable AI Infra & Deployment, Equity

Fundamental • United States

Remote
USD 180,000 - 260,000
Salary + equity
Health coverage for you and dependents
Parental leave for all
+2
MLOps Engineer: Deploy AI at Scale with Kubernetes
MLOps Engineer: Deploy AI at Scale with Kubernetes

Akvelon • United States

Remote
USD 120,000 - 160,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer: Build & Scale Production ML Pipelines
MLOps Engineer: Build & Scale Production ML Pipelines

Remote DXB • United States

Remote
USD 120,000 - 180,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer
MLOps Engineer

Evlo AI • Seattle (WA)

On-site
USD 130,000 - 190,000