ML DevOps Architect: Cloud & Large-Scale Compute (Remote)

Ignite Next GmbH

Palo Alto, Northern (CA, KY)

Hybrid

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work
Office visits in Palo Alto, Paris, orW
Flexible relocation options

Job summary

Pathway is seeking a Machine Learning DevOps engineer to optimize ML infrastructure in the cloud, automate ML/LLM pipelines, and support scalable model deployment. The role focuses on reliable, scalable, and auditable ML lifecycle operations across distributed systems.

You will manage GPUs and distributed compute, monitor performance, and collaborate with ML engineers, software engineers, and platform teams to scale production workloads.

Qualifications

  • BSc in Computer Science or Information Technology is required.
  • Hands-on experience in Linux, shell scripting, and cluster management.
  • Strong knowledge of Python and ML pipelines.
  • Proficiency with CI/CD tools and workflows.
  • Experience with ML pipeline orchestration tools (MLflow, Kubeflow, Airflow, Metaflow).
  • Familiarity with cloud services across AWS/GCP/Azure and monitoring stacks.

Responsibilities

  • Optimize infrastructure for ML training and inference (GPUs, distributed compute).
  • Automate and maintain ML/LLM pipelines (data ingestion, training, validation, deployment).
  • Manage model versioning, reproducibility, and traceability.
  • Work with terabyte-scale datasets.
  • Implement ML-centric CI/CD practices.
  • Monitor model performance and data drift in production.

Skills

Linux
Shell scripting
Python
CI/CD
ML pipelines
Cloud platforms
Team collaboration

Education

BSc in Computer Science or Information Technology

Tools

Slurm
Docker
Kubernetes
Terraform
MLflow
Kubeflow
Airflow
Metaflow
Grafana
CloudWatch
Prometheus
Loki

Job description

Pathway is seeking a Machine Learning DevOps engineer to optimize ML infrastructure in the cloud, automate ML/LLM pipelines, and support scalable model deployment. The role focuses on reliable, scalable, and auditable ML lifecycle operations across distributed systems.

You will manage GPUs and distributed compute, monitor performance, and collaborate with ML engineers, software engineers, and platform teams to scale production workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote ML DevOps Engineer – Cloud & Compute Clusters
Remote ML DevOps Engineer – Cloud & Compute Clusters

Pathway Genomics Corp. • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Remote ML DevOps Engineer for Scalable AI Infra
Remote ML DevOps Engineer for Scalable AI Infra

Pathway • Palo Alto (CA)

Hybrid
USD 150,000 - 190,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
ML Ops / Production Engineer — Remote, Equity, Flexible PTO
ML Ops / Production Engineer — Remote, Equity, Flexible PTO

Ent Security • United States

Remote
USD 140,000 - 190,000
Equity
Remote work across North America
Medical, dental, and vision coverage
+4
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Staff ML Engineer - Scalable LLM Infra
Senior Staff ML Engineer - Scalable LLM Infra

Moveworks • Mountain View (CA), Northern (KY)

Hybrid
USD 190,000 - 280,000
Lead ML Cloud DevOps Engineer, Azure & AWS Patterns
Lead ML Cloud DevOps Engineer, Azure & AWS Patterns

London Stock Exchange Group • Creve Coeur (MO)

On-site
USD 120,000 - 180,000
ML Ops Senior Engineer
ML Ops Senior Engineer

Compunnel, Inc. • California (MO)

On-site
USD 120,000 - 160,000