Machine Learning Ops & Infrastructure Engineer

Noble Machines

Sunnyvale (CA)

On-site

USD 160,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Noble Machines in Sunnyvale is seeking an experienced ML Ops & Infrastructure Engineer to develop the foundational systems essential for AI development. This critical role intersects both Research and Engineering teams, architecting high-performance ML infrastructure for scalable and reliable machine learning.

The ideal candidate will have over three years of experience in the industry, expertise in cloud platforms (AWS or GCP), and strong programming skills. The position offers a base salary range of $160,000 – $300,000 along with additional bonuses and benefits.

Qualifications

  • 3+ years of hands-on industry experience building scalable ML infrastructure.
  • Deep expertise in AWS or GCP and modern orchestration tools.
  • Strong programming skills in Python and experience with Git.

Responsibilities

  • Design and maintain scalable ML infrastructure.
  • Manage data ingestion and processing pipelines.
  • Optimize environments for distributed model training.

Skills

Cloud platforms expertise (AWS or GCP)
Kubernetes (K8s)
Python programming
Data management pipelines
Containerization

Tools

Docker
Airflow
Terraform

Job description

About Noble Machines

Noble Machines (formerly Under Control Robotics) builds multipurpose robots to support human workers in the world's toughest jobs—turning dangerous work from a necessity into a choice. Our work demands reliability, robustness, and readiness for the unexpected—on time, every time. We're assembling a mission‑driven team focused on delivering real impact in heavy industry, from construction and mining to energy. If you're driven to build rugged, reliable products that solve real‑world problems, we'd love to talk.

Position Overview

At Noble Machines AI we are pushing the boundaries of machine learning and artificial intelligence. We are looking for an experienced ML Ops & Infrastructure Engineer to build the foundational systems that power our AI development. In this role you will sit at the critical intersection of our Research and Engineering teams. You won’t just be maintaining systems; you will be architecting the high‑performance ML infrastructure that enables our researchers to seamlessly transition from data collection and model training to evaluation and production. If you are passionate about scalable compute, elegant data platforms, and robust deployment pipelines, we want you on our team.

Responsibilities
  • End‑to‑End ML Infrastructure: Design, build, and maintain a highly scalable and reliable machine learning infrastructure that accelerates the research and development lifecycle.
  • Data Platform & Management: Architect and manage robust data ingestion, collection, and processing pipelines. Own the data platforms that ensure our models are trained on high‑quality, perfectly versioned datasets.
  • Training & Evaluation Pipelines: Build and optimize the environments used for distributed model training, hyperparameter tuning, and automated model evaluation.
  • Cloud Compute Orchestration: Manage and orchestrate heavy compute workflows across AWS and/or Google Cloud Platform (GCP), optimizing for both performance and cost.
  • Containerization & Kubernetes: Take full ownership of containerizing ML workloads and orchestrating them via Kubernetes (K8s) to ensure high availability, scalability, and reproducibility.
  • Cross‑Functional Collaboration: Partner closely with ML Researchers and Software Engineers to understand their bottlenecks, gather requirements, and build tooling that makes their workflows frictionless.
Requirements
  • Proven Industry Experience: 3+ years of hands‑on industry experience building scalable ML infrastructure, MLOps platforms, or data engineering systems.
  • Cloud & Orchestration Mastery: Deep expertise in cloud platforms (AWS or GCP) and modern orchestration tools, specifically Docker and Kubernetes (K8s).
  • Software Engineering Fundamentals: Strong programming skills in Python, alongside experience with bash scripting and version control (Git).
  • Data & Pipeline Expertise: Hands‑on experience building large‑scale data management pipelines and using workflow orchestration tools (e.g., Airflow, Argo, Kubeflow, or similar).
  • Relevant Domain Background: While explicit robotics experience is not required, we highly value candidates with backgrounds in hardware‑interfacing AI, autonomous driving, computer vision, or other high‑complexity ML fields.
Nice to Have
  • Experience with Infrastructure as Code (IaC) tools like Terraform.
  • Familiarity with distributed training frameworks (e.g., PyTorch DDP, Horovod, Ray).
  • Experience implementing model observability, monitoring, and data drift detection in production environments.
  • A background handling large volumes of unstructured data (video, sensor data, spatial data).

The base salary range for this full‑time position is $160,000 – $300,000, in addition to bonus, equity and benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops & Infra Engineer — Scale AI Pipelines
ML Ops & Infra Engineer — Scale AI Pipelines

Noble Machines • Sunnyvale (CA)

On-site
USD 160,000 - 300,000
Senior ML Ops Engineer (Machine Learning Infrastructure)
Senior ML Ops Engineer (Machine Learning Infrastructure)

PVH (Tommy Hilfiger/Calvin Klein) • Los Angeles (CA)

Hybrid
USD 150,000 - 250,000
1.6 Machine Learning Operations Engineer
1.6 Machine Learning Operations Engineer

Field AI • Mission Viejo (CA)

On-site
USD 70,000 - 300,000
Software Engineer, MLOps
Software Engineer, MLOps

Medium • Irvine (CA)

On-site
Competitive salary
Flexible work arrangement
Collegial team environment
Software Engineer, MLOps
Software Engineer, MLOps

FieldAI • Irvine (CA)

On-site
Competitive salary
Flexible work options
Diverse and inclusive environment
Machine Learning Engineer, Applied AI Infrastructure
Machine Learning Engineer, Applied AI Infrastructure

Orbifold AI • Palo Alto (CA)

On-site
USD 170,000 - 230,000
AI Engineer, AIOps & Infrastructure
AI Engineer, AIOps & Infrastructure

eloquentai • San Francisco (CA)

On-site
USD 130,000 - 160,000
ML Infrastructure Engineer
ML Infrastructure Engineer

echo inc. • San Francisco (CA)

On-site
USD 180,000 - 230,000
Stock options
Competitive compensation
401(k) program with matching
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Machine Learning Engineer, Applied AI Infrastructure
Machine Learning Engineer, Applied AI Infrastructure

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 210,000