ML Infrastructure Engineer

Mach9 Robotics Inc.

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mach9 Robotics Inc. is seeking an ML infrastructure engineer to build and maintain systems powering production AI models for civil engineering and surveying. You will manage training pipelines, data generation, and inference that integrates with CAD software.

The role targets mid-career engineers experienced with training and inference, focusing on scalable, reliable pipelines and real-time serving capabilities for field deployment.

Qualifications

  • 3+ years of work experience in ML infra or related fields.
  • Bachelor's or Master's degree in Computer Science, Engineering, or equivalent experience.
  • Strong communication with ML researchers/engineers to translate workflows into robust systems.
  • Experience designing data/versioning/artifact management or dataset lineage systems (e.g., DVC, LakeFS, Weights & Biases).
  • Hands-on experience with ML pipeline orchestration tools (Airflow, Prefect, Metaflow).
  • Experience with model serving and inference optimization (latency, memory footprint, scaling).
  • Ability to read and refactor ML training code to improve reliability.
  • Proficiency in Python and PyTorch.

Responsibilities

  • Design and build a centralized system for versioning training data, datasets, and model artifacts with full lineage.
  • Develop and maintain reliable ML training and data generation pipelines.
  • Refactor training/data scripts into modular, testable components.
  • Create CI/CD workflows for validating data pipelines and training runs with automated checks.
  • Build tooling to launch, monitor, and debug training jobs with minimal friction.
  • Optimize and scale real-time inference services to meet latency/throughput goals.
  • Own deployment path from trained model to production endpoint with reliable rollouts and monitoring.

Skills

3+ years exp
Communication
Python
PyTorch
Team collaboration
Data versioning
ML pipeline tools
Model serving
Code readability

Education

BSc/MS in CS/Engineering

Tools

Airflow
Prefect
Metaflow
DVC
LakeFS
Weights & Biases
Terraform

Job description

The role

At Mach9, ML infrastructure engineers build and maintain the systems that power production AI models for civil engineering and surveying. Our ML pipeline spans 10,000+ miles of labeled survey data, image segmentation networks, and 3D prediction models serving real-time inference to surveyors and engineers in the field.

This role is ideal for mid-career ML infrastructure engineers with experience building for both training and inference.

You'll build training pipelines that handle deep transformer models on hundreds of terabytes of 3D point cloud and image data. You'll also architect our inference infrastructure, delivering both heavy offline detection algorithms and real-time responsive inference that integrates directly with our CAD software.

Responsibilities
  • Design and build a centralized system for versioning training data, generated datasets, and model artifacts, with full lineage tracking from raw source data through to trained model outputs.
  • Develop and maintain reliable, reproducible ML training and data generation pipelines.
  • Refactor and harden existing training and data generation scripts into composable, testable, and maintainable components.
  • Create CI/CD workflows for validating data pipelines and model training runs, including automated correctness checks and regression detection.
  • Build tooling that enables ML engineers to launch, monitor, and debug training jobs with minimal friction.
  • Optimize and scale real-time model inference services to meet latency and throughput requirements in production, including profiling, batching strategies, and resource-efficient serving.
  • Own the deployment path from trained model artifact to production endpoint, ensuring reliable rollouts, rollback, and monitoring.
Requirements
  • 3+ years of work experience in relevant fields.
  • Bachelor's or Master's degree in Computer Science, Engineering, or equivalent experience.
  • Strong communication skills and the ability to work closely with ML researchers and engineers to understand their workflows and translate them into robust systems.
  • Experience designing and building data versioning, artifact management, or dataset lineage systems (e.g., DVC, LakeFS, Weights & Biases, or custom solutions).
  • Hands‑on experience with ML pipeline orchestration tools (e.g., Airflow, Prefect, Metaflow, or similar).
  • Experience with model serving and inference optimization — profiling latency, reducing memory footprint, or scaling serving infrastructure to meet real‑time constraints.
  • Ability to read and refactor ML training code — you don't need to design model architectures, but you need to understand what training pipelines are doing well enough to make them reliable.
  • Proficient with Python, PyTorch.
Bonus qualifications
  • Familiarity with AWS infrastructure services.
  • Experience with containerized ML workflows and GPU-accelerated training environments.
  • Experience with model optimization techniques (e.g., quantization, TensorRT, ONNX Runtime, distillation).
  • Knowledge of infrastructure-as-code tools (e.g., AWS CDK, Terraform).
  • Experience building or operating ML systems that handle large unstructured datasets (imagery, 3D data, sensor data).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer
ML Engineer

Tensec • San Francisco (CA)

On-site
USD 120,000 - 150,000
ML Engineer, Research
ML Engineer, Research

Quiet Capital • San Francisco (CA)

On-site
USD 150,000 - 190,000
ML Infrastructure Engineer: Training & Real-Time Inference
ML Infrastructure Engineer: Training & Real-Time Inference

Mach9 Robotics Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Engineer
ML Engineer

Mach9 • San Francisco (CA)

On-site
USD 100,000 - 140,000
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Alexander Chapman Ltd • New York (NY)

On-site
USD 120,000 - 170,000
Health insurance
Competitive equity
ML Engineer, Product
ML Engineer, Product

Mach9 Robotics Inc. • San Francisco (CA)

On-site
USD 120,000 - 190,000
Head of ML
Head of ML

Tensec • San Francisco (CA)

On-site
USD 150,000 - 200,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000