Applied Machine Learning Platform Engineer

Buzz Solutions

United States

On-site

USD 80,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Buzz Solutions is seeking an entry/mid-level Applied Machine Learning Platform Engineer to join their computer vision team in the United States. This role involves designing and maintaining scalable training infrastructure, implementing distributed training pipelines, and improving data workflows. Candidates should have 2-4 years of experience in MLOps or backend engineering, along with strong Python skills and familiarity with cloud platforms like AWS and GCP. Join a dynamic team to drive impactful machine learning projects.

Qualifications

  • 2-4 years of industry experience in platform, backend, data, or MLOps engineering roles.
  • Strong software engineering fundamentals including testing and code review.
  • Experience with cloud machine learning infrastructure.

Responsibilities

  • Design, build, and maintain scalable training infrastructure for computer vision workloads.
  • Implement and manage distributed training pipelines.
  • Build and maintain robust data pipelines for ML development.

Skills

Python proficiency
Distributed cloud machine learning infrastructure
Database design for ML workloads
Workflow orchestration and automation

Tools

AWS
GCP
Kubernetes
Docker
Postgres

Job description

About Us

Buzz is revolutionizing the analytics and maintenance of power grid infrastructure through our advanced AI solutions. Our computer vision systems analyze critical infrastructure to enhance safety, reliability, and operational efficiency across the power grid network.

Job Description

We're looking for an entry/mid-level Applied Machine Learning Platform Engineer to join our computer vision team and help improve the databases, cloud infrastructure, and tooling our team builds on. You'll build tooling and infrastructure to help scale our training and data pipelines. You'll work within a team of experienced ML engineers with the autonomy to drive your own projects and the support to keep growing.

Responsibilities
  • Design, build, and maintain scalable training infrastructure for computer vision workloads
  • Implement and manage distributed training pipelines (multi-GPU, multi-node) to support large-scale model training and hyperparameter tuning
  • Build and maintain robust data pipelines for ML development
  • Design database schemas and storage strategies for managing large training datasets, annotations, and model artifacts
  • Implement and manage feature stores, data versioning, and experiment tracking to support reliable model iteration
  • Automate existing analysis workflows
  • Maintain clear documentation for platform components, data contracts, and deployment processes
  • Communicate infrastructure decisions, tradeoffs, and system limitations clearly to ML engineers and stakeholders
  • Conduct thorough code reviews and write integration tests for ML pipelines
Qualifications & Experience
  • 2-4 years of industry experience in platform, backend, data, or MLOps engineering roles
  • Python proficiency — idiomatic code, type hints, async patterns, packaging, and performance-aware implementation
  • Strong software engineering fundamentals — testing, code review, API design, component-level system design
  • Hands-on experience building and operating distributed cloud machine learning infrastructure
  • Designing and maintaining scalable training infrastructure, managing ML platform reliability, optimizing data pipelines for throughput at scale
  • Experience with database design and data systems for ML workloads — schema design, query optimization, and storage strategies for large-scale datasets
  • Excels at workflow orchestration and automation
  • Solid proficiency in Python and core ML tooling:
    • Python ecosystem: Pytest, UV, FastAPI, Pydantic
    • Tooling: Git, Docker, UV
    • Tracking: MLflow, Weights & Biases, or equivalent
    • Automation: Github Actions, CI/CD, Prefect or equivalent
    • Infrastructure: AWS, GCP, Kubernetes, Helm, Terraform or equivalent
    • Databases: Postgres, DynamoDB, Bigtable
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Computer Vision & Machine Learning Engineer
Computer Vision & Machine Learning Engineer

Buzz Solutions • United States

On-site
USD 100,000 - 130,000
Senior Computer Vision & Machine Learning Engineer
Senior Computer Vision & Machine Learning Engineer

Voiceflow • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 210,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
Developer/Engineer
Developer/Engineer

Sbhonline • New York (NY)

On-site
USD 100,000 - 130,000
Machine Learning Engineer
Machine Learning Engineer

ControlRooms.ai • United States

On-site
USD 120,000 - 170,000
Machine Learning Engineer (GCP, Vertex AI, Dataproc, Apache Iceberg)
Machine Learning Engineer (GCP, Vertex AI, Dataproc, Apache Iceberg)

TechDigital Group • Charlotte (NC)

On-site
USD 120,000 - 190,000
ML Platform Engineer
ML Platform Engineer

Jobtailor • North Carolina

On-site
USD 100,000 - 180,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site