Senior Software Engineer, ML Infrastructure

Voxel Sa

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity via equity plan
Health benefits
Unlimited PTO
Daily meals in-office
Parental leave

Job summary

Voxel, based in San Francisco, seeks a strong software engineer to own the ML infrastructure that powers how we train and ship vision models. You’ll build scalable training systems, manage experiments, and export models to optimized formats like TensorRT/ONNX, collaborating with CV, ML Data, and Platform engineers.

You will establish ML experiment tracking, embrace DevOps-for-ML on AWS, and design scalable solutions that accelerate research-to-production.

Qualifications

  • 4+ years of experience building and shipping large scale software solutions.
  • Hands-on experience building ML training pipelines in PyTorch.
  • Experience with ML experiment tracking tools (Weights & Biases, MLflow, ClearML).
  • Experience with AWS (S3, EC2, EKS) for ML workloads.
  • Strong Python; write production-grade code.
  • Track record of owning infrastructure end-to-end.
  • Bias toward shipping and practical results.

Responsibilities

  • Build and maintain ML training infrastructure to train multiple models concurrently and manage experiments.
  • Own the train-to-deploy handoff—export trained models to optimized formats and quantify impact for production deployment.
  • Establish ML experiment tracking and lifecycle management with tools like Weights & Biases, MLflow, or ClearML.
  • Implement DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring).
  • Design scalable infra solutions that support model development for CV teams.

Skills

ML training pipelines
Python development
DevOps practices
Communication
End-to-end ownership

Tools

PyTorch
Weights & Biases
MLflow
ClearML
AWS

Job description

Who We Are

Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs.

About the Role

Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team.

We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers.

What You'll Do
  • Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures.

  • Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment.

  • Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently.

  • Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely.

  • Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development.

What We're Looking For
  • 4+ years of experience building and shipping large scale software solutions.

  • Hands-on experience building ML training pipelines in PyTorch.

  • Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar).

  • <
  • Experience with AWS (S3, EC2, EKS, or similar) for ML workloads.

  • Strong Python. Write performant code that scales well in production environments.

  • Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on.

  • Bias toward shipping. You'd rather ship something good this week than something perfect next quarter.

  • Strong communication skills.

Nice to Have
  • Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar)

  • Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar)

  • Background in computer vision model training

Compensation & Benefits
  • Equity through Voxel’s Equity Incentive Plan

  • Total compensation includes base salary, annual bonus, and equity

  • Comprehensive health, dental, and vision insurance

  • Competitive paid parental leave

  • Unlimited PTO and flexible work arrangements

  • Daily meals in-office, team events, annual company onsite

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, ML Infrastructure
Senior Software Engineer, ML Infrastructure

Voxel • San Francisco (CA)

On-site
USD 200,000 - 240,000
Equity through Voxel’s Equity Incentive Plan
Health, dental, and vision insurance
Unlimited PTO and flexible work arrangements
+1
Senior Software Engineer, ML Systems
Senior Software Engineer, ML Systems

Socket.dev • San Francisco (CA)

On-site
USD 150,000 - 210,000
Health insurance
Equity plan
PTO
+2
Software Engineer, Perception
Software Engineer, Perception

Voxel • San Francisco (CA)

On-site
USD 100,000 - 130,000
Extensive health, dental, and vision insurance
Generous paid parental leave
Equity Incentive Plan
+3
Tier 3 Support Engineer
Tier 3 Support Engineer

Voxel • United States

On-site
USD 80,000 - 120,000
Health, dental, and vision insurance
Competitive paid parental leave
Ownership through an Equity Incentive Plan
+3
Tier 3 Support Engineer
Tier 3 Support Engineer

Voxel • San Francisco (CA)

On-site
USD 90,000 - 120,000
Health insurance
Paid parental leave
Equity Incentive Plan
+3
Tier 2 Support Engineer
Tier 2 Support Engineer

Voxel • United States

Remote
USD 90,000 - 140,000
Extensive health, dental, and vision保险
Equity Incentive Plan
Generous paid time off
+1
Computer Vision Engineer
Computer Vision Engineer

Voxel51 • United States

Remote
USD 190,000 - 225,000
Equity in the form of options
Variety of benefits
Opportunity for growth in a collaborative environment
Staff ML Ops Engineer
Staff ML Ops Engineer

LVT (LiveView Technologies) • Seattle (WA)

On-site
USD 213,000 - 272,000
Health, dental, and vision coverage
401k with match
Flexible PTO
Staff ML Ops Engineer
Staff ML Ops Engineer

LiveView Technologies • Seattle (WA)

On-site
USD 213,000 - 272,000
Health, dental, vision coverage
401k match up to 4%
Flexible PTO
Senior ML Infra Engineer — Train & Deploy Vision Models
Senior ML Infra Engineer — Train & Deploy Vision Models

Voxel Sa • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Equity via equity plan
Health benefits
Unlimited PTO
+2