Senior ML Infrastructure Engineer - Train & Deploy Systems
Voxel
San Francisco (CA)
On-site
USD 200,000 - 240,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Equity through Voxel’s Equity Incentive Plan
Health, dental, and vision insurance
Unlimited PTO and flexible work arrangements
Daily meals in‑office
Job summary
Voxel is seeking a Software Engineer to develop and optimize ML Infrastructure in San Francisco. You will build systems for training multiple vision models, manage ML experiments, and establish best practices on AWS. The role demands strong Python skills and the capacity to handle infrastructure needs effectively. Compensation ranges from $200K to $240K, with a total package including bonuses and equity, comprehensive health benefits, and flexible working arrangements.
Qualifications
4+ years of experience building and shipping large‑scale software solutions.
Hands‑on experience building ML training pipelines in PyTorch.
Strong communication skills.
Responsibilities
Build and maintain training infrastructure for applied ML team.
Establish ML experiment tracking and lifecycle management.
Design scalable solutions for model development.
Skills
Python
ML training pipelines in PyTorch
AWS for ML workloads
DevOps-for-ML best practices
Tools
Weights & Biases
MLflow
ClearML
Job description
Voxel is seeking a Software Engineer to develop and optimize ML Infrastructure in San Francisco. You will build systems for training multiple vision models, manage ML experiments, and establish best practices on AWS. The role demands strong Python skills and the capacity to handle infrastructure needs effectively. Compensation ranges from $200K to $240K, with a total package including bonuses and equity, comprehensive health benefits, and flexible working arrangements.