Machine Learning Operations

voltai-com

Edison (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
Professional Growth
Visa Sponsorship

Job summary

Voltai is seeking engineers to design, build, and operate scalable ML pipelines for training, evaluation, and deployment of LLMs and retrieval-augmented systems. You will automate the ML lifecycle, monitor model quality, and optimize performance across diverse hardware and cloud environments.

Join a team blending software and hardware expertise to productionize foundation models for enterprise customers, while balancing latency, cost, and security in fast-paced settings.

Qualifications

  • Experience building reliable, scalable ML systems.
  • End-to-end ML lifecycle management experience.
  • Ability to translate research into production-ready systems.

Responsibilities

  • Design, build, and maintain scalable ML pipelines for training, evaluation, and deployment of LLMs.
  • Operationalize evaluation workflows using synthetic and human-labeled data across deployments.
  • Automate the ML Developer lifecycle with data versioning and CI/CD pipelines.
  • Optimize model training and inference for latency, throughput, and cost.
  • Collaborate cross-functionally to productionize foundation models.
  • Deploy and manage both open-source and proprietary models with latency, security, and compliance constraints.
  • Implement real-time monitoring to detect drift and bottlenecks in live systems.
  • Work directly with enterprise customers on deployment strategies and feedback loops.

Skills

Python
Go
Rust
ML Ops
System design
Cross-functional
Performance

Tools

MLflow
Kubeflow
SageMaker
Vertex AI
Apache Airflow
Docker
Kubernetes
Pulumi
Terraform
Weights & Biases
CometML

Job description

About Voltai

Voltai is the leading AI company building agentic systems and frontier foundation models for semiconductor and electronics design. Backed by Sequoia Capital, we’re putting AI in the hands of hardware engineers in over 70% of the world’s largest semiconductor and electronics companies to have effortless control over their next-generation chip and board designs, powering the future of automotive, industrial automation, consumer electronics, IoT, and semiconductor manufacturing.

About the Team

Our founding team consists of IOI/IPhO olympiad medalists, Stanford professors, ex-CTO of Synopsys, and our business leadership has scaled revenue in their previous companies to over $1.5bn. At Voltai, we are combining the world’s best talent in the intersection of software and hardware.

Key Responsibilities
  • Design, build, and maintain scalable ML pipelines for training, evaluation, and deployment of LLMs and retrieval-augmented systems, optimized for performance, traceability, and reproducibility
  • Operationalize evaluation workflows using both synthetic and human-labeled datasets to monitor model quality at scale across multiple downstream tasks and customer deployments
  • Automate the ML Developer lifecycle by implementing robust data versioning, model tracking, and CI/CD pipelines using modern ML Ops tooling
  • Optimize model training and inference, focusing on reducing latency, maximizing throughput, and controlling cost across heterogeneous hardware environments.
  • Collaborate cross-functionally with research, infrastructure, and product teams to productionize foundation models and integrate them into customer-facing AI products
  • Deploy and manage both open-source and proprietary models within stringent constraints on latency, security, and compliance—balancing reliability with innovation.
  • Implement real-time monitoring and alerting systems to detect model/data drift, quality regressions, and infrastructure bottlenecks in live environments.
  • Work directly with enterprise customers, supporting deployment strategies, ensuring production readiness, and creating tight feedback loops from real-world usage to continuous model improvement
Required Skill Sets
  • Software Engineering Expertise: Proven experience in building reliable and scalable systems, with a strong foundation in software engineering principles and expertise with Python, Go, or Rust
  • ML Ops Platforms: Hands‑on experience with ML Ops platforms such as MLflow, Kubeflow, SageMaker, Vertex AI, or Apache Airflow, facilitating efficient model lifecycle management
  • Cloud-Native Tools: Proficiency in cloud-native tools including Docker, Kubernetes, storage optimizers. Experience with major cloud providers like AWS, GCP or other compute providers for deploying and managing ML workloads with a focus on cost optimization
  • Experiment Tracking & Model Management: Proficiency in tools like Weights & Biases, MLflow, and/or CometML for tracking experiments, managing model metadata, and facilitating collaboration
  • Infrastructure: Familiarity with infrastructure tools like Terraform, Pulumi, and Chronosphere for monitoring and alerting.
  • Demonstrated ability to translate research findings into robust, production‑ready systems, bridging the gap between experimentation and deployment.
Bonus Points
  • Some background in hardware/electronics, gained through professional, academic, or personal projects
  • Contributions to open‑source initiatives
  • Notable awards or publications in leading journals/conferences
  • Experience thriving in a fast‑paced, hyper‑growth startup environment
Our Benefits
  • Unlimited PTO: Recharge when you need it, no questions asked.
  • Comprehensive Health Coverage: Medical, dental, and vision insurance for you and your dependents.
  • Free Meals and Snacks: Daily lunches, dinners, and snacks in the office.
  • Professional Growth: We invest in your continuous learning and offer opportunities to expand your skills.
  • Visa Sponsorship: We welcome global talent and provide visa sponsorship to support qualified candidates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Systems Engineer
Machine Learning Systems Engineer

voltai-com • Edison (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Machine Learning Research Engineer
Machine Learning Research Engineer

voltai-com • Edison (CA)

On-site
USD 180,000 - 250,000
Unlimited PTO
Health coverage
Meals & snacks
+2
Software Engineer
Software Engineer

voltai-com • Edison (CA)

On-site
USD 140,000 - 210,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Machine Learning Research Scientist
Machine Learning Research Scientist

voltai-com • Edison (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Software Engineering - Data Engineer
Software Engineering - Data Engineer

voltai-com • Edison (CA)

On-site
USD 140,000 - 220,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Strategic Projects Lead
Strategic Projects Lead

voltai-com • Edison (CA)

On-site
USD 120,000 - 180,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Software Engineer - Backend Engineer
Software Engineer - Backend Engineer

voltai-com • Edison (CA)

On-site
USD 180,000 - 240,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Chief of Staff
Chief of Staff

voltai-com • Edison (CA)

On-site
USD 120,000 - 190,000
Unlimited PTO
Health coverage
Free meals and snacks
+2
Applied AI Engineer
Applied AI Engineer

voltai-com • Edison (CA)

On-site
USD 140,000 - 210,000
Unlimited PTO
Health coverage
Free meals & snacks
+2
Software Engineering - Infrastructure
Software Engineering - Infrastructure

voltai-com • Edison (CA)

On-site
USD 130,000 - 170,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2