Production ML Engineer: Scalable Pipelines and Low-Latency

Evlo AI

Austin (TX)

On-site

USD 120,000 - 180,000

Full time

10 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evlo AI is seeking an experienced ML Engineer to own end-to-end architecture and deployment of high-throughput machine learning systems. You will design scalable pipelines in Python using PyTorch, deploy and monitor models in cloud environments, and optimize latency for demanding production endpoints.

The role requires 3–6 years in software and ML engineering, strong Python proficiency, and experience with containers and MLOps tools.

Qualifications

  • 3–6 years of professional software engineering experience with at least 3 years dedicated to machine learning engineering.
  • Strong proficiency in Python and deep experience with PyTorch, TensorFlow, or equivalent ML frameworks.
  • Production experience with containerization, orchestration, and MLOps tools like Docker, Kubernetes, and MLflow.
  • Solid foundation in software design principles, API development, and distributed computing.
  • Bonus: Master’s or PhD in Computer Science, ML, or related field, with open-source ML contributions.

Responsibilities

  • Design and implement scalable machine learning pipelines using Python, PyTorch, and distributed data processing frameworks.
  • Deploy, monitor, and scale models in production using cloud infrastructure such as AWS or GCP.
  • Optimize model inference latency, memory footprint, and throughput for high-traffic endpoints.
  • Collaborate with data engineers to establish robust data quality checks across feature stores and training pipelines.
  • Conduct rigorous code reviews, establish engineering best practices, and contribute to system architecture discussions.

Skills

Python
PyTorch
TensorFlow
API development
Distributed computing

Education

Master's or PhD (bonus)

Tools

Docker
Kubernetes
MLflow

Job description

Evlo AI is seeking an experienced ML Engineer to own end-to-end architecture and deployment of high-throughput machine learning systems. You will design scalable pipelines in Python using PyTorch, deploy and monitor models in cloud environments, and optimize latency for demanding production endpoints.

The role requires 3–6 years in software and ML engineering, strong Python proficiency, and experience with containers and MLOps tools.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production ML Engineer: Scalable Pipelines & Uptime
Production ML Engineer: Scalable Pipelines & Uptime

Evlo AI • Miami (FL)

On-site
USD 110,000 - 170,000
Production ML Engineer: Build, Deploy & Monitor AI
Production ML Engineer: Build, Deploy & Monitor AI

Evlo AI • Raleigh (NC)

On-site
USD 120,000 - 180,000
Machine Learning Engineer
Machine Learning Engineer

Evlo AI • Austin (TX)

On-site
USD 120,000 - 180,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
ML Engineer, Production Pipelines & MLOps
ML Engineer, Production Pipelines & MLOps

Pattern AI • Hayward Park (CA)

On-site
USD 90,000 - 130,000
AI/ML Engineer — Production Pipelines & MLOps
AI/ML Engineer — Production Pipelines & MLOps

veritone • Irvine (CA)

On-site
USD 175,000 - 200,000
Machine Learning Engineer
Machine Learning Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision insurance
401(k)
Equity
+2
Senior ML Ops Engineer: Build Scalable ML Pipelines
Senior ML Ops Engineer: Build Scalable ML Pipelines

Myparadigm • Town of Middleton (WI)

On-site
USD 140,000 - 190,000