ML Inference Infrastructure Engineer — Scale & GPU

Talanto

Northern (KY)

Hybrid

USD 221,000 - 260,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous Time Off
Comprehensive Health Plans
Paid Parental Leave
Family Forming Benefits
401(k) Matching
Personal Device Allowance
Pre-tax Benefits
Lifestyle Wallet
Mental Health Support
Sabbatical Leave
Compensation and Equity

Job summary

Talanto in the United States is seeking an ML Infrastructure Engineer, Model Inference to build and optimize our core inference infrastructure powering AI models. You will work with Infrastructure and Research to deploy, optimize, and orchestrate AI models across scalable systems.

The role requires 5+ years in production ML, strong Kubernetes skills, and experience with model serving frameworks and GPU optimization. Hybrid options may apply with a focus on scalable backend infrastructure.

Qualifications

  • 5+ years of experience building and deploying ML models in production.
  • Deep understanding of container orchestration and distributed systems.
  • Expertise in Kubernetes administration, CRDs, operators, and cluster management.

Responsibilities

  • Design, deploy and maintain scalable Kubernetes clusters for AI model inference and training.
  • Develop, optimize, and maintain ML model serving infrastructure with high performance and low latency.
  • Scale backend infrastructure for AI-driven products, focusing on deployment and throughput.
  • Optimize compute-heavy workflows and GPU utilization for ML workloads.
  • Build a robust model API orchestration system.
  • Collaborate with leadership to scale infrastructure as the company grows.

Skills

Container orchestration
Distributed systems
API development
Communication skills

Tools

Kubernetes
NVIDIA Triton Server
PyTorch
TensorFlow
Terraform
Ansible
GitOps
CUDA optimization

Job description

Talanto in the United States is seeking an ML Infrastructure Engineer, Model Inference to build and optimize our core inference infrastructure powering AI models. You will work with Infrastructure and Research to deploy, optimize, and orchestrate AI models across scalable systems.

The role requires 5+ years in production ML, strong Kubernetes skills, and experience with model serving frameworks and GPU optimization. Hybrid options may apply with a focus on scalable backend infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Infrastructure Engineer
ML Inference Infrastructure Engineer

Baseten • United States

Remote
USD 120,000 - 190,000
Equity
Medical coverage for employee and dep.
Flexible PTO including Winter Break
+3
ML Inference Platform Engineer — Scale Production AI
ML Inference Platform Engineer — Scale Production AI

The Consensus • New York (NY)

On-site
USD 120,000 - 150,000
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
+2
ML Infra Engineer — Scale GPU ML Platform & Equity
ML Infra Engineer — Scale GPU ML Platform & Equity

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
ML Infrastructure Engineer - Model Inference & Scale
ML Infrastructure Engineer - Model Inference & Scale

Abridge • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Generous Time Off
Comprehensive Health Plans
401(k) Matching
+2
Senior ML Engineer - Remote, Lead Production-Scale Models
Senior ML Engineer - Remote, Lead Production-Scale Models

Talanto • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 250,000
Fully remote
Performance-based bonus
Learning & development stipend
+2
ML Inference Engineer — Scale & Reliability
ML Inference Engineer — Scale & Reliability

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1