ML Platform Engineer

Foundry AI Partners

Northern (KY)

On-site

USD 120,000 - 160,000

Full time

25 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Foundry AI Partners seeks a Platform Engineer to own and scale the ML infrastructure powering model deployment and serving. You will craft scalable platforms, implement robust monitoring, and optimize end-to-end workflows across client environments.

The role emphasizes building CI/CD for training and deployment, reducing latency, and managing costs on GPU/CPU workloads while ensuring security and governance for enterprise clients.

Qualifications

  • 4+ years of experience in platform engineering or DevOps.
  • Strong knowledge of Kubernetes, Docker, and cloud infrastructure (AWS/GCP/Azure).
  • Experience with ML model serving frameworks (TorchServe, TensorFlow Serving, etc.).
  • Proficiency in Python and infrastructure-as-code tools (Terraform, Pulumi).
  • Understanding of ML workflows: training, evaluation, deployment, monitoring.
  • Excellent debugging skills and operational mindset.

Responsibilities

  • Design and maintain ML infrastructure for model deployment and serving
  • Build monitoring and observability systems for production AI agents
  • Implement CI/CD pipelines for model training, evaluation, and deployment
  • Optimize inference latency and cost across GPU/CPU workloads
  • Develop tooling for experiment tracking, versioning, and reproducibility
  • Ensure security, compliance, and data governance for enterprise clients

Skills

Platform engineering
DevOps
Kubernetes
Python
Cloud infrastructure
ML model serving

Education

Bachelor's degree or equivalent in CS/Math/Engineering

Tools

Kubernetes
Docker
Terraform
Pulumi
TorchServe
TensorFlow Serving

Job description

Own the infrastructure that powers our AI deployments. Build scalable platforms for model serving, monitoring, and continuous improvement across client environments.

Responsibilities
  • Design and maintain ML infrastructure for model deployment and serving
  • Build monitoring and observability systems for production AI agents
  • Implement CI/CD pipelines for model training, evaluation, and deployment
  • Optimize inference latency and cost across GPU/CPU workloads
  • Develop tooling for experiment tracking, versioning, and reproducibility
  • Ensure security, compliance, and data governance for enterprise clients
Qualifications
  • 4+ years of experience in platform engineering or DevOps
  • Strong knowledge of Kubernetes, Docker, and cloud infrastructure (AWS/GCP/Azure)
  • Experience with ML model serving frameworks (TorchServe, TensorFlow Serving, etc.)
  • Proficiency in Python and infrastructure-as-code tools (Terraform, Pulumi)
  • Understanding of ML workflows: training, evaluation, deployment, monitoring
  • Excellent debugging skills and operational mindset
Nice to Have
  • Experience with GPU orchestration and optimization
  • Knowledge of vector databases (Pinecone, Weaviate, Qdrant)
  • Familiarity with MLOps tools (MLflow, Weights & Biases, Kubeflow)
  • Background in distributed systems or database internals
  • Contributions to open-source infrastructure projects
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Alexander Chapman Ltd • New York (NY)

On-site
USD 120,000 - 170,000
Health insurance
Competitive equity
Consultant Machine Learning Engineer
Consultant Machine Learning Engineer

Careervitablr • United States

Remote
USD 140,000 - 210,000
ML-Ops / Platform Engineer
ML-Ops / Platform Engineer

Veriipro • Charlotte (NC)

On-site
USD 120,000 - 180,000
MLOps & Agentic Platform Engineer (AI Infrastructure)
MLOps & Agentic Platform Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 110,000 - 150,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
ML Platform Engineer — Scale AI Deployments
ML Platform Engineer — Scale AI Deployments

Foundry AI Partners • Northern (KY)

On-site
USD 120,000 - 160,000
AI/ML Platform Engineer
AI/ML Platform Engineer

Surge IT • Alexandria (VA)

On-site
USD 120,000 - 150,000
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
MLOps & Agentic Platform Engineer (AI Infrastructure)
MLOps & Agentic Platform Engineer (AI Infrastructure)

Hyphen Connect Limited • Seattle (WA)

On-site
USD 100,000 - 130,000