MLOps Platform Engineer (SageMaker)

TPI Global (formerly Tech Providers, Inc.)

Plano (TX)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

TPI Global in Plano, TX is seeking a Senior Lead AI Platform Engineer to build an enterprise AI platform from the ground up. You will enable internal engineering teams to securely deploy and operate agentic AI workloads at scale, focusing on platform engineering, cloud infrastructure, Kubernetes, DevOps, and end-to-end automation.

You will design production-ready AI infrastructure, establish engineering standards, and drive reliable, governed AI deployments across cloud environments.

Qualifications

  • 8–10 years of software or platform engineering experience, including 7+ years building and operating production systems in AWS, GCP, or hybrid cloud environments.
  • Hands-on experience deploying and supporting AI/ML services in production.
  • 2+ years using AI-assisted development tools such as Claude Code, Codex, or Cursor.
  • Strong experience with Docker, Kubernetes, GitHub Actions, Terraform, and Helm.
  • Experience designing secure authentication and authorization solutions using OAuth 2.0, OIDC, SAML, JWT, RBAC, and IAM.
  • Strong troubleshooting skills across Linux, Kubernetes, networking, containers, and production environments.
  • Experience designing secure, cost-effective infrastructure for hosting large language models (LLMs).

Responsibilities

  • Design and implement secure, scalable AI platform architectures and reusable deployment patterns.
  • Automate infrastructure provisioning, CI/CD pipelines, and platform deployments using Infrastructure as Code.
  • Build and operate highly available AI platform services with monitoring, observability, disaster recovery, and incident response practices.
  • Partner with cloud, DevOps, platform, and security teams to establish AI governance, deployment standards, and operational best practices.
  • Lead technical design reviews, mentor engineers, and promote scalable, secure engineering solutions.
  • Improve platform reliability, deployment consistency, and operational efficiency through automation and standardization.

Skills

Platform engineering
CI/CD
Cloud security
Linux troubleshooting
Networking
RBAC/IAM

Tools

Docker
Kubernetes
GitHub Actions
Terraform
Helm

Job description

Senior Lead AI Platform Engineer Build the Future of Enterprise AI Platforms

In this role you'll build an enterprise AI Platform from the ground up that will enable internal engineering teams to securely deploy and operate agentic AI workloads at scale. The focus is on platform engineering, cloud infrastructure, Kubernetes, DevOps, and end-to-end automation rather than developing AI models. In this role, you'll design and automate production-ready AI infrastructure, establish engineering standards, and help drive reliable, governed AI deployments across cloud environments. You'll work closely with platform, security, and cloud engineering teams while mentoring engineers and shaping best practices for modern AI operations.

Key Responsibilities
  • Design and implement secure, scalable AI platform architectures and reusable deployment patterns.
  • Automate infrastructure provisioning, CI/CD pipelines, and platform deployments using Infrastructure as Code.
  • Build and operate highly available AI platform services with monitoring, observability, disaster recovery, and incident response practices.
  • Partner with cloud, DevOps, platform, and security teams to establish AI governance, deployment standards, and operational best practices.
  • Lead technical design reviews, mentor engineers, and promote scalable, secure engineering solutions.
  • Improve platform reliability, deployment consistency, and operational efficiency through automation and standardization.
Required Qualifications
  • 8–10 years of software or platform engineering experience, including 7+ years building and operating production systems in AWS, GCP, or hybrid cloud environments.
  • Hands-on experience deploying and supporting AI/ML services in production.
  • 2+ years using AI-assisted development tools such as Claude Code, Codex, or Cursor.
  • Strong experience with Docker, Kubernetes, GitHub Actions, Terraform, and Helm.
  • Experience designing secure authentication and authorization solutions using OAuth 2.0, OIDC, SAML, JWT, RBAC, and IAM.
  • Strong troubleshooting skills across Linux, Kubernetes, networking, containers, and production environments.
  • Experience designing secure, cost-effective infrastructure for hosting large language models (LLMs).
Preferred Qualifications
  • Experience with agentic AI frameworks such as LangGraph or Google ADK.
  • Knowledge of AI orchestration, RAG architectures, tool calling, evaluation frameworks, and responsible AI practices.
  • Experience with AI observability tools such as LangSmith, Grafana, or LGTM.
  • Familiarity with GPU infrastructure, Vertex AI, NVIDIA GPU Operator, or model serving platforms.
  • Experience with workflow orchestration tools such as Dagster, Prefect, or Airflow.
  • Strong communication, mentoring, stakeholder management, and technical leadership skills.
Why Join Us
  • Build enterprise-scale AI platforms using modern cloud and automation technologies.
  • Influence engineering standards and best practices for AI infrastructure.
  • Work with cutting-edge AI technologies in a collaborative, innovation-focused environment.
  • Mentor talented engineers while helping shape the future of responsible AI delivery.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Senior AI DevOps Engineer (AI Ops / Platform Engineering)
Senior AI DevOps Engineer (AI Ops / Platform Engineering)

DeepCamp • Tucker (GA)

On-site
USD 96,000 - 165,000
Senior AI Engineer: Scalable LLMs & MLOps Leader
Senior AI Engineer: Scalable LLMs & MLOps Leader

Compunnel, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer – AI Platform
Senior Software Engineer – AI Platform

PRI Technology • New York (NY)

On-site
USD 160,000 - 240,000
AI Platform Engineer
AI Platform Engineer

Apollo Solutions • Boston (MA)

On-site
USD 140,000 - 200,000
Lead Engineer – AI Agent Platform
Lead Engineer – AI Agent Platform

Stellar Consulting Solutions, LLC • Tennessee

On-site
USD 150,000 - 230,000
AI Platform Lead
AI Platform Lead

SeekUp • New York (NY)

On-site
USD 130,000 - 170,000
AI Platform Engineer
AI Platform Engineer

The Aspen Group • Chicago (IL)

On-site
USD 130,000 - 180,000
Paid time off
Health insurance
401(k) with match
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
Senior Software Engineer, AI Platform
Senior Software Engineer, AI Platform

iSolved HCM • United States

On-site
USD 120,000 - 150,000