Infrastructure Engineer

ITCAPS LLC

St. Louis (MO)

On-site

USD 140,000 - 200,000

Full time

14 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ITCAPS LLC in St. Louis, MO is seeking an AI Infrastructure Engineer to design, build, and operate scalable GPU-enabled Kubernetes platforms for AI/ML workloads. You will shape the on-prem and cloud-native infrastructure to support advanced analytics and model deployment.

The role requires hands-on experience with Kubernetes, Linux (Ubuntu), NVIDIA stacks, and Terraform, plus a passion for automation, reliability, and cost optimization in AI environments.

Qualifications

  • Experience in designing and managing GPU-enabled Kubernetes clusters.
  • Strong Linux (Ubuntu) expertise.
  • Proficient with NVIDIA GPUs and stack.
  • Hands-on with Terraform and IaC.
  • Experience with GPU AI/ML workloads.

Responsibilities

  • Design and manage Kubernetes clusters.
  • Build and manage GPU-enabled infrastructure.
  • Deploy and manage Longhorn storage.
  • Automate infrastructure using Terraform.
  • Monitor platforms with Prometheus and Grafana.
  • Troubleshoot and optimize AI/ML infrastructure.

Skills

Kubernetes
Linux/Ubuntu
NVIDIA stack
Terraform
GPU infrastructure

Tools

Longhorn
Ceph
MAAS/Juju/Charmed Kubernetes
Kubeflow
Prometheus
Grafana

Job description

Role Overview:

Seeking a highly skilled AI Infrastructure Engineer to design, build, and operate scalable GPU-enabled Kubernetes platforms for AI/ML workloads.

Must-Have Skills
  • Strong Kubernetes experience
  • Linux/Ubuntu expertise
  • NVIDIA GPUs and NVIDIA stack
  • Terraform
  • GPU-enabled Kubernetes infrastructure
Responsibilities
  • Design and manage Kubernetes clusters
  • Build and manage GPU-enabled infrastructure
  • Deploy and manage Longhorn storage
  • Automate infrastructure using Terraform
  • Monitor platforms using Prometheus and Grafana
  • Troubleshoot and optimize AI/ML infrastructure
Preferred Skills
  • Longhorn or Ceph
  • Canonical MAAS, Juju, or Charmed Kubernetes
  • Kubeflow or other AI/ML platforms
  • CKA or NVIDIA certifications
Knowledge Transfer & Client Enablement
  • Conduct KT sessions on Kubernetes, GPU infrastructure, NVIDIA stack, and storage
  • Create technical documentation, runbooks, and training materials
  • Conduct hands-on workshops for client teams
  • Advise on cloud-native infrastructure, AI/ML operations, reliability, performance, and cost optimization
  • Ensure smooth production handoff and operational readiness
Soft Skills
  • Strong troubleshooting and problem-solving
  • Excellent communication and documentation
  • Ability to collaborate with ML engineers and data scientists
  • Strong focus on automation and platform scalability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Platform Engineer
Platform Engineer

AMroute LLC • St. Louis (MO)

On-site
USD 15,429,000 - 26,450,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 150,000 - 210,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Platform Engineer (GPU)
Platform Engineer (GPU)

Vero • United States

On-site
USD 136,000 - 160,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000