AI Infrastructure Ops Engineer (Kubernetes & GPUs)

Mirantis

United States

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Professional development and training
Conferences and working groups
Company outings and social events
Competitive compensation and benefits

Job summary

Mirantis is seeking an Americas-based AI Infrastructure & Platform Operations engineer to manage expansive AI ecosystems using NVIDIA GPU acceleration, Kubernetes, and cutting-edge platform frameworks. You will operate and optimize production infrastructure across data centers, cloud, and edge environments.

You will collaborate with hardware vendors, engineering teams, and service delivery teams to ensure high availability, performance, and reliability of AI service platforms, while contributing

Qualifications

  • 3+ years in infrastructure/platform/related technical roles.
  • Strong Linux administration and troubleshooting.
  • Good networking concepts understanding.
  • Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem-solving skills.
  • Experience with incident management processes.
  • Excellent communication and collaboration skills.
  • Ability to work in a shift-based operational environment.

Responsibilities

  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform‑related incidents.
  • Support NVIDIA GPU infrastructure and platform services.
  • Monitor and troubleshoot Kubernetes‑based environments.
  • Investigate performance, availability, and reliability issues across components.
  • Collaborate with engineering teams, hardware vendors, DC personnel, and service delivery teams.
  • Participate in incident response, root‑cause analysis, and operational improvement activities.
  • Contribute to improvements in monitoring, observability, automation, and processes.
  • Maintain operational documentation, runbooks, and knowledge articles.

Skills

Linux administration
Networking concepts
Kubernetes
Incident management
Analytical skills
Communication skills
Shift-based work

Tools

Grafana
Prometheus
ELK
OpenTelemetry

Job description

Mirantis is seeking an Americas-based AI Infrastructure & Platform Operations engineer to manage expansive AI ecosystems using NVIDIA GPU acceleration, Kubernetes, and cutting-edge platform frameworks. You will operate and optimize production infrastructure across data centers, cloud, and edge environments.

You will collaborate with hardware vendors, engineering teams, and service delivery teams to ensure high availability, performance, and reliability of AI service platforms, while contributing

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure & Platform Ops Engineer — Remote EU
AI Infrastructure & Platform Ops Engineer — Remote EU

Front Door Defense • Town of Poland (NY)

Remote
USD 60,000 - 67,000
Remote AI Infra & Platform Ops Engineer (EU)
Remote AI Infra & Platform Ops Engineer (EU)

Mirantis, Inc. • Union (NJ)

Remote
USD 60,000 - 67,000
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
AI Infrastructure & Platform Operations Engineer (remote in the EU)
AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis, Inc. • Union (NJ)

Remote
USD 60,000 - 67,000
Senior AI Infra Engineer, Kubernetes LLM Serving & GPU
Senior AI Infra Engineer, Kubernetes LLM Serving & GPU

Mirantis • United States

On-site
USD 180,000 - 230,000
Professional development
Conferences & talks
Hackathons
+2
Senior AI Cloud Deployment Engineer (SRE)
Senior AI Cloud Deployment Engineer (SRE)

Worky • Austin (TX)

On-site
USD 140,000 - 190,000
Competitive compensation
Strong benefits plan
Professional development
Senior Kubernetes DevOps Engineer | AI Infra & Open Source
Senior Kubernetes DevOps Engineer | AI Infra & Open Source

Mirantis • San Jose (CA)

On-site
USD 120,000 - 160,000
Competitive compensation package
Professional development and training
Company outings and hackathons
AI Infrastructure Product Marketing Lead
AI Infrastructure Product Marketing Lead

Mirantis • United States

On-site
USD 120,000 - 180,000
Senior Go Engineer — Remote Kubernetes Storage for AI
Senior Go Engineer — Remote Kubernetes Storage for AI

JobCubby • Northern (KY)

Hybrid
USD 150,000 - 230,000
Competitive compensation and benefits
Professional development
Conference attendance
+1
Senior Storage Systems Engineer - Kubernetes & AI Data
Senior Storage Systems Engineer - Kubernetes & AI Data

Doist • United States

Remote
USD 130,000 - 170,000
Professional development
Conference participation
Competitive compensation