AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis, Inc.

Union (NJ)

Remote

USD 60,000 - 67,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mirantis is assembling a European AI Infrastructure & Platform Operations team to operate large-scale AI infrastructure with NVIDIA GPUs, Kubernetes, and high-speed networking. The role focuses on availability, performance, and operational stability across datacenters for AI workloads.

You will work at the intersection of infrastructure, networking, and platform operations, contributing to k0rdent AI development and expanding production platforms in a dynamic environment.

Qualifications

  • 3+ years in infrastructure/platform operations or similar
  • Strong Linux administration and troubleshooting
  • Good networking understanding and issue diagnosis
  • Production Kubernetes experience
  • Experience with incident management processes
  • Excellent communication and collaboration skills
  • Ability to work in a shift-based environment

Responsibilities

  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform incidents.
  • Support NVIDIA GPU infrastructure and related services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Analyze performance, availability, and reliability of infrastructure components.
  • Collaborate with engineering, datacenter, and service teams to resolve issues.
  • Participate in incident response and post-mortem analyses.
  • Improve monitoring, observability, and automation.
  • Maintain runbooks and knowledge articles.

Skills

Linux administration
Networking concepts
Kubernetes in production
Incident management
Analytical problem solving
Communication & collaboration
Shift-based operations

Tools

Grafana
Prometheus
ELK
OpenTelemetry

Job description

AI Infrastructure & Platform Operations Engineer (remote in the EU)
  • Full-time

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on‑premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI‑driven workloads, Mirantis delivers the automation, GPU orchestration, and policy‑driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock‑in, Mirantis ensures that customers retain full control of their infrastructure strategy.

Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Société Generale, Splunk, and Volkswagen. Learn more at www.mirantis.com.

We are building a European AI Infrastructure & Platform Operations team responsible for operating large‑scale AI infrastructure environments powered by NVIDIA GPUs, high‑performance networking, Kubernetes, and next‑generation platform technologies.

The team is responsible for ensuring the availability, performance, and operational stability of critical AI infrastructure platforms deployed across multiple datacenters. Working at the intersection of infrastructure, networking, and platform operations, you will help support the environments that power modern AI workloads.

This is an opportunity to work with some of the latest technologies in AI infrastructure while contributing to the evolution of AI‑powered operational services through platforms such as k0rdent AI.

Responsibilities
  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform‑related incidents.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Monitor and troubleshoot Kubernetes‑based environments.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues.
  • Participate in incident response, root cause analysis, and operational improvement activities.
  • Contribute to improvements in monitoring, observability, automation, and operational processes.
  • Maintain operational documentation, runbooks, and knowledge articles.
Requirements
  • 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and experience diagnosing infrastructure‑related issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem‑solving skills.
  • Experience working within structured operational and incident management processes.
  • Excellent communication and collaboration skills.
  • Ability to work within a shift‑based operational environment.
Highly desirable skills
  • NVIDIA GPU infrastructure and accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • Kubernetes platform operations.
  • AI infrastructure or HPC environments.
  • Site Reliability Engineering (SRE) or Platform Engineering.
  • Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Large‑scale distributed systems and production platforms.
What does Mirantis offer you?
  • Work with some of the most advanced AI infrastructure environments in production today.
  • Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high‑performance networking environments.
  • Help define how next‑generation AI infrastructure is operated and supported.
  • Be part of a team shaping the future of AI‑powered operations through k0rdent AI.
  • Join a growing organisation investing heavily in AI infrastructure and platform services.
  • Salary range: $60000-$67000 gross per year

It is understood that Mirantis, Inc. may use automated decision‑making technology (ADMT) for specific employment‑related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

By clicking the link above or any third‑party link within this posting, you are leaving this site and going to a third‑party website where the third‑party website's terms and privacy policy apply

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Worky • Austin (TX)

On-site
USD 140,000 - 190,000
Competitive compensation
Strong benefits plan
Professional development
Remote AI Infra & Platform Ops Engineer (EU)
Remote AI Infra & Platform Ops Engineer (EU)

Mirantis, Inc. • Union (NJ)

Remote
USD 60,000 - 67,000
AI Infrastructure & Platform Ops Engineer — Remote EU
AI Infrastructure & Platform Ops Engineer — Remote EU

Front Door Defense • Town of Poland (NY)

Remote
USD 60,000 - 67,000
AI Infrastructure Ops Engineer (Kubernetes & GPUs)
AI Infrastructure Ops Engineer (Kubernetes & GPUs)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
Technical Product Marketer, k0rdent AI - remote in the US
Technical Product Marketer, k0rdent AI - remote in the US

Mirantis • United States

On-site
USD 90,000 - 150,000
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events
Product Manager - AI Inference & Model Serving
Product Manager - AI Inference & Model Serving

Mirantis • Austin (TX)

On-site
USD 120,000 - 160,000
Professional development and training
Customized workstation
Competitive compensation package
Senior Kubernetes DevOps Engineer - K0rdent AI Apps/Core Services
Senior Kubernetes DevOps Engineer - K0rdent AI Apps/Core Services

Mirantis • San Jose (CA)

On-site
USD 120,000 - 160,000
Competitive compensation package
Professional development and training
Company outings and hackathons
Enterprise Architect (Professional Services)
Enterprise Architect (Professional Services)

Mirantis • Northern (KY)

Hybrid
USD 150,000 - 210,000
Silicon Valley leader
Professional development and training
Conferences and tech talks
+1