Remote AI Infra & Platform Ops Engineer (EU)

Mirantis, Inc.

Union (NJ)

Remote

USD 60,000 - 67,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mirantis is assembling a European AI Infrastructure & Platform Operations team to operate large-scale AI infrastructure with NVIDIA GPUs, Kubernetes, and high-speed networking. The role focuses on availability, performance, and operational stability across datacenters for AI workloads.

You will work at the intersection of infrastructure, networking, and platform operations, contributing to k0rdent AI development and expanding production platforms in a dynamic environment.

Qualifications

  • 3+ years in infrastructure/platform operations or similar
  • Strong Linux administration and troubleshooting
  • Good networking understanding and issue diagnosis
  • Production Kubernetes experience
  • Experience with incident management processes
  • Excellent communication and collaboration skills
  • Ability to work in a shift-based environment

Responsibilities

  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform incidents.
  • Support NVIDIA GPU infrastructure and related services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Analyze performance, availability, and reliability of infrastructure components.
  • Collaborate with engineering, datacenter, and service teams to resolve issues.
  • Participate in incident response and post-mortem analyses.
  • Improve monitoring, observability, and automation.
  • Maintain runbooks and knowledge articles.

Skills

Linux administration
Networking concepts
Kubernetes in production
Incident management
Analytical problem solving
Communication & collaboration
Shift-based operations

Tools

Grafana
Prometheus
ELK
OpenTelemetry

Job description

Mirantis is assembling a European AI Infrastructure & Platform Operations team to operate large-scale AI infrastructure with NVIDIA GPUs, Kubernetes, and high-speed networking. The role focuses on availability, performance, and operational stability across datacenters for AI workloads.

You will work at the intersection of infrastructure, networking, and platform operations, contributing to k0rdent AI development and expanding production platforms in a dynamic environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure & Platform Ops Engineer — Remote EU
AI Infrastructure & Platform Ops Engineer — Remote EU

Front Door Defense • Town of Poland (NY)

Remote
USD 60,000 - 67,000
AI Infrastructure Ops Engineer (Kubernetes & GPUs)
AI Infrastructure Ops Engineer (Kubernetes & GPUs)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
AI Infrastructure & Platform Operations Engineer (remote in the EU)
AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis, Inc. • Union (NJ)

Remote
USD 60,000 - 67,000
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
Senior Kubernetes DevOps Engineer | AI Infra & Open Source
Senior Kubernetes DevOps Engineer | AI Infra & Open Source

Mirantis • San Jose (CA)

On-site
USD 120,000 - 160,000
Competitive compensation package
Professional development and training
Company outings and hackathons
Senior AI Cloud Deployment Engineer (SRE)
Senior AI Cloud Deployment Engineer (SRE)

Worky • Austin (TX)

On-site
USD 140,000 - 190,000
Competitive compensation
Strong benefits plan
Professional development
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events
Senior AI Infra & Platform Reliability Engineer
Senior AI Infra & Platform Reliability Engineer

United States Digital Space LLC • United States

Remote
USD 140,000 - 190,000
Technical Product Manager for AI Cloud Networking & GPU
Technical Product Manager for AI Cloud Networking & GPU

Doist • United States

Remote
USD 140,000 - 195,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Worky • Austin (TX)

On-site
USD 140,000 - 190,000
Competitive compensation
Strong benefits plan
Professional development