AI Infrastructure Platform Operations Engineer remote in the US

Mirantis

United States

Remote

USD 110,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Professional development
Conferences attendance
Team events

Job summary

PVH (Tommy Hilfiger/Calvin Klein) is building an Americas-based AI Infrastructure & Platform Operations unit to manage large AI ecosystems with NVIDIA GPUs, Kubernetes and cutting-edge frameworks. You will maintain reliability and architectural integrity across a global datacenter footprint.

You will work at the nexus of infrastructure and network engineering, driving automated operational capabilities and contributing to incident response, observability enhancements and runbooks for production

Qualifications

  • 3+ years in infrastructure/operations, platform or SRE roles.
  • Strong Linux admin and troubleshooting skills.
  • Good networking knowledge and incident management experience.
  • Production Kubernetes experience and collaboration across teams.

Responsibilities

  • Monitor, operate and support AI infrastructure platforms.
  • Investigate incidents: infra, networking, hardware and platform issues.
  • Support NVIDIA GPU infrastructure and related services.
  • Triage and resolve issues in Kubernetes-based environments.
  • Collaborate with data center, hardware and engineering teams to fix problems.
  • Participate in incident response and root cause analysis.
  • Improve monitoring, observability, automation and runbooks.
  • Maintain documentation and knowledge articles.

Skills

Linux administration
Networking concepts
Kubernetes production
Incident management
Analytical thinking
Communication

Tools

Grafana
Prometheus
ELK
OpenTelemetry

Job description

Our organization is establishing an Americas-based AI Infrastructure & Platform Operations unit dedicated to the management of expansive AI ecosystems utilizing NVIDIA GPU acceleration, high-speed interconnects, Kubernetes, and bleeding-edge platform frameworks. This team maintains the reliability, efficiency, and architectural integrity of vital AI service platforms across a global datacenter footprint. Positioned at the nexus of core infrastructure and network engineering, you will sustain the high-performance environments essential for contemporary AI application suites. This position offers the chance to engage with pioneering AI hardware while driving the development of automated operational capabilities via the k0rdent AI platform.

Responsibilities
  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, Data Center personnel, and service delivery teams to resolve technical issues.
  • Participate in incident response, root cause analysis, and operational improvement activities.
  • Contribute to improvements in monitoring, observability, automation, and operational processes.
  • Maintain operational documentation, runbooks, and knowledge articles.
Required Experience
  • 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and experience diagnosing infrastructure-related issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem-solving skills.
  • Experience working within structured operational and incident management processes.
  • Excellent communication and collaboration skills.
  • Ability to work within a shift-based operational environment.
Preferred Experience
  • Experience in NVIDIA GPU infrastructure and accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • Kubernetes platform operations.
  • AI infrastructure or HPC environments.
  • Site Reliability Engineering (SRE) or Platform Engineering.
  • Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Large-scale distributed systems and production platforms.
Benefits
  • Professional development and training.
  • Attend conferences and working groups.
  • Company outings, hackathons, and tech talks.
  • Competitive compensation package with a strong benefits plan.
  • Opportunity to work with advanced AI infrastructure environments in production today.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
AI Infrastructure & Platform Operations Engineer (remote in the EU)
AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis, Inc. • Union (NJ)

Remote
USD 60,000 - 67,000
Infrastructure Engineer
Infrastructure Engineer

HCLTech • California (MO)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior Staff Platform Engineer
Senior Staff Platform Engineer

Nvidia Corporation • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Comprehensive benefits
Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

JobCubby • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Principal Software Engineer - Compute Infrastructure
Principal Software Engineer - Compute Infrastructure

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 248,000 - 391,000
Equity
Benefits