GPU-Accelerated AI Infra & Platform Ops Engineer

Mirantis

United States

Remote

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Professional development
Conferences attendance
Team events

Job summary

PVH (Tommy Hilfiger/Calvin Klein) is building an Americas-based AI Infrastructure & Platform Operations unit to manage large AI ecosystems with NVIDIA GPUs, Kubernetes and cutting-edge frameworks. You will maintain reliability and architectural integrity across a global datacenter footprint.

You will work at the nexus of infrastructure and network engineering, driving automated operational capabilities and contributing to incident response, observability enhancements and runbooks for production

Qualifications

  • 3+ years in infrastructure/operations, platform or SRE roles.
  • Strong Linux admin and troubleshooting skills.
  • Good networking knowledge and incident management experience.
  • Production Kubernetes experience and collaboration across teams.

Responsibilities

  • Monitor, operate and support AI infrastructure platforms.
  • Investigate incidents: infra, networking, hardware and platform issues.
  • Support NVIDIA GPU infrastructure and related services.
  • Triage and resolve issues in Kubernetes-based environments.
  • Collaborate with data center, hardware and engineering teams to fix problems.
  • Participate in incident response and root cause analysis.
  • Improve monitoring, observability, automation and runbooks.
  • Maintain documentation and knowledge articles.

Skills

Linux administration
Networking concepts
Kubernetes production
Incident management
Analytical thinking
Communication

Tools

Grafana
Prometheus
ELK
OpenTelemetry

Job description

PVH (Tommy Hilfiger/Calvin Klein) is building an Americas-based AI Infrastructure & Platform Operations unit to manage large AI ecosystems with NVIDIA GPUs, Kubernetes and cutting-edge frameworks. You will maintain reliability and architectural integrity across a global datacenter footprint.

You will work at the nexus of infrastructure and network engineering, driving automated operational capabilities and contributing to incident response, observability enhancements and runbooks for production

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra & Platform SRE — GPU & Kubernetes
Senior AI Infra & Platform SRE — GPU & Kubernetes

PVH (Tommy Hilfiger/Calvin Klein) • United States

Remote
USD 140,000 - 210,000
GPU-Driven AI Infrastructure & Platform Engineer
GPU-Driven AI Infrastructure & Platform Engineer

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Competitive compensation package
Professional development and training
Conferences and working groups
+1
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events
Senior Data Engineer - Cloud Pipelines & AI Enablement
Senior Data Engineer - Cloud Pipelines & AI Enablement

PVH (Tommy Hilfiger/Calvin Klein) • Boston (MA)

On-site
USD 117,000 - 146,000
Lead Enterprise AI Deployment Engineer
Lead Enterprise AI Deployment Engineer

PVH (Tommy Hilfiger/Calvin Klein) • New York (NY)

On-site
USD 80,000 - 110,000
AI Infrastructure & Platform Engineering Leader
AI Infrastructure & Platform Engineering Leader

Koitecc Solutions • Connecticut

On-site
USD 175,000 - 335,000
Competitive pay
Bonus program
Equity awards
+1
AI Infrastructure Ops Engineer (Kubernetes & GPUs)
AI Infrastructure Ops Engineer (Kubernetes & GPUs)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
InfraOps Engineer - AI GPU Platform
InfraOps Engineer - AI GPU Platform

Lightning-Ai • San Francisco (CA)

Hybrid
USD 160,000 - 200,000
Medical, dental and vision coverage (U
Pension contribution
Generous paid time off
+3
Chief AI Infra & Platform Engineering
Chief AI Infra & Platform Engineering

Koitecc Solutions • Kansas

On-site
USD 175,000 - 335,000
Medical, dental, vision coverage
Retirement savings options
Paid time off