Platform Engineer (GPU)

Vero

United States

On-site

USD 144,000 - 176,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
Flexible Spending Account
Flexible PTO
Company-paid Life Insurance

Job summary

Vero is seeking a skilled Platform Engineer (GPU) to ensure the reliability and performance of its cutting-edge GPU clusters tailored for AI and HPC workloads.

This role includes working closely with key technologies like Kubernetes, Terraform, and Ansible, focusing on day-to-day operations and system management within a modern, liquid-cooled data center.

Offering a competitive salary up to $160,000 with equity opportunities, medical benefits, and a flexible PTO policy.

Qualifications

  • 3+ years of experience in Platform Engineering, SRE, DevOps or infrastructure roles.
  • Strong experience with GPU infrastructure & HPC clusters.
  • Proven experience operating and scaling large distributed systems.

Responsibilities

  • Support the reliability and performance of large-scale GPU infrastructure.
  • Optimize Kubernetes platforms for efficiency and stability.
  • Develop reusable Terraform and Ansible modules.
  • Maintain high availability through observability and incident response.
  • Troubleshoot complex issues and manage platform lifecycle.

Skills

Platform Engineering
SRE
DevOps
GPU infrastructure
Terraform
Ansible
Monitoring
Incident response

Tools

Prometheus
Grafana

Job description

Platform Engineer (GPU)

Required by one of the most exciting, well-resourced AI infrastructure startups in the world. The company works in close partnership with NVIDIA and other key organisations shaping the future of data centres and AI infrastructure.

You will play a key role in the operation, optimization, and reliability of large-scale GPU clusters supporting AI/ML and HPC workloads. The focus is on Day 2 operations, performance tuning, and systems management within a state-of-the-art liquid-cooled data center environment.

What we offer
  • Salary up to $160,000 + 20% Bonus
  • Huge equity upside
Responsibilities
  • Support the reliability, performance, and day-to-day operations of large-scale GPU infrastructure supporting AI/ML and HPC workloads
  • Optimize Kubernetes platforms to maximize efficiency, utilization, and stability in production
  • Develop reusable Terraform and Ansible modules to enable scalable, low-drift deployments
  • Maintain high availability through strong observability, SLO/SLI ownership, and incident response practices
  • Troubleshoot complex cross-layer issues and manage platform lifecycle (upgrades, scaling, security, multi-tenancy) in production environments
Requirements
  • 3+ years of experience in Platform Engineering, SRE, DevOps or infrastructure roles
  • Robust experience with GPU infrastructure & HPC clusters
  • Proven experience operating and scaling large distributed systems in high-availability environments
  • Terraform & Ansible
  • Strong background in monitoring, observability and incident response (Prometheus, Grafana, etc.)
Bonus
  • Medical, dental, and vision insurance for the employee and family
  • Equity Scheme
  • Bonus
  • 401(k) with a generous employer match
  • Company-paid Life Insurance
  • Flexible Spending Account
  • Flexible PTO
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
GPU Cluster Architect
GPU Cluster Architect

Jobgether SRL • United States

Remote
USD 184,000 - 318,000
Medical, dental, vision insurance
Remote work reimbursement
RSUs may be available
+3
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Systems and Platform Engineer
Senior Systems and Platform Engineer

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
GPU Platform Engineer for AI/ML Infra
GPU Platform Engineer for AI/ML Infra

Vero • United States

On-site
USD 144,000 - 176,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
Staff Senior Virtualization & Orchestration Engineer
Staff Senior Virtualization & Orchestration Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision insurance
401(k) plan with company match
Paid holidays