Platform Engineer (GPU)

Vero

United States

On-site

USD 144,000 - 176,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
Flexible Spending Account
Flexible PTO
Company-paid Life Insurance

Job summary

Vero is seeking a skilled Platform Engineer (GPU) to ensure the reliability and performance of its cutting-edge GPU clusters tailored for AI and HPC workloads.

This role includes working closely with key technologies like Kubernetes, Terraform, and Ansible, focusing on day-to-day operations and system management within a modern, liquid-cooled data center.

Offering a competitive salary up to $160,000 with equity opportunities, medical benefits, and a flexible PTO policy.

Qualifications

  • 3+ years of experience in Platform Engineering, SRE, DevOps or infrastructure roles.
  • Strong experience with GPU infrastructure & HPC clusters.
  • Proven experience operating and scaling large distributed systems.

Responsibilities

  • Support the reliability and performance of large-scale GPU infrastructure.
  • Optimize Kubernetes platforms for efficiency and stability.
  • Develop reusable Terraform and Ansible modules.
  • Maintain high availability through observability and incident response.
  • Troubleshoot complex issues and manage platform lifecycle.

Skills

Platform Engineering
SRE
DevOps
GPU infrastructure
Terraform
Ansible
Monitoring
Incident response

Tools

Prometheus
Grafana

Job description

Platform Engineer (GPU)

Required by one of the most exciting, well-resourced AI infrastructure startups in the world. The company works in close partnership with NVIDIA and other key organisations shaping the future of data centres and AI infrastructure.

You will play a key role in the operation, optimization, and reliability of large-scale GPU clusters supporting AI/ML and HPC workloads. The focus is on Day 2 operations, performance tuning, and systems management within a state-of-the-art liquid-cooled data center environment.

What we offer
  • Salary up to $160,000 + 20% Bonus
  • Huge equity upside
Responsibilities
  • Support the reliability, performance, and day-to-day operations of large-scale GPU infrastructure supporting AI/ML and HPC workloads
  • Optimize Kubernetes platforms to maximize efficiency, utilization, and stability in production
  • Develop reusable Terraform and Ansible modules to enable scalable, low-drift deployments
  • Maintain high availability through strong observability, SLO/SLI ownership, and incident response practices
  • Troubleshoot complex cross-layer issues and manage platform lifecycle (upgrades, scaling, security, multi-tenancy) in production environments
Requirements
  • 3+ years of experience in Platform Engineering, SRE, DevOps or infrastructure roles
  • Robust experience with GPU infrastructure & HPC clusters
  • Proven experience operating and scaling large distributed systems in high-availability environments
  • Terraform & Ansible
  • Strong background in monitoring, observability and incident response (Prometheus, Grafana, etc.)
Bonus
  • Medical, dental, and vision insurance for the employee and family
  • Equity Scheme
  • Bonus
  • 401(k) with a generous employer match
  • Company-paid Life Insurance
  • Flexible Spending Account
  • Flexible PTO
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Solution Architect - AI Infrastructure
Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 283,500 - 346,500
Equity (RSUs)
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Customer Solution Architect - Systems Integrator
Customer Solution Architect - Systems Integrator

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 225,000 - 275,000
RSU equity
20% bonus
HPC Engineer - AI Infrastructure
HPC Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership