GPU Platform Engineer — Large-Scale Compute & Kubernetes

Bake AI

Hillsboro (OR)

On-site

USD 140,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Bake AI in Hillsboro, Oregon is seeking a hands-on Infrastructure Engineer to build and operate a large-scale GPU compute platform. You will work with research and engineering teams to ensure compute reliability and ease of use while supporting rapid iteration.

The role spans Linux, Kubernetes, networking, automation, observability, and GPU hardware, including troubleshooting Kubernetes workloads and deploying bare-metal servers across multi-node GPU systems.

Qualifications

  • Strong experience administering Linux systems and troubleshooting performance issues.
  • Hands-on experience operating Kubernetes and container infrastructure in a compute-heavy environment.
  • Deep understanding of TCP/IP, VLANs, routing, DNS, NAT, and firewalls.
  • Proficiency in automating infrastructure with Go, Bash, Ansible, or Terraform.
  • Experience with monitoring/observability tools such as Prometheus, Grafana, or Loki.
  • Ability to troubleshoot across hardware, operating systems, networks, containers, and applications.
  • Willingness to work on-site in Hillsboro and handle equipment as needed.

Responsibilities

  • Deploy, operate, expand, and troubleshoot bare-metal NVIDIA and AMD GPU clusters.
  • Build and maintain Kubernetes platforms for training, inference, and research computing.
  • Manage Linux, container runtimes, NVIDIA drivers, CUDA, and the GPU Operator.
  • Configure and troubleshoot VLANs, routing, DNS, firewalls, and high-speed networks.
  • Automate server provisioning, configuration, upgrades, monitoring, and recovery.
  • Build monitoring and alerting for GPUs, servers, networks, storage, and Kubernetes.
  • Improve resource scheduling, GPU utilization, platform reliability, and user isolation.
  • Build internal tools that help research teams access compute and troubleshoot workloads.
  • Perform hands-on rack installation, cabling, BMC/IPMI management, and hardware troubleshooting.
  • Participate in on-call support for critical infrastructure and contribute to incident reviews.

Skills

Linux administration
Kubernetes
Networking (TCP/IP, VLANs, routing,DNS
Automation (Go,Bash,Ansible,Terraform)
Monitoring/observability (Prometheus,G

Tools

NVIDIA drivers
CUDA
GPU Operator

Job description

Bake AI in Hillsboro, Oregon is seeking a hands-on Infrastructure Engineer to build and operate a large-scale GPU compute platform. You will work with research and engineering teams to ensure compute reliability and ease of use while supporting rapid iteration.

The role spans Linux, Kubernetes, networking, automation, observability, and GPU hardware, including troubleshooting Kubernetes workloads and deploying bare-metal servers across multi-node GPU systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU AI Platform Engineer - Kubernetes & DevOps
GPU AI Platform Engineer - Kubernetes & DevOps

AMD • San Jose (CA)

On-site
USD 140,000 - 170,000
AMD benefits
GPU AI Platform Engineer | Kubernetes & DevOps Leader
GPU AI Platform Engineer | Kubernetes & DevOps Leader

Socket.dev • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
SRE / Platform Engineer, GPU Infrastructure
SRE / Platform Engineer, GPU Infrastructure

Bake AI • Hillsboro (OR)

On-site
USD 140,000 - 210,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Kubernetes Engineer for GPU AI Compute Platform
Senior Kubernetes Engineer for GPU AI Compute Platform

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000
Senior Systems Software Engineer - GPU-Driven Kubernetes
Senior Systems Software Engineer - GPU-Driven Kubernetes

NVIDIA AI • Seattle (WA)

On-site
USD 180,000 - 260,000
Equity
Benefits
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute

Kindredventures • United States

On-site
USD 140,000 - 190,000
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5