SRE / Platform Engineer, GPU Infrastructure

Bake AI

Hillsboro (OR)

On-site

USD 140,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bake AI in Hillsboro, Oregon is seeking a hands-on Infrastructure Engineer to build and operate a large-scale GPU compute platform. You will work with research and engineering teams to ensure compute reliability and ease of use while supporting rapid iteration.

The role spans Linux, Kubernetes, networking, automation, observability, and GPU hardware, including troubleshooting Kubernetes workloads and deploying bare-metal servers across multi-node GPU systems.

Qualifications

  • Strong experience administering Linux systems and troubleshooting performance issues.
  • Hands-on experience operating Kubernetes and container infrastructure in a compute-heavy environment.
  • Deep understanding of TCP/IP, VLANs, routing, DNS, NAT, and firewalls.
  • Proficiency in automating infrastructure with Go, Bash, Ansible, or Terraform.
  • Experience with monitoring/observability tools such as Prometheus, Grafana, or Loki.
  • Ability to troubleshoot across hardware, operating systems, networks, containers, and applications.
  • Willingness to work on-site in Hillsboro and handle equipment as needed.

Responsibilities

  • Deploy, operate, expand, and troubleshoot bare-metal NVIDIA and AMD GPU clusters.
  • Build and maintain Kubernetes platforms for training, inference, and research computing.
  • Manage Linux, container runtimes, NVIDIA drivers, CUDA, and the GPU Operator.
  • Configure and troubleshoot VLANs, routing, DNS, firewalls, and high-speed networks.
  • Automate server provisioning, configuration, upgrades, monitoring, and recovery.
  • Build monitoring and alerting for GPUs, servers, networks, storage, and Kubernetes.
  • Improve resource scheduling, GPU utilization, platform reliability, and user isolation.
  • Build internal tools that help research teams access compute and troubleshoot workloads.
  • Perform hands-on rack installation, cabling, BMC/IPMI management, and hardware troubleshooting.
  • Participate in on-call support for critical infrastructure and contribute to incident reviews.

Skills

Linux administration
Kubernetes
Networking (TCP/IP, VLANs, routing,DNS
Automation (Go,Bash,Ansible,Terraform)
Monitoring/observability (Prometheus,G

Tools

NVIDIA drivers
CUDA
GPU Operator

Job description

You will build and operate a large-scale GPU compute platform in Hillsboro, Oregon, while remotely maintaining the GPU cluster at our Santa Clara data center.

This is a highly hands-on role spanning Linux, Kubernetes, networking, automation, observability, and GPU hardware. You will work closely with research and engineering teams to make compute reliable, easy to use, and ready for rapid iteration. The work ranges from troubleshooting Kubernetes workloads and Linux networking to deploying bare-metal servers and diagnosing performance bottlenecks in multi-node GPU systems.

You will normally work from our Hillsboro office. This role does not require being stationed full-time in a data center; on-site equipment work is performed as needed.

What you’ll do
  • Deploy, operate, expand, and troubleshoot bare-metal NVIDIA and AMD GPU clusters.
  • Build and maintain Kubernetes platforms for training, inference, and research computing.
  • Manage Linux, container runtimes, NVIDIA drivers, CUDA, and the GPU Operator.
  • Configure and troubleshoot VLANs, routing, DNS, firewalls, and high-speed networks.
  • Automate server provisioning, configuration, upgrades, monitoring, and recovery.
  • Build monitoring and alerting for GPUs, servers, networks, storage, and Kubernetes.
  • Improve resource scheduling, GPU utilization, platform reliability, and isolation between users.
  • Build internal tools that help research teams access compute and troubleshoot their workloads.
  • Perform hands-on rack installation, cabling, BMC/IPMI management, and hardware troubleshooting.
  • Participate in on-call support for critical infrastructure and contribute to incident reviews.
What you’ll bring
  • Experience with Linux administration, performance analysis, and troubleshooting.
  • Practical experience operating Kubernetes and container infrastructure.
  • A working understanding of TCP/IP, VLANs, routing, DNS, NAT, and firewalls.
  • Experience automating infrastructure with tools such as Go, Bash, Ansible, or Terraform.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, or Loki.
  • The ability to troubleshoot across hardware, operating systems, networks, containers, and applications.
  • Willingness to work on-site in Hillsboro and handle equipment when needed. Your usual workplace will be the Hillsboro office, with data center equipment work as required rather than a permanent data center assignment.
Nice to have
  • Experience with NVIDIA or AMD GPU clusters.
  • Familiarity with NCCL, NVLink, RDMA, InfiniBand, RoCE, or GPUDirect.
  • Experience deploying or troubleshooting 100G–800G networks.
  • Experience with schedulers and distributed compute tools such as Slurm, Ray, KubeRay, or Kueue.
  • Experience with bare-metal lifecycle management using PXE, Redfish, or IPMI.
  • Familiarity with inference systems such as vLLM, SGLang, or Triton.
  • Experience with virtualization technologies such as KVM, Proxmox, KubeVirt, Incus, or Firecracker.
  • Experience with Ceph, ZFS, NVMe, or distributed storage.
  • Experience with eBPF, the Linux kernel, or performance optimization.
  • Familiarity with data center power, cooling, rack planning, and structured cabling.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
GPU Infrastructure Engineer
GPU Infrastructure Engineer

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000