Infrastructure / Cluster Engineer

Linuxcareers

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Linuxcareers is seeking an Infrastructure/Cluster Engineer to design and operate large-scale clusters that enable AI inference at scale. The role focuses on managing diverse hardware architectures and building robust infrastructure.

The ideal candidate will possess deep expertise in Linux systems, automation tools, and orchestration technologies. Responsibilities include debugging performance issues and designing observability systems for cluster health. Experience with GPU infrastructure is a plus.

Qualifications

  • Experience in infrastructure, cluster engineering, platform engineering, or distributed systems.
  • Deep Linux systems experience with an emphasis on debugging.
  • Strong automation skills using Terraform, Ansible, or Python.

Responsibilities

  • Design and operate large-scale CPU and GPU clusters.
  • Build automation solutions for infrastructure management.
  • Debug complex production issues across various technology layers.

Skills

Infrastructure engineering
Linux systems
Kubernetes
Automation with Terraform
GPU infrastructure

Tools

Ansible
Python
Go

Job description

Gimlet is building AI infrastructure and orchestration platforms for large-scale AI datacenters. This Infrastructure/Cluster Engineer role involves designing, building, and operating heterogeneous cluster infrastructure that intelligently routes workloads across diverse hardware architectures to enable production AI inference at scale.

What You'll Do
  • Design, deploy, and operate large-scale CPU, GPU, and accelerator clusters powering production AI inference workloads
  • Build automation for provisioning, configuration, upgrades, validation, and lifecycle management across heterogeneous bare-metal infrastructure
  • Debug complex production issues spanning Linux, networking, storage, drivers, firmware, and orchestration layers
  • Build and operate high-performance networking infrastructure including RDMA-enabled environments and accelerator interconnects
  • Design and scale observability systems for cluster health, capacity, performance, failures, and workload behavior
What You Need
  • Experience in infrastructure, cluster engineering, platform engineering, SRE, HPC, or distributed systems
  • Deep Linux systems experience including debugging performance, networking, storage, processes, and kernel-level issues
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration and scheduling systems
  • Strong automation skills using Terraform, Ansible, Helm, Python, Go, or equivalent
  • Experience with GPU or accelerator infrastructure including drivers, firmware, CUDA/ROCm stacks, or hardware validation
Nice to Have
  • Experience building or operating AI inference, training, HPC, or neocloud infrastructure
  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up
  • Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting
  • Experience building observability platforms using Prometheus, OpenTelemetry, Grafana, or similar technologies
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Heterogeneous AI Infra & Cluster Engineer
Heterogeneous AI Infra & Cluster Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Network Engineer
Network Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 250,000 - 320,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000