Data Center Engineer (Fleet Operations) - AI Infrastructure

Hamilton Barnes Associates Limited

United States

On-site

USD 120,000 - 180,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Full Benefits

Job summary

Hamilton Barnes Associates Limited is seeking a Fleet Operations Engineer to own the health, performance, and availability of the hardware underpinning data-center AI workloads. You will work hands-on with production data center and HPC-style infrastructure to ensure reliability.

The role requires building tooling for fleet health, automating workflows, and responding to incidents, with on-call rotations and cross-team collaboration across platform and control-plane groups.

Qualifications

  • Experience operating production data center or HPC infrastructure at scale.
  • Strong systems fundamentals across macOS, Linux, and Unix.
  • Programming/scripting in Python, Go, Rust, or Bash.
  • Experience with telemetry, monitoring, and observability tooling for hardware fleets.
  • Comfort with on-call responsibilities and on-site operational work.
  • Builder mentality for a small, fast-moving team.

Responsibilities

  • Operate and maintain production data center / HPC infrastructure running macOS, Linux, and Unix systems.
  • Build and run tooling for fleet health, telemetry, and observability across the hardware fleet.
  • Diagnose and resolve hardware and system-level issues, including on-site and remote troubleshooting.
  • Automate fleet operations workflows using Python, Go, Rust, or Bash.
  • Participate in on-call rotations and respond to production incidents affecting the fleet.
  • Work cross-functionally with the platform/control-plane team to coordinate capacity and lifecycle events.
  • Support scaling operations as the fleet grows, including new hardware onboarding and deployment.

Skills

Data center operations
macOS/Linux/Unix
Scripting languages
Telemetry/Observability
On-call readiness
Startup mindset

Job description

Keen to join a company that champions growth and development?

Join a high-growth infrastructure company that makes Apple hardware available and performant at data center scale for AI workloads, building cloud platforms that expose high-performance hardware as elastic compute through developer-friendly interfaces.

The organization is currently on the lookout for a Fleet Operations Engineer to join the team responsible for the physical layer of the compute platform. The ideal candidate will own the health, performance, and availability of the machines underpinning every customer workload, working hands-on with production data center and HPC-style infrastructure.

Responsibilities:
  • Operate and maintain production data center / HPC infrastructure running macOS, Linux, and Unix systems
  • Build and run tooling for fleet health, telemetry, and observability across the hardware fleet
  • Diagnose and resolve hardware and system-level issues, including on-site and remote troubleshooting
  • Automate fleet operations workflows using Python, Go, Rust, or Bash
  • Participate in on-call rotations and respond to production incidents affecting the fleet
  • Work cross-functionally with the platform/control-plane team to coordinate capacity and lifecycle events
  • Support scaling operations as the fleet grows, including new hardware onboarding and deployment
Skills/Must Have:
  • Experience operating production data center or HPC infrastructure at scale
  • Strong systems fundamentals across macOS, Linux, and Unix
  • Programming/scripting ability in Python, Go, Rust, or Bash
  • Experience with telemetry, monitoring, and observability tooling for hardware fleets
  • Comfort with on-call responsibilities and hands-on/on-site operational work
  • A builder mentality suited to a small, fast-moving team scaling from the ground up
Benefits:
  • Full Benefits
Salary:
  • $120,000 - $180,000 Base Salary
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Platform - AI Infrastructure
Member of Technical Staff, Platform - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 213,000 - 288,000
Equity
Health care
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
AI DataCenter Network Production Engineer
AI DataCenter Network Production Engineer

Intelletec Energy • San Francisco (CA)

On-site
USD 200,000 - 275,000
Fleet Operations Engineer — Data Center & HPC Infra
Fleet Operations Engineer — Data Center & HPC Infra

Hamilton Barnes Associates Limited • United States

On-site
USD 120,000 - 180,000
Full Benefits
Big Data Systems Engineer
Big Data Systems Engineer

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 170,000
Software Engineer, Fleet Management
Software Engineer, Fleet Management

OpenAI • San Francisco (CA)

Hybrid
USD 230,000 - 490,000
Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering
Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Infrastructure Engineer
Infrastructure Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 390,000
Equity
Principal Data Center Infrastructure Software Engineer
Principal Data Center Infrastructure Software Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 150,000 - 210,000
Hardware Operations Engineer
Hardware Operations Engineer

OpenAI • Seattle (WA)

On-site
USD 150,000 - 200,000