Senior AI Platform Reliability Engineer

Matchbox

City of Melbourne

Hybrid

AUD 180,000 - 260,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Firmus Technologies in Melbourne, Australia, seeks a Senior Platform Reliability Engineer to own and operate the AI FactoryOS production estate, including GPU clusters, exabyte-scale storage, and shared services. You will drive automation, CI/CD, and strong incident response within a 24/7 on-site team.

You will contribute to architecture, tooling, and runbooks, ensuring service levels and security requirements across the estate are met with rigorous, evidence-based improvements.

Qualifications

  • 8+ years in infrastructure, systems and platform engineering in 24/7 environments.
  • Expertise with scale-out storage (VAST/WEKA/Ceph/Lustre/GPFS/NetApp) and performance diagnosis.
  • Experience operating virtualization platforms and shared services at defined service levels.
  • Deep Linux knowledge: storage, kernel, networking, I/O, drivers.
  • Observability as a service: metrics, logs, traces, alerting, runbooks.

Responsibilities

  • Operate multi-tenant control plane and access controls for tenant workloads.
  • Manage exabyte-scale storage and S3-compatible object storage with automation.
  • Run GPU compute fleet: health, fault handling, firmware baselines, vendor coordination.
  • Build guarded automation and remediation tooling; deliver changes as code.
  • Diagnose data-path performance; apply OS, network and storage tuning.
  • Maintain observability infrastructure; design actionable alerts tied to runbooks.
  • Lead incident recovery, on-call rotations, and post-incident reviews.

Skills

Experience 8+ yrs
Infrastructure
Storage systems
Linux systems
Kubernetes
Automation / IaC
CI/CD
Incident response
Networking
Programming (Python/Go)

Tools

Proxmox
VMware
KVM
Terraform/OpenTofu
Argo CD
Prometheus/Grafana/OpenTelemetry

Job description

Firmus Technologies in Melbourne, Australia, seeks a Senior Platform Reliability Engineer to own and operate the AI FactoryOS production estate, including GPU clusters, exabyte-scale storage, and shared services. You will drive automation, CI/CD, and strong incident response within a 24/7 on-site team.

You will contribute to architecture, tooling, and runbooks, ensuring service levels and security requirements across the estate are met with rigorous, evidence-based improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Reliability Engineer - AI FactoryOS, 24/7
Senior Platform Reliability Engineer - AI FactoryOS, 24/7

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Senior GPU Interconnect & Platform Reliability Engineer
Senior GPU Interconnect & Platform Reliability Engineer

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 240,000
Site Reliability Engineer: AI Infra & HPC
Site Reliability Engineer: AI Infra & HPC

Firmus Technologies • City of Melbourne

On-site
AUD 120,000 - 180,000
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Senior AI Platform Engineer - Build the AI Factory
Senior AI Platform Engineer - Build the AI Factory

Matchbox • Sydney

Hybrid
AUD 180,000 - 240,000
ESOP – Employee Stock Ownership Plan
AI Productivity Benefit – up to USD 1,
Parental Leave – 12 weeks paid, plus 6
+1
Senior Platform Security Engineer – 24/7 AI Infra Protection
Senior Platform Security Engineer – 24/7 AI Infra Protection

Firmus • Sydney

On-site
AUD 180,000 - 260,000
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Matchbox • City of Melbourne

Hybrid
AUD 180,000 - 260,000
Site Reliability Engineer, AI Infrastructure
Site Reliability Engineer, AI Infrastructure

Firmus Technologies • City of Melbourne

On-site
AUD 120,000 - 180,000
Senior AI Infra Engineer — Kubernetes for GPU Clusters
Senior AI Infra Engineer — Kubernetes for GPU Clusters

Matchbox • Sydney

Hybrid
AUD 180,000 - 240,000
Senior Platform Reliability Engineer (Fabric and Interconnect)
Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies • Sydney

On-site
AUD 140,000 - 210,000