Principal Ai Infrastructure Engineer, Kubernetes

Uncover

Sydney

On-site

AUD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Firmus Technologies is seeking an experienced Senior Infrastructure/Platform Engineer to own the Kubernetes platform across multi-cluster environments. You will design, build, and operate production-grade back-end services, with emphasis on automation, GitOps, and scalable infrastructure for AI workloads.

You will work on bare-metal Kubernetes deployment, cluster networking, storage patterns, GPU integration, and security governance while mentoring peers and influencing architecture decisions.

Qualifications

  • 10+ years in infrastructure/platform engineering with production Kubernetes ownership.
  • Experience designing multi-cluster Kubernetes on bare metal/private/hybrid infra.
  • Strong Go and scripting skills; Linux systems expertise.

Responsibilities

  • Define and own Kubernetes platform reference architecture across management and workload clusters.
  • Build and maintain backend services, APIs, controllers, operators to provision, configure, upgrade, scale clusters.
  • Engineer bare-metal Kubernetes deployment and lifecycle workflows using cluster tooling.
  • Design and operate cluster networking: CNI, ingress, DNS, load balancing, service mesh.
  • Define storage patterns using CSI, Ceph, NVMe, object storage, backup and restore.
  • Integrate NVIDIA GPU and related operators, telemetry, and topology-aware placement.
  • Establish GitOps and CI/CD patterns for platform software and release management.
  • Security governance: IAM, RBAC, secrets, policy-as-code, image security.
  • Define SLAs and observability; diagnose complex distributed failures.
  • Set engineering standards and mentor senior engineers; participate in implementation.

Skills

Kubernetes platform ownership
Go programming
Python
Linux systems
GitOps
Terraform
Ansible
CI/CD tooling
Networking (CNI, BGP)
Observability tooling

Education

Bachelor's degree in computer science or engineering

Tools

ClusterAPI
kubeadm
Redfish
Ironic
Metal3
CRI

Job description

Job Description

Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting‑edge technology with a steadfast commitment to sustainability. At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model‑to‑grid technology approach, we have pushed the boundaries of multi‑generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low‑cost AI tokens globally.

Firmus AI Cloud Our large‑scale GPU cloud platform, Firmus AI Cloud, is purpose‑built to deliver energy‑efficient AI compute at scale to customers. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever‑growing suite of services and applications, we are committed to delivering a cloud experience that is market‑leading, proprietary, and built to scale.

Key Responsibilities
  • Define and own the Kubernetes platform reference architecture across management and workload clusters, including control‑plane topology, cluster lifecycle, multi‑tenancy, workload isolation, and failure‑domain design.
  • Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably.
  • Engineer repeatable bare‑metal Kubernetes deployment and lifecycle workflows using infrastructure‑as‑code and automated provisioning technologies such as ClusterAPI, kubeadm, Redfish, PXE, Ironic, or Metal3.
  • Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR‑IOV, BGP, InfiniBand, or RoCE where required for high‑performance AI workloads.
  • Define persistent‑storage and data‑service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster‑recovery mechanisms appropriate for stateful platform and AI workloads.
  • Integrate and productionise NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology‑aware placement for multi‑node accelerated workloads.
  • Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management, with safe testing, progressive rollout, rollback, and upgrade practices.
  • Build platform security into the architecture through identity and access control, RBAC, secrets management, policy‑as‑code, image and software‑supply‑chain controls, tenant isolation, and auditable change management.
  • Define service‑level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures and eliminate recurring operational toil.
  • Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross‑team technical decisions while remaining directly involved in implementation.
Skills & Experience
  • 10+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years operating at senior staff, principal, or equivalent level.
  • Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, upgrades, and control‑plane failure modes.
  • Demonstrated experience designing, building, and operating highly available, multi‑cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
  • Strong software engineering ability in Go, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
  • Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel and container runtime behaviour, performance analysis, and low‑level troubleshooting.
  • Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi‑network architectures.
  • Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
  • Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
  • Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies, and using telemetry to manage reliability, capacity, and performance.
  • Experience with GPU‑enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling, RDMA networking, and distributed AI workload requirements.
  • Experience with distributed storage and data services such as Ceph, CSI‑backed storage, object storage, backup and restore, and disaster recovery.
  • CKA‑level expertise is expected; CKA, CKS, or relevant cloud‑native certifications are strongly preferred.
  • Bachelor's degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
  • Clear technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams.
Location & Reporting

Australia (Sydney, NSW or Launceston, TAS) Reporting to Head of AI Platform

Employment Basis

Full‑time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions. Join us in our mission to revolutionize the AI industry through sustainable practices and cutting‑edge engineering.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

Firmus Technologies • Sydney

On-site
AUD 150,000 - 210,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

Firmus • Sydney

On-site
AUD 130,000 - 180,000
Senior HPC Infrastructure Engineer
Senior HPC Infrastructure Engineer

Matchbox • Council of the City of Sydney

On-site
AUD 180,000 - 240,000
Senior Software Engineer, AI Job Orchestration
Senior Software Engineer, AI Job Orchestration

Matchbox • Sydney

On-site
AUD 150,000 - 210,000
Principal AI Infrastructure Engineer (Kubernetes)
Principal AI Infrastructure Engineer (Kubernetes)

Firmus Technologies • Sydney, City of Launceston

On-site
AUD 180,000 - 240,000
Site Reliability Engineer, AI Infrastructure
Site Reliability Engineer, AI Infrastructure

Firmus • City of Melbourne

On-site
AUD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Firmus Technologies • City of Launceston

On-site
AUD 90,000 - 120,000
Senior DevOps Engineer, AI & Applications
Senior DevOps Engineer, AI & Applications

Firmus Technologies • Sydney

On-site
AUD 100,000 - 130,000
Diversity and inclusion initiatives
Senior AI DevOps Engineer - Scale GPU Deployments Safely
Senior AI DevOps Engineer - Scale GPU Deployments Safely

Firmus Technologies • Sydney

On-site
AUD 100,000 - 130,000
Diversity and inclusion initiatives
Site Reliability Engineer
Site Reliability Engineer

Matchbox • City of Launceston

On-site
AUD 110,000 - 170,000