Senior Kubernetes Platform Engineer

Firmus Technologies Pty Ltd.

Sydney

On-site

AUD 180,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Firmus Technologies is seeking a Senior Kubernetes Platform Engineer to own and operate the multi-tenant Kubernetes estate across Australia and Singapore. You will manage fleet-wide upgrades, guard remediation and the tenant onboarding pipeline while ensuring production reliability and resilience.

The role demands hands-on expertise in GPU-enabled Kubernetes, IaC/GitOps, and strong incident response capabilities.

Qualifications

  • 8+ years in platform and infrastructure engineering.
  • Production Kubernetes on 24/7 environments.
  • Kubernetes cluster lifecycle, upgrades, and multi-cluster management.
  • Deep knowledge of etcd, API server, controllers and rotation.
  • Experience writing Kubernetes controllers or admission logic.
  • Multi-tenant/virtual cluster patterns familiar.
  • GPU-enabled Kubernetes experience with NVIDIA Operator.
  • IaC/GitOps with OpenTofu, Terraform, Ansible, Argo CD.
  • Scripting in Go/Python/Bash for automation.
  • Senior escalation point with runbooks and post-incident reviews.
  • Kubernetes security fundamentals and policy controls.
  • Strong technical communication and documentation.

Responsibilities

  • Responsible for reliable operation, automation and continuous improvement of the multi-tenant Kubernetes platform.
  • Build operational tooling, guarded remediation and orchestration for fleet-scale operations.
  • Execute cluster lifecycle: provisioning, patching, upgrades and decommissioning.
  • Operate and recover Kubernetes control planes carrying live tenant workloads.
  • Operate the virtual cluster platform and automated tenant onboarding/release pipelines.
  • Operate the GPU integration layer (NVIDIA Operator) including device plugins and driver coordination.
  • Diagnose scheduling, CNI/CSI faults and drive CI/CD and testing framework improvements.
  • Run baselines with admission policy, workload identity and network policy checks.
  • Provide deep Kubernetes fault expertise and mentor frontline engineers.
  • Lead major incident recovery and coordinate after-hours escalation roster.

Skills

Kubernetes platform ownership
Fleet-scale Kubernetes
Control plane internals
Controllers / operators
Multi-tenant / virtual clusters
GPU-enabled Kubernetes
IaC & GitOps
Automation scripting (Go/Python/Bash)
Incident response / on-call
Kubernetes security fundamentals
Documentation & runbooks

Education

Bachelor's degree in CS/Engineering or equivalent

Tools

OpenTofu
Terraform
Ansible
Argo CD
NVIDIA GPU Operator
vCluster (or equivalents)

Job description

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability. It also builds the shared services the estate's own operation depends on, and runs them. Operating the estate every day is what shows how the platform behaves under real load and under failure, and the function works with the engineering teams that build it to turn what it finds into permanent fixes and design improvements.

Role Summary

Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Kubernetes Platform Engineer runs the Kubernetes estate that every product and every tenant runs on: the platform controller layer, the virtual cluster platform tenants are provisioned onto, the Kubernetes environment baselines used across the estate, the GPU integration layer, and the automated tenant onboarding and release pipeline.

This is a hands-on senior role with deep technical expertise . This role owns the platform lifecycle execution across the fleet : keeping clusters healthy and current across the fleet, keeping tenant workloads running through upgrade and failure, and restoring control planes to service when they degrade . Automation is a first-class part of the role , in the operational tooling, guarded remediation and fleet orchestration that make estate-scale operation possible, delivered as controlled code and reviewed by AI Infrastructure where it affects service behaviour.

Key Responsibilities
  • Responsible for the reliable operation, automation and continuous improvement of the multi-tenant Kubernetes platform that Firmus' products and tenants run on, spanning every site in the estate.
  • Build the operational tooling, guarded remediation and orchestration that automate operations at fleet scale, and contribute operator and controller requirements, and code where agreed, to AI Infrastructure's platform backlog with the production evidence behind them.
  • Execute the Kubernetes cluster lifecycle across the fleet, including provisioning, patching, upgrade and decommissioning, running the deployment and upgrade mechanisms built by AI Infrastructure through the agreed staged or canary path, and holding estate version compliance and retirement coordination.
  • Operate and recover Kubernetes control planes carrying live tenant workload, including etcd state, certificate rotation, failed upgrades and corrupted resources.
  • Operate the virtual cluster platform and multi-tenant isolation patterns that tenants are provisioned onto, and the automated tenant onboarding and release pipeline that lands new tenants safely and repeatably.
  • Operate the GPU integration layer for Kubernetes (for example the NVIDIA GPU Operator), including device plugins, GPU scheduling and driver coordination.
  • Diagnose and resolve scheduling failures, CNI and CSI faults, admission rejections and resource contention from first principles, and drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks.
  • Run the Kubernetes baselines in production carrying the admission policy, workload identity and network policy content set by the Senior Platform Security Engineer, and hold the operational acceptance requirements those baselines have to meet before they enter production.
  • Provide the deepest technical expertise for Kubernetes faults across the estate, diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure , and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results.
  • Lead technical recovery during major Kubernetes incidents under the incident commander, drive the changes that remove repeat causes through the problem record, and share the after-hours escalation roster for the Kubernetes estate.
Skills & Experience
Required Skills
  • Strong skills in platform and infrastructure engineering, with 8+ years of experience overall and substantial ownership of production Kubernetes platforms in a 24/7 environment.
  • Deep experience operating Kubernetes at fleet scale, including cluster lifecycle, upgrades and multi-cluster management.
  • Strong experience with Kubernetes control plane internals, including etcd, the API server, controllers, schedulers, and certificate and credential rotation.
  • Strong experience writing Kubernetes controllers, operators or admission logic in a production setting.
  • Experience with multi-tenant or virtual cluster patterns (for example vCluster or equivalent), including tenant isolation at the Kubernetes layer.
  • Experience operating GPU-enabled Kubernetes, including device plugins, GPU scheduling and driver coordination (for example the NVIDIA GPU Operator).
  • Strong skills in infrastructure automation, infrastructure-as-code and GitOps practices (for example OpenTofu or Terraform, Ansible, Argo CD), with change delivered through peer review, automated testing and progressive rollout.
  • Strong experience with scripting or programming for operational automation and tooling, such as Go, Python or Bash.
  • Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, post-incident review, and the production of runbooks that others can execute successfully.
  • Solid understanding of Kubernetes security fundamentals, including admission control, workload identity and network policy.
  • Clear technical judgement and communication skills, with the ability to produce documentation, design notes and escalations that other engineers can act on.
Preferred Experience
  • Experience operating Kubernetes for GPU or HPC workloads at scale.
  • Experience with automated tenant or customer onboarding pipelines in a multi-tenant platform.
  • Experience with vendor Kubernetes distributions or reference architectures for accelerated computing.
  • Familiarity with DPU or SmartNIC-based networking as it relates to Kubernetes CNI design.
  • Experience contributing to open-source Kubernetes ecosystem projects.
  • A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training.
Expected Outcomes
  • Fleet-wide cluster upgrades executed through the staged path with no unplanned tenant-visible outage, and version compliance held across the estate.
  • Tenant onboarding running as a routine operation rather than a project.
  • Control plane recovery tested against live-equivalent conditions and executable by an engineer who did not write the procedure.
  • Kubernetes escalations falling as operational automation and guarded remediation take on the common faults.
  • Platform defects and operability gaps evidenced into the platform engineering backlog and closed permanently, with repeat causes falling.
Location & Reporting

Location : Based in Australia or Singapore, with travel to Australian AI Factory sites as required.

On-call: The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function.

About Firmus Technologies

Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs and engineers passionate about sustainable computing infrastructure.

Firmus builds and operates AI infrastructure across Asia-Pacific, utilising its proprietary AI Factory platform to deliver transformative cost-effective GPU clusters and AI cloud services for developers, enterprise, education and government users.

We are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Kubernetes Platform Engineer
Senior Kubernetes Platform Engineer

Firmus Technologies • Sydney

On-site
AUD 180,000 - 240,000
Senior Kubernetes Platform Engineer
Senior Kubernetes Platform Engineer

Firmus • Sydney

On-site
AUD 150,000 - 210,000
Senior Platform Security Engineer
Senior Platform Security Engineer

Firmus Technologies Pty Ltd. • Sydney

Hybrid
AUD 180,000 - 210,000
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

Firmus Technologies • Sydney

On-site
AUD 150,000 - 210,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

re-zoo-me • Sydney

Hybrid
AUD 180,000 - 240,000
Senior Platform Reliability Engineer (Fabric and Interconnect)
Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus • Sydney

On-site
AUD 140,000 - 190,000
Senior Platform Security Engineer
Senior Platform Security Engineer

Firmus Technologies • Sydney

On-site
AUD 170,000 - 260,000
null
Service Delivery Manager
Service Delivery Manager

Firmus Technologies Pty Ltd. • Sydney

On-site
AUD 120,000 - 180,000
Senior Platform Reliability Engineer (Fabric and Interconnect)
Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies • Sydney

On-site
AUD 140,000 - 210,000