Senior Kubernetes Platform Engineer

Firmus Technologies

Sydney

On-site

AUD 180,000 - 240,000

Full time

35 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Firmus Technologies is seeking a Senior Kubernetes Platform Engineer to own the multi-tenant Kubernetes estate that powers AI FactoryOS and related platforms. You will manage the platform lifecycle, automation tooling and guardrails to keep clusters healthy across a large fleet.

The role requires deep Kubernetes expertise, experience with GPU-enabled workloads, and strong automation skills in a 24/7 production setting. Based in Australia or Singapore with travel to sites as needed.

Qualifications

  • 8+ years in platform/infrastructure engineering with production Kubernetes in 24/7 environments.
  • Experience operating Kubernetes at fleet scale and multi-cluster management.
  • Strong knowledge of Kubernetes control plane internals and upgrades.
  • Experience writing controllers or admission logic in production.
  • Familiarity with multi-tenant patterns and GPU-enabled Kubernetes.

Responsibilities

  • Ensure reliable operation and automation of a multi-tenant Kubernetes platform across the fleet.
  • Develop operational tooling, guarded remediation and orchestration for fleet-scale operations.
  • Manage cluster lifecycle: provisioning, patching, upgrade and decommissioning.
  • Operate and recover Kubernetes control planes and tenant onboarding pipelines.
  • Lead incident response and drive permanent fixes through runbooks and knowledge sharing.

Skills

Kubernetes expertise
Platform engineering
Automation & IaC
Scripting for automation
Incident response
Security fundamentals
Documentation & design notes
Strong communication

Education

Bachelor's degree in computer science or related

Tools

OpenTofu
Terraform
Ansible
Argo CD
NVIDIA GPU Operator
vCluster

Job description

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability. It also builds the shared services the estate's own operation depends on, and runs them. Operating the estate every day is what shows how the platform behaves under real load and under failure, and the function works with the engineering teams that build it to turn what it finds into permanent fixes and design improvements.

Role Summary

Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Kubernetes Platform Engineer runs the Kubernetes estate that every product and every tenant runs on: the platform controller layer, the virtual cluster platform tenants are provisioned onto, the Kubernetes environment baselines used across the estate, the GPU integration layer, and the automated tenant onboarding and release pipeline.

This is a hands-on senior role with deep technical expertise . This role owns the platform lifecycle execution across the fleet : keeping clusters healthy and current across the fleet, keeping tenant workloads running through upgrade and failure, and restoring control planes to service when they degrade . Automation is a first-class part of the role , in the operational tooling, guarded remediation and fleet orchestration that make estate-scale operation possible, delivered as controlled code and reviewed by AI Infrastructure where it affects service behaviour.

Key Responsibilities
  • Responsible for the reliable operation, automation and continuous improvement of the multi-tenant Kubernetes platform that Firmus' products and tenants run on, spanning every site in the estate.
  • Build the operational tooling, guarded remediation and orchestration that automate operations at fleet scale, and contribute operator and controller requirements, and code where agreed, to AI Infrastructure's platform backlog with the production evidence behind them.
  • Execute the Kubernetes cluster lifecycle across the fleet, including provisioning, patching, upgrade and decommissioning, running the deployment and upgrade mechanisms built by AI Infrastructure through the agreed staged or canary path, and holding estate version compliance and retirement coordination.
  • Operate and recover Kubernetes control planes carrying live tenant workload, including etcd state, certificate rotation, failed upgrades and corrupted resources.
  • Operate the virtual cluster platform and multi-tenant isolation patterns that tenants are provisioned onto, and the automated tenant onboarding and release pipeline that lands new tenants safely and repeatably.
  • Operate the GPU integration layer for Kubernetes (for example the NVIDIA GPU Operator), including device plugins, GPU scheduling and driver coordination.
  • Diagnose and resolve scheduling failures, CNI and CSI faults, admission rejections and resource contention from first principles, and drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks.
  • Run the Kubernetes baselines in production carrying the admission policy, workload identity and network policy content set by the Senior Platform Security Engineer, and hold the operational acceptance requirements those baselines have to meet before they enter production.
  • Provide the deepest technical expertise for Kubernetes faults across the estate, diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure , and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results.
  • Lead technical recovery during major Kubernetes incidents under the incident commander, drive the changes that remove repeat causes through the problem record, and share the after-hours escalation roster for the Kubernetes estate.
Skills & Experience
Required Skills
  • Strong skills in platform and infrastructure engineering, with 8+ years of experience overall and substantial ownership of production Kubernetes platforms in a 24/7 environment.
  • Deep experience operating Kubernetes at fleet scale, including cluster lifecycle, upgrades and multi-cluster management.
  • Strong experience with Kubernetes control plane internals, including etcd, the API server, controllers, schedulers, and certificate and credential rotation.
  • Strong experience writing Kubernetes controllers, operators or admission logic in a production setting.
  • Experience with multi-tenant or virtual cluster patterns (for example vCluster or equivalent), including tenant isolation at the Kubernetes layer.
  • Experience operating GPU-enabled Kubernetes, including device plugins, GPU scheduling and driver coordination (for example the NVIDIA GPU Operator).
  • Strong skills in infrastructure automation, infrastructure-as-code and GitOps practices (for example OpenTofu or Terraform, Ansible, Argo CD), with change delivered through peer review, automated testing and progressive rollout.
  • Strong experience with scripting or programming for operational automation and tooling, such as Go, Python or Bash.
  • Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, post-incident review, and the production of runbooks that others can execute successfully.
  • Solid understanding of Kubernetes security fundamentals, including admission control, workload identity and network policy.
  • Clear technical judgement and communication skills, with the ability to produce documentation, design notes and escalations that other engineers can act on.
Preferred Experience
  • Experience operating Kubernetes for GPU or HPC workloads at scale.
  • Experience with automated tenant or customer onboarding pipelines in a multi-tenant platform.
  • Experience with vendor Kubernetes distributions or reference architectures for accelerated computing.
  • Familiarity with DPU or SmartNIC-based networking as it relates to Kubernetes CNI design.
  • Experience contributing to open-source Kubernetes ecosystem projects.
  • A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training.
Expected Outcomes
  • Fleet-wide cluster upgrades executed through the staged path with no unplanned tenant-visible outage, and version compliance held across the estate.
  • Tenant onboarding running as a routine operation rather than a project.
  • Control plane recovery tested against live-equivalent conditions and executable by an engineer who did not write the procedure.
  • Kubernetes escalations falling as operational automation and guarded remediation take on the common faults.
  • Platform defects and operability gaps evidenced into the platform engineering backlog and closed permanently, with repeat causes falling.
Location & Reporting
Location

Based in Australia or Singapore, with travel to Australian AI Factory sites as required.

On-call

The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Security Engineer
Senior Platform Security Engineer

Firmus Technologies • Sydney

On-site
AUD 170,000 - 260,000
null
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

Firmus Technologies • Sydney

On-site
AUD 180,000 - 260,000
Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

United States Digital Space LLC • Sydney

On-site
AUD 220,000 - 280,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

Firmus Technologies • Sydney

On-site
AUD 150,000 - 210,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

Firmus • Sydney

On-site
AUD 130,000 - 180,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

re-zoo-me • Sydney

Hybrid
AUD 180,000 - 240,000
Principal Platform Identity Engineer
Principal Platform Identity Engineer

Firmus Technologies • Sydney

On-site
AUD 180,000 - 230,000
Principal Platform Identity EngineerNew
Principal Platform Identity EngineerNew

Firmus Technologies • Sydney

On-site
AUD 180,000 - 260,000
Engineering Manager, Platform
Engineering Manager, Platform

Firmus Technologies • Sydney

On-site
AUD 120,000 - 190,000