Senior AI Infrastructure Engineer, Kubernetes

Firmus Technologies Pty Ltd.

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Firmus Technologies seeks a Senior Kubernetes Engineer to own backend infrastructure for the AI Factory fleet. You will design, build, and operate production-grade clusters and automation across GPU-accelerated bare-metal environments, driving security, observability, and reliability at scale.

You will lead platform engineering standards, collaborate with AI platforms, security, and networking teams, and mentor engineers while delivering resilient multi-tenant solutions.

Qualifications

  • 7+ years in infrastructure/platform engineering with production Kubernetes ownership.
  • Deep knowledge of Kubernetes internals (API server, etcd, scheduler, controller manager, kubelet).
  • Experience designing and operating multi-cluster Kubernetes on bare metal/private/hybrid infra.
  • Go and/or Rust with practical Python and Bash skills; building Kubernetes operators recommended.
  • Strong Linux systems knowledge and networking expertise (CNI, DNS, LB, policy).
  • GitOps and CI/CD expertise with Terraform, Ansible, Argo CD, Flux, and CI tools.
  • Security/governance: RBAC, secrets, image security, policy-as-code, isolation.
  • Observability: Prometheus, Grafana, OpenTelemetry, logs, traces; capacity planning.

Responsibilities

  • Define and own the Kubernetes platform reference architecture across management and workload clusters.
  • Build and maintain backend services, APIs, controllers, operators, and automation to manage clusters.
  • Automate bare-metal Kubernetes deployment and lifecycle with IaaC and provisioning tools.
  • Design and operate cluster networking, including CNI, ingress, DNS, load balancing, and service mesh.
  • Define persistent storage and data-service patterns (Ceph, CSI, object storage, DR).
  • Integrate NVIDIA GPU and Network Operators for multi-node AI workloads.
  • Establish GitOps and CI/CD patterns for platform software with safe testing and rollouts.
  • Incorporate security with IAM, RBAC, secrets, policy-as-code, and image security.
  • Define SLOs, observability, and diagnostics for complex distributed systems.
  • Mentor senior engineers and participate in cross-team architectural decisions.

Skills

Kubernetes
Go
Python
Bash
Linux
GitOps
CI/CD
Cluster API
kubeadm
Redfish
PXE
Ironic
Metal3
NVIDIA GPU Operator
RBAC

Education

Bachelor’s degree in computer science, engineering, or related discipline

Tools

Terraform
Ansible
Argo CD
Flux
GitHub Actions
GitLab CI
Jenkins
Cluster API
kubeadm

Job description

Firmus Technologies

Firmus Technologies is a globalleader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combiningcutting-edgetechnology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, andoperate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AIcomputeat scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

Role Summary

The Senior Kubernetes Engineer, AI Infrastructure owns the technical design and delivery of the backend infrastructure that powers the Firmus Kubernetes platform. This is a hands‑on principal-level individual contributor role, responsible for building production‑grade cluster lifecycle, control‑plane, networking, storage, security, observability, and automation capabilities across GPU-accelerated bare‑metal environments.

They solve the hardest platform engineering problems, set Kubernetes engineering standards, and provide domain-level technical sign‑off for platform designs. They work across AI Platforms, Solutions Architecture & Delivery, networking, security, and operations to create a secure, resilient, multi‑tenant platform that can be deployed and operated consistently at AI‑factory scale.

Key Responsibilities
  • Define and own the Kubernetes platform reference architecture across management and workload clusters, including control‑plane topology, cluster lifecycle, multi‑tenancy, workload isolation, and failure‑domain design.
  • Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably.
  • Engineer repeatable bare‑metal Kubernetes deployment and lifecycle workflows using infrastructure‑as‑code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, or Metal3.
  • Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR‑IOV, BGP, InfiniBand, or RoCE where required for high‑performance AI workloads.
  • Define persistent‑storage and data‑service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster‑recovery mechanisms appropriate for stateful platform and AI workloads.
  • Integrate and productionise NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology‑aware placement for multi‑node accelerated workloads.
  • Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management, with safe testing, progressive rollout, rollback, and upgrade practices.
  • Build platform security into the architecture through identity and access control, RBAC, secrets management, policy‑as‑code, image and software‑supply‑chain controls, tenant isolation, and auditable change management.
  • Define service‑level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures and eliminate recurring operational toil.
  • Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross‑team technical decisions while remaining directly involved in implementation.
Skills & Experience
  • 7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years operating at senior staff, principal, or equivalent level.
  • Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, cluster performance, upgrades, and control‑plane failure modes.
  • Demonstrated experience designing, building, and operating highly available, large scale and multi‑cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
  • Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
  • Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel, host networking and container runtime behaviour, performance analysis, and low‑level troubleshooting.
  • Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi‑network architectures.
  • Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
  • Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
  • Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies, and using telemetry to manage reliability, capacity, and performance.
  • Experience with GPU‑enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements.
  • Experience with distributed storage and data services such as Ceph, CSI‑backed storage, object storage, backup and restore, and disaster recovery.
  • CKA‑level expertise is expected; CKA, CKS, or relevant cloud‑native certifications are strongly preferred.
  • Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
  • Clear technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams.
Location & Reporting
  • San Francisco Bay Area
  • Reporting to Head of AI Platform
Employment Basis

Full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting‑edge engineering.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

Firmus • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

Engg • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Kubernetes Architect, AI Infrastructure
Senior Kubernetes Architect, AI Infrastructure

Firmus • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Networking Engineer
Senior Networking Engineer

Firmus Technologies Pty Ltd. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Kubernetes Architect for AI Infrastructure
Senior Kubernetes Architect for AI Infrastructure

Firmus Technologies • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Networking Engineer
Senior Networking Engineer

Firmus Technologies • San Francisco (CA)

On-site
USD 140,000 - 210,000
Strategic Account Lead
Strategic Account Lead

Firmus Technologies • San Francisco (CA)

On-site
USD 160,000 - 220,000
Head of Customer Success
Head of Customer Success

Firmus Technologies • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Kubernetes Administrator – AI Infrastructure
Kubernetes Administrator – AI Infrastructure

Sira Consulting, an Inc 5000 company • United States

On-site
USD 140,000 - 210,000