Senior Platform Engineer

Firmus Technologies

Singapore

On-site

SGD 140,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Firmus Technologies is seeking a Senior Platform Engineer to drive the design and implementation of MLOps capabilities for a cutting-edge AI factory platform. You will collaborate with engineers to scale the platform from model to grid, focusing on IaC, container orchestration, observability, security, and self-service experiences.

You will build, operate, and secure production Kubernetes infrastructure and contribute to reliability and incident response, while working closely with founders in a

Qualifications

  • Bachelor's degree in computer science or related technical field.
  • 7+ years of experience as Platform Engineer, Site Reliability Engineer, DevOps engineer, MLOps Engineer or Observability Engineer.
  • Infrastructure-as-Code, configuration management and CI/CD (Terraform, Ansible, GitHub Actions, Jenkins, ArgoCD).
  • Containerization technologies (Docker), Kubernetes networking and cluster management, including upgrades and troubleshooting.
  • Observability stack design and scaling (Loki, Grafana, Tempo, Prometheus, Thanos, ClickHouse).
  • Telemetry solutions using Redfish, gNMI, SNMP, eBPF, streaming analytics; OpenTelemetry.
  • Compliance automation (OPA, Kyverno).
  • Scripting and programming skills (Bash, Python, Go).
  • Linux internals, networking, and distributed storage.
  • Clear and effective English communication, written and spoken.
  • Bonus: SOC 2 Type 2 and ISO 27001 experience.

Responsibilities

  • Build MLOps capabilities from the ground up across environments.
  • Improve DevOps platform for reliability, scalability, security, and CI/CD integration.
  • Design, operate and secure Kubernetes-based production infra, including NVL and InfiniBand/Spectrum-X components.
  • Develop observability platforms for internal and external users.
  • Integrate Firmus central services with NVIDIA software stack.
  • Lead enhancement and evangelism of internal platform products with secure self-service experiences.
  • Drive incident response, participate in on‑call, and perform detailed RCA to improve reliability.

Skills

7+ years experience
SRE/DevOps background
MLOps
English communication

Education

Bachelor's degree in computer science or related field

Tools

Terraform
Ansible
GitHub Actions
Jenkins
ArgoCD
Docker
Kubernetes
Loki
Grafana
Prometheus
Tempo
Thanos
ClickHouse
OpenTelemetry
OPA
Kyverno

Job description

Firmus Technologies is a global leader pioneering the development and operation of efficient AIinfrastructure from model to grid. Founded in Australia in 2019, our mission is to create the mostefficient AI infrastructure by combining cutting-edge technology with a steadfast commitment tosustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digitalinfrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed theboundaries of multi-generational liquid cooling systems, energy management, AI softwareorchestration, and construction. This co-designed approach from model to grid allows us to makeevery watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Why you’ll love working here

At Firmus, you’ll work at the intersection of sustainability and artificial intelligence in a fast-pacedenvironment powered by next-generation technology. You’ll be helping to transform an entireindustry — and you’ll feel it every day.

Our team is made up of true innovators and leaders in their fields, and as an emerging company, youwon’t be lost in a crowd. You’ll work closely with the founders, build a strong network, and see theimpact of your work first-hand as we democratise AI tools for everyone — more sustainably andmore affordably.

We believe great things happen when people from diverse backgrounds come together to do theirbest work and be their authentic selves. We are proud to be an equal opportunity employer.

ROLE

Firmus Technologies is seeking a Senior Platform Engineer to join our Engineering and Technology team. You will drive the design and implementation of ou r MLO ps capability. You will also collaborate with other engineers and make technical decision s on scal ing Firmus AI factory platform engineering capabilities to planet sca le , from IaC , container orchestration, observability , self-service portal to platform security. This role is ideal for a self-starter with passion for building things from first principles. You naturally break down complex problems into their fundamental truths to uncover novel and elegant solutions - rather than relying on conventional patterns.

KEY RESPONSIBILITIES

  • Build MLOps capabilities from the ground up, enabling reproducible, scalable, and secure ML workflows across internal and customer-facing environments.
  • Continuously improve our DevOps platform to ensure reliability, scalability, security, and seamless integration with CI/CD pipelines and infrastructure services.
  • Design, implement, operate and secure Kubernetes-based production infrastructure for high reliability, performance and security, including clusters supporting NVIDIA GB300 NVL72 systems with NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet.
  • Develop world-class observability platforms for internal and external customers
  • Integrate Firmus central services with NVIDIA’s software stack, including Mission Control, NETQ, UFM, and NMX.
  • Lead the enhancement and evangelism of internal platform products that provide cohesive, composable, secure-by-default, and low-friction self-service experiences that accelerates time to market and reduce engineers' cognitive load.
  • Drive incident response efforts, participate actively in the on-call rotation, and lead detailed Root Cause Analysis (RCA) to continuously improve system reliability, operational maturity, and incident handling processes.

SKILLS AND EXPERIENCE

  • Bachelor's degree in computer science or a related technical field.
  • 7+ years of experience as Platform Engineer, Site Reliability Engineer, DevOps engineer, MLOps Engineer or Observability Engineer.
  • Demonstrated strong proficiency on the following areas:
    • Infrastructure-as-Code, configuration management and CI/CD (e.g., Terraform, Ansible, GitHub Actions, Jenkins, ArgoCD).
    • Containerization technologies (e.g., Docker), Kubernetes networking and cluster management, including upgrades and troubleshooting.
    • Observability stack design and scaling (e.g., Loki, Grafana, Tempo, Prometheus, Thanos, ClickHouse).
    • Telemetry solutions using various technology (e.g., Redfish, gNMI, SNMP, eBPF, streaming analytics).
    • Unified telemetry collection with OpenTelemetry.
    • Compliance automation (e.g., OPA, Kyverno).
  • Competent in scripting and programming skills (e.g., Bash, Python, Go).
  • Systems knowledge on Linux internals, networking stacks, and distributed storage.
  • Clear and effective English communication, written and spoken.
  • Bonus: Experience in high-growth startups or regulated industries with robust security and data privacy requirements, including SOC 2 Type 2 and ISO 27001.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer, MLOps & Observability
Senior Platform Engineer, MLOps & Observability

Firmus Technologies • Singapore

On-site
SGD 140,000 - 260,000
Network Engineer
Network Engineer

Firmus Technologies • Singapore

On-site
SGD 120,000 - 180,000
Principal Solutions Architect, AI Data Infrastructure
Principal Solutions Architect, AI Data Infrastructure

Firmus Technologies • Singapore

On-site
SGD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Firmus Technologies • Singapore

On-site
SGD 120,000 - 160,000
Senior AI Infrastructure Engineer (Virtualisation)
Senior AI Infrastructure Engineer (Virtualisation)

Firmus Technologies • Singapore

On-site
SGD 232,378 - 335,657
DevOps Engineer (AL-FNC260724 007/01)
DevOps Engineer (AL-FNC260724 007/01)

United States Digital Space LLC • Singapore

On-site
SGD 90,000 - 150,000
Senior Devops Engineer
Senior Devops Engineer

luxoft information technology (singapore) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Project Manager, MLOps
Project Manager, MLOps

Hyundai Motor Group Innovation Center Singapore (HMGICS) • Singapore

On-site
SGD 140,000 - 220,000
Senior Platform Engineer, Cloud (Xora Portfolio Company)
Senior Platform Engineer, Cloud (Xora Portfolio Company)

United States Digital Space LLC • Singapore

Hybrid
SGD 180,000 - 282,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Ensign InfoSecurity • Singapore

On-site
SGD 120,000 - 180,000