An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Firmus Technologies seeks a Service Delivery Manager to own incident and change management across our AI infrastructure portfolio. You will lead major incidents, coordinate runbooks, and drive production readiness with engineering teams while reporting on service health and reliability.
The role requires strong ITSM discipline, governance mindset and experience in 24/7 operations within a cloud/hyperscaler context. Based in AU or SG with travel to Australia sites.
Firmus Technologies is a globalleader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combiningcutting-edgetechnology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, andoperate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AIcomputeat scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.
AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.
The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.
Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Service Delivery Manager owns the incident and change management practices this operation runs on: the severity model and major incident command, the change calendar and change records, the runbook programme, production readiness review, and the operational reporting that keeps leadership and customers informed of service health, so that our services remain reliable and secure.
The role sits at the fusion of ITIL and SRE practices. Incident, problem and change are managed with the rigour of formal service management, and delivered with the methods of reliability engineering: measured against service level objectives, automated wherever automation makes response faster and safer, and continuously improved from what incidents reveal.
The role owns the processes, not the technical decisions inside them. This role owns the standard, the record and the discipline that make those decisions consistent, visible and auditable, and it owns the authorisation path that turns a technically complete release into an authorised production change.
Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
The function runs 24/7. This role provides major incident command cover across regions alongside peer Service Delivery Managers, on a published roster.
Reports to the Head of AI FactoryOS Operations while the function is being established, working under broad direction with a high degree of autonomy and direct access to the decision makers. As the function reaches its planned structure, the role will report to the Service Reliability Manager, with the Head of AI FactoryOS Operations remaining accountable for the function. The scope, level and remit of the role do not change under either arrangement.
Permanent full-time
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering.