Practice Manager - Platform Reliability, Operations Hub & Automation

Datacom

City of Melbourne

Hybrid

AUD 180,000 - 250,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote working
Flexi-hours
Professional development

Job summary

Datacom in Australia is seeking a seasoned technology leader to transition from manual operations to an automated, AI-led Mode 2 platform engineering model for our Reliability Operations Hub (ROH). You will lead a 24x7 operation across hybrid and cloud environments, shaping an SRE-as-a-Service model and driving significant toil reduction.

You will combine strategic vision, people leadership and commercial acumen to deliver measurable outcomes, including faster incident resolution, improved

Qualifications

  • Bachelor's degree in a relevant field (IT/Engineering) and 15+ years in large-scale tech ops.
  • Proven experience leading SRE/Platform Engineering teams across regions.
  • Strong track record delivering automation-first transformations.

Responsibilities

  • Lead 24x7 technical operations with follow-the-sun coverage.
  • Own ROH Front Door operating model and improve intake/assignment of work.
  • Drive automation through Ansible and AI-assisted workflows.
  • Establish SRE practices: SLOs, error budgets, and continuous improvement.
  • Lead major incidents and executive communications during crises.
  • Build and standardise runbooks and knowledge docs.

Skills

Team leadership
SRE/Platform Engineering
Automation strategy
Incident management
Commercial leadership
Cross-region coordination

Education

Bachelor's degree in IT/Engineering

Tools

Ansible Automation Platform
ITSM tooling
Observability platforms
Cloud platforms (AWS/Azure/GCP)

Job description

Our Why

Datacom works with organisations and communities across Australia and New Zealand to make a difference in people’s lives and help organisations use the power of tech to innovate and grow.


About the role (your why)

We're seeking an exceptional technology leader to transition from manual operations to an automated and AI led Mode 2 platform engineering model for our Reliability Operations Hub (ROH), a critical capability at the centre of our transformation from traditional IT operations to an AI-enabled, automation-first Platform Engineering model.


This is more than an operational leadership role. You'll lead the evolution of a 24x7 Reliability Operations Hub that not only drives operational excellence across complex hybrid and cloud environments but also helps shape the future commercialisation of our operational services through an innovative \"SRE-as-a-Service\" model.


You'll combine strategic vision, technical expertise, people leadership, and commercial acumen to deliver measurable outcomes including significant toil reduction, enhanced observability, accelerated incident resolution, and increased automation across the enterprise.


Clearance:This role works with sensitive data and requires NV1 clearance. Applicants must be Australian citizens and must hold, or be eligible to obtain, this clearance.

In this role you will

Lead 24x7 Technical Operations


  • Lead and scale regional 24x7 technical operations, ensuring effective follow-the-sun support, triage, handovers, and operational excellence.

  • Own and continuously optimise the ROH Front Door operating model, ensuring efficient intake, routing, segmentation, and tracking of operational work.

  • Apply Lean principles to improve operational flow, reducing queue wait times, cycle times, and delivery bottlenecks.

  • Define clear engagement models between the ROH and specialist engineering teams, protecting engineering capacity from low-value operational noise.

  • Drive integrated customer outcomes across modern platform environments.


Drive SRE & Automation Excellence


  • Lead Automation Swarms focused on converting repetitive operational work into reusable automated solutions using Ansible platform.

  • Champion Site Reliability Engineering (SRE) practices including Service Level Objectives (SLOs) and Error Budget frameworks.

  • Expand self-service capabilities and automate common operational tasks through standardised Golden Paths.

  • Drive adoption of automation-first and AI-assisted operational practices across the organisation.


Own Major Incident & Platform Resilience


  • Act as escalation leader during critical incidents, ensuring rapid service restoration and effective executive communication.

  • Lead root cause analysis and problem management processes that convert operational learnings into long-term improvements.

  • Ensure high-severity incidents result in identified and prioritised preventative automation opportunities.


Build Operational Standards & Knowledge Management


  • Govern the development and continuous improvement of operational runbooks and AI-ready documentation.

  • Reduce operational variance through standardisation, automated patching, role-based access controls, and gold-standard platform configurations.

  • Foster a culture focused on knowledge sharing and continuous improvement.


Shape Commercial Service Growth


  • Lead the transition of operational capabilities from a traditional cost centre to a scalable, productised service offering.

  • Partner with product and commercial teams to package observability, automated triage, and SRE capabilities into customer-facing services.

  • Drive operational efficiency, service margin growth, and the creation of repeatable, high-value offerings.


Inspire Teams & Transformation


  • Lead, coach, and develop a distributed Platform Reliability Engineering (PRE) team across multiple regions and drive process standardisation and unification.

  • Champion cross-skilling of PRE’s to build capabilities across infrastructure, cloud, database, middleware, and platform technologies.

  • Champion a culture of psychological safety, accountability, innovation, and continuous improvement.

  • Lead organisational change initiatives that shift teams from reactive operations to engineering-led automation practices.

  • Support right-shoring strategies across onshore, nearshore, and offshore delivery teams.


What You'll Bring

Experience


  • 15+ years' experience in large-scale technology operations environments.

  • At least 3 years leading Site Reliability Engineering, Platform Engineering, or similar operational engineering teams.

  • Proven success leading 24x7 operational functions across multiple regions and time zones.

  • Experience delivering transformational operating model change and driving automation-first ways of working.

  • Commercial leadership experience with service-based delivery models, service pricing, margin optimisation, and operational economics.

  • Strong experience managing major incidents, crisis response, and enterprise resilience programmes.


Technical Expertise


  • Deep understanding of SRE principles, platform reliability, and operational engineering.

  • Experience with enterprise observability platforms and ITSM tooling.

  • Knowledge of automation and orchestration technologies, including Ansible Automation Platform and AI-assisted operational workflows.

  • Strong understanding of APIs, systems integration, scripting, coding, and database automation.

  • Experience working across hybrid technology environments, including:


    • Windows and Linux platforms

    • Databases

    • Middleware technologies

    • Containers and virtualisation

    • AWS, Azure, and Google Cloud


  • Familiarity with CI/CD, version control, cloud-native operations, automation frameworks, and modern infrastructure platforms.


Culture and Benefits

Datacom is one of Australia and New Zealand’s largest suppliers of Information Technology professional services. We have managed to maintain a dynamic, agile, small business feel that is often diluted in larger organisations of our size. It's our people that give Datacom its unique culture and energy that you can feel from the moment you meet with us.


We care about our people and provide a range of perks such as social events, chill-out spaces, remote working, flexi-hours and professional development courses to name a few. You’ll have the opportunity to learn, develop your career, connect and bring your true self to work. You will be recognised and valued for your contributions and be able to do your work in a collegial, flat-structured environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Practice Manager - Automation and Operations
Practice Manager - Automation and Operations

Datacom • Sydney

Hybrid
AUD 250,000 - 350,000
Remote working
Flexi-hours
Professional development
Solution Specialist
Solution Specialist

Datacom • Sydney

Hybrid
AUD 110,000 - 150,000
Remote working
Flexi-hours
Professional development courses
+1
Technical Delivery Manager
Technical Delivery Manager

Datacom • Canberra

On-site
AUD 140,000 - 190,000
Social events
Chill‑out spaces
Remote working
+2
Desktop Support Team Lead
Desktop Support Team Lead

Datacom • Sydney

On-site
AUD 100,000 - 150,000
Parental leave
Employee Assistance (DataCare)
Learning & development
+3
Endpoint Engineering Specialist
Endpoint Engineering Specialist

Datacom • City of Brisbane

On-site
AUD 100,000 - 140,000
Social events
Remote working
Flexi-hours
+1
DevOps Engineer - Azure
DevOps Engineer - Azure

Datacom • Sydney

On-site
AUD 90,000 - 150,000
Social events
Chill-out spaces
Remote working
+2
General Manager - Product - 6 Month Fixed Term
General Manager - Product - 6 Month Fixed Term

Datacom • Sydney

Hybrid
AUD 180,000 - 260,000
Remote working
Flexi-hours
Professional development courses
+1
Business Improvement Analyst - Fixed Term
Business Improvement Analyst - Fixed Term

Datacom • City of Melbourne

Hybrid
AUD 90,000 - 130,000
Paid parental leave
24/7 Employee Assistance Programme (D​
Senior Mobility Engineer – Remote, Flexible Hours
Senior Mobility Engineer – Remote, Flexible Hours

Datacom • Sydney

On-site
AUD 80,000 - 100,000
Flexible working hours
Professional development courses
Health insurance
+2
Senior App Support Analyst - 24/7 Incident Lead
Senior App Support Analyst - 24/7 Incident Lead

Datacom • Sydney

On-site
AUD 100,000 - 130,000
Paid parental leave
DataCare – 24/7 Employee Assistance
Career development & certifications
+4