Site Reliability Engineer (Sre) / Senior Devops Engineer - Azure (4 Positions)

9X5 Consulting

City of Melbourne

Hybrid

AUD 140,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

9X5 Consulting seeks experienced Site Reliability Engineers / Senior DevOps Engineers to join clients in Melbourne. You will drive reliability, scalability and operational performance of critical platforms with strong Azure, automation, observability and modern infrastructure practices.

Applicants should demonstrate SRE principles, SLIs/SLOs, and incident post-mortems, with hands-on Azure across compute, AKS, storage and databases. Hybrid Melbourne CBD arrangement offered.

Qualifications

  • Minimum 3 years in SRE/DevOps engineering.
  • Hands-on Azure across compute, networking, App Service/AKS.
  • Experience with SLIs, SLOs and error budgets.
  • Experience with Azure Monitor, Log Analytics and Application Insights.
  • IaC using Terraform or Bicep.
  • CI/CD pipelines with Azure DevOps or GitHub Actions.
  • Automation using PowerShell, Python, or Go.
  • Networking basics: DNS, TCP/IP, load balancing, NSGs.
  • Linux/Windows Server administration.
  • Incident management and post-incident reviews.

Responsibilities

  • Ensure reliability, availability, performance and scalability of production services.
  • Define, implement and manage SLIs, SLOs and error budgets.
  • Design and support solutions across Microsoft Azure (compute, networking, App Services, AKS, storage, databases).
  • Build and maintain monitoring, logging and observability using Azure Monitor, Log Analytics and Application Insights.
  • Develop and maintain IaC using Terraform and/or Bicep.
  • Build, maintain and improve CI/CD pipelines using Azure DevOps and/or GitHub Actions.
  • Automate operational and infrastructure processes using PowerShell, Python and/or Go.
  • Monitor production environments, troubleshoot issues and respond to incidents.
  • Lead root-cause analysis and blameless post-mortems.
  • Collaborate with development, infrastructure, security teams to embed reliability practices.

Skills

Azure experience
SRE DevOps
Terraform
CI/CD
PowerShell
Python
Go
Kubernetes
Monitoring Observability
Networking

Tools

Azure Monitor
Log Analytics
Application Insights
Terraform
Bicep
Azure DevOps
GitHub Actions
Docker
Kubernetes

Job description

Job Description

Site Reliability Engineer (SRE) / Senior DevOps Engineer – Azure (4 Positions)

Site Reliability Engineer (SRE) / Senior DevOps Engineer – Azure (4 Roles)

Melbourne CBD + Hybrid

Introduction

9X5 Consulting is seeking experienced Site Reliability Engineer (SRE) / Senior DevOps Engineers to join one of our clients and support the reliability, scalability and operational performance of business‑critical production platforms.

This role would suit an experienced SRE, DevOps, Platform or Cloud/Infrastructure Engineer who has worked with production systems at scale and has strong hands‑on experience across Microsoft Azure, automation, observability and modern infrastructure practices.

You don't necessarily need to have held the title of "Site Reliability Engineer". We are particularly interested in candidates who can demonstrate strong production operations experience and an understanding of SRE principles, including SLIs, SLOs, error budgets, observability and continuous improvement.

About 9X5 Consulting

We are transforming the way business and technology is managed by putting real‑time data into the hands of every decision maker across organisations. The insight garnered from diverse backgrounds, perspectives and lived experiences results in pioneering innovations across the organisation and better experiences for our customers. The more diverse our talent, the more impact we have on each other and on our valued clients.

Our clients range from small SME's right through to some of the largest ASX listed organisations in Australia, so experience at selling to stakeholders at this level is required.

A bit about you

You are an experienced SRE, DevOps, Platform or Infrastructure Engineer who enjoys working with complex production environments and finding ways to make systems more reliable, scalable and easier to operate.

You have strong hands‑on experience with Microsoft Azure and are comfortable working across cloud infrastructure, automation, CI/CD, monitoring and observability. You understand that reliability is more than simply responding to incidents — it's about proactively identifying issues, automating repetitive tasks and continually improving the environment.

Ideally, you will bring:
  • 3+ years' experience in SRE, DevOps, Platform or Infrastructure Engineering, supporting production services at scale.
  • Strong hands‑on Microsoft Azure experience across compute, networking, App Service/AKS, storage and databases.
  • Experience working with SLIs, SLOs and error budgets in production environments.
  • Experience with Azure Monitor, Log Analytics and Application Insights.
  • Infrastructure as Code experience using Terraform and/or Bicep.
  • CI/CD experience using Azure DevOps and/or GitHub Actions.
  • Automation and scripting skills using PowerShell, Python and/or Go.
  • A solid understanding of networking, including DNS, TCP/IP, load balancing and NSGs.
  • Experience supporting Linux and/or Windows Server environments.
  • Experience with incident management, root‑cause analysis and blameless post‑mortems.
  • Exposure to Docker, Kubernetes/AKS, Prometheus and Grafana would be highly regarded.
  • An understanding of ITIL and experience working within Scrum or Kanban environments.
Key Responsibilities
  • Ensure the reliability, availability, performance and scalability of business‑critical production services.
  • Define, implement and manage SLIs, SLOs and error budgets to measure and improve service reliability.
  • Design, deploy and support solutions across Microsoft Azure, including compute, networking, App Services, AKS, storage and databases.
  • Build and maintain monitoring, logging and observability solutions using Azure Monitor, Log Analytics and Application Insights.
  • Develop and maintain infrastructure using Terraform and/or Bicep and Infrastructure as Code best practices.
  • Build, maintain and improve CI/CD pipelines using Azure DevOps and/or GitHub Actions.
  • Automate operational and infrastructure processes using PowerShell, Python and/or Go.
  • Monitor production environments, troubleshoot complex issues and respond to incidents.
  • Lead or contribute to root‑cause analysis and blameless post‑mortems, ensuring lessons learned result in measurable improvements.
  • Work across networking and infrastructure technologies including DNS, TCP/IP, load balancing, NSGs, Linux and Windows Server.
  • Support containerised environments using Docker and Kubernetes/AKS, where required.
  • Identify opportunities to reduce manual intervention, improve resilience and increase operational efficiency through automation.
  • Collaborate closely with development, infrastructure, security and operational teams to embed reliability and operational best practices throughout the delivery lifecycle.
Key Deliverables
  • Reliable, scalable and highly available production services and infrastructure.
  • Clearly defined and measurable SLIs and SLOs, supported by appropriate error budgets.
  • Effective monitoring, alerting and observability across critical applications and infrastructure.
  • Automated and repeatable infrastructure deployments using Terraform and/or Bicep.
  • Reliable and efficient CI/CD pipelines supporting application and infrastructure deployments.
  • Increased automation of operational processes, reducing manual effort and improving consistency.
  • Improved production stability through proactive monitoring, performance analysis and remediation.
  • Effective incident response, root‑cause analysis and blameless post‑mortems, with identified improvements followed through to completion.
  • Improved platform resilience, performance and scalability across the Azure environment.
  • Clear operational documentation, procedures and knowledge transfer to support ongoing service management.
  • Continuous improvements that strengthen reliability, operational efficiency and overall service performance.
Job Requirements
Essential
  • Minimum 3 years' experience in SRE, DevOps, Platform Engineering, Infrastructure Engineering or a similar role supporting production environments.
  • Strong hands‑on experience with Microsoft Azure, including compute, networking, App Service, storage and databases.
  • Demonstrated experience supporting and operating production services at scale.
  • Experience defining, implementing or working with SLIs, SLOs and error budgets.
  • Strong experience with Azure monitoring and observability tools, including Azure Monitor, Log Analytics and Application Insights.
  • Experience with Infrastructure as Code (IaC) using Terraform and/or Bicep.
  • Experience building and maintaining CI/CD pipelines using Azure DevOps and/or GitHub Actions.
  • Strong scripting and automation capability using PowerShell and Python, or a similar language such as Go.
  • Solid understanding of networking concepts including DNS, TCP/IP, load balancing and Network Security Groups (NSGs).
  • Experience administering or supporting Linux and/or Windows Server environments.
  • Practical experience with incident management, troubleshooting, root‑cause analysis and post‑incident reviews.
  • Strong problem‑solving, communication and stakeholder engagement skills.
Desirable
  • Hands‑on experience with Docker and containerised applications.
  • Experience with Kubernetes and Azure Kubernetes Service (AKS).
  • Experience with Prometheus and Grafana for monitoring and observability.
  • Understanding of distributed systems and highly available architectures.
  • Experience implementing or improving formal Site Reliability Engineering (SRE) practices.
  • Knowledge of ITIL principles and service management practices.
  • Experience working within Scrum and/or Kanban delivery environments.
  • Relevant Microsoft Azure, DevOps, Kubernetes or cloud certifications would be advantageous.
Personal attributes
  • A genuine passion for technology, automation, reliability and continuous improvement.
  • A proactive, solutions‑focused approach to identifying and resolving complex technical issues.
  • Strong analytical and troubleshooting skills with the ability to remain methodical when responding to production incidents.
  • A strong sense of ownership and accountability for the reliability and performance of the environments you support.
  • The ability to work calmly and effectively when managing critical production issues and incidents.
  • A mindset focused on automation, with a desire to eliminate repetitive manual processes wherever possible.
  • Strong attention to detail and a commitment to building reliable, secure and scalable solutions.
  • The ability to work independently while contributing positively to collaborative engineering and operational teams.
  • Excellent communication skills with the ability to explain complex infrastructure, cloud and reliability concepts to both technical and non‑technical stakeholders.
  • A continuous improvement mindset, with a willingness to challenge existing processes and identify better ways of working.
  • The ability to adapt to changing priorities and work effectively within dynamic production environments.
  • A professional, reliable and accountable approach to your work.
  • A collaborative and blameless approach to incident management, focused on learning, improvement and prevention rather than assigning fault.
  • A willingness to share knowledge, support colleagues and contribute to the ongoing development of the broader team.
  • A commitment to applying SRE, DevOps, security and cloud engineering best practices to improve platform reliability and operational performance.
  • A professional, reliable and accountable approach to your work.
  • A positive attitude and willingness to share knowledge, mentor others and contribute to team success.
  • A commitment to delivering solutions that are secure, scalable and aligned with best practice engineering principles.
Deal-breakers (stuff we don't want in a new team member)
  • A "that'll do" approach to our clients' projects;
  • A closed mind to change;
  • Big egos. We're really, really not keen on big egos here.

Only applicants with citizenship, permanent visa will be considered.

No agencies please!

Get your free, confidential resume review.
or drag and drop your file here.