Site Reliability Engineer (Azure)

University of New South Wales

Sydney

Hybrid

AUD 120,000 - 160,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible Working Options
17% Superannuation
Extra 3 days leave

Job summary

UNSW is seeking a Site Reliability Engineer (Azure) to ensure reliable Azure-based integration platforms across production and non-production environments. You will own service reliability, set SLAs, and lead incident reviews while collaborating with developers, security, and platform teams.

The role requires hands-on Azure expertise, strong monitoring capabilities, and IaC/integration experience. Hybrid working with 2–3 days in the Kensington office is offered in Sydney.

Qualifications

  • Strong hands-on experience with Microsoft Azure integration services (ADF, Logic Apps, Function Apps, API Management, Event Hub).
  • Experience applying Site Reliability Engineering (SRE) principles in cloud environments.
  • Proficiency in monitoring/observability tools and automation/scripting.
  • Experience with Infrastructure as Code and incident management.

Responsibilities

  • Own the reliability and availability of Azure integration services across production and non-production environments.
  • Monitor system health and performance using Azure Monitor, Log Analytics and Application Insights.
  • Define and maintain Service Level Objectives and metrics for key services.
  • Lead incident response, triage, root cause analysis, and post-incident reviews.
  • Reduce operational toil through automation (PowerShell, Azure CLI) and IaC (Bicep/Terraform).
  • Maintain CI/CD pipelines for deployments and config changes.
  • Ensure environments are consistent, secure, and supportable across dev/test/prod.
  • Oversee configuration management and logging enhancements.
  • Engage with security, platform teams, and vendors on improvements and compliance.

Skills

Azure integration services
SRE principles
Monitoring & observability
Automation scripting
Infrastructure as Code
Incident management
High availability
Azure Monitor
PowerShell
Azure CLI
Bicep/Terraform

Tools

Azure Monitor
Log Analytics
Application Insights
PowerShell
Azure CLI
Bicep
Terraform

Job description

At UNSW, we take pride in the broad range and high quality of our teaching programs. Our teaching gains strength and currency from our research activities, strong industry links and our international nature; UNSW has strong regional and...

  • Full-time continuing role within UNSW IT as a Site Reliability Engineer ( Azure )
  • Starting Salary $132,445 plus generous superannuation and leave loading
  • Kensington, Sydney location, 2-3 days in the office, Hybrid working
About UNSW:

UNSW isn’t like other places you’ve worked. Yes, we’re a large organisation with a diverse and talented community; a community doing extraordinary things. Together, we are driven to be thoughtful, practical, and purposeful in all we do. Taking this combined approach is what makes our work matter. It’s the reason we’re one of the top 20 universities in the world (QS top 20) and a member of Australia’s prestigious Group of Eight. If you want a career where you can thrive, be challenged and do meaningful work, you’re in the right place.

The Site Reliability Engineer (Azure) is a specialist role within the Integration Technologies team, responsible for ensuring the reliability, availability, performance, and operational excellence of Azure-based integration platforms and services. This role focuses on the production and non-production environments supporting enterprise integrations (e.g., Azure Data Factory, Logic Apps, API Management, Event Hub), applying SRE principles to proactively manage system health, automate operations, reduce operational toil, and continuously improve service resilience and observability. The role bridges engineering and operations, working closely with integration developers, platform teams, and vendors to ensure solutions are supportable, scalable, and aligned with operational acceptance standards.

Specific accountabilities for this role include:
  • Own the reliability and availability of Azure integration services across production and non production environments
  • Monitor system health, performance, and availability using Azure Monitor, Log Analytics, Application Insights, DataDog and other observability tooling
  • Define, implement, and maintain Service Level Objectives and other metrics for key services
  • Lead incident response, triage, root cause analysis, and post-incident reviews (PIRs)
  • Reduce operational toil through automation using scripting (PowerShell, Azure CLI), IaC (Bicep/Terraform), and pipelines
  • Implement self-healing mechanisms, alerting improvements, and proactive remediation patterns o Maintain and improve CI/CD pipelines for operational deployments and configuration changes
  • Manage and maintain Azure integration platform components including API Management, Logic Apps, Function Apps, Data Factory, SQL, Databrick, Service Bus/Event Hub o Ensure environments (dev/test/prod) are consistent, secure, and supportable
  • Oversee configuration management including Key Vault integration, secrets rotation, certificates, and connectivity (e.g., private endpoints, hybrid connections) o Enhance logging, monitoring, and tracing (e.g., distributed tracing with tools like App Insights)
  • Analyse trends in incidents, performance, and usage to identify systemic improvements
  • Maintain operational dashboards and reporting for service health and performance
  • Contribute to and enforce Operational Acceptance Criteria (OAC) for new and changed integration solutions
  • Be available for On-Call / After Hours support on a rotating basis, typically one week per month as well as weekend and After Hours work as required
  • Ensure solutions meet standards for supportability, monitoring, alerting, and documentation before production release
  • Maintain runbooks, support documentation, and knowledge base articles o Work closely with Integration Engineers, Data Engineers, and Architects to embed reliability into solution design
  • Partner with security and platform teams on compliance, patching, and platform upgrades
  • Engage with vendors and third-parties for issue resolution and service improvements
  • Ensure adherence to security, governance, and compliance requirements (e.g., identity, access, secrets management)
  • Support audit activities and implement remediation actions where required
  • Align with and actively demonstrate the Code of Conduct and Values o Ensure hazards and risks psychosocial and physical are identified and controlled for tasks, projects, and activities that pose a health and safety risk within your area of responsibility.
Who you are:
  • Essential - Strong hands-on experience with Microsoft Azure, particularly integration services (ADF, Logic Apps, Function Apps, API Management, Event Hub/Service Bus, SQL, Databricks) - Experience applying Site Reliability Engineering (SRE) principles in cloud environments - Proficiency in monitoring and observability tools (Azure Monitor, Log Analytics, Application Insights) - Experience with automation and scripting (PowerShell, Azure CLI) and Infrastructure as Code (Bicep, Terraform) - Strong incident management and problem-solving skills, including root cause analyse - Experience managing production systems with high availability and reliability requirements.
  • Desirable - Experience with distributed tracing and advanced observability (e.g., OpenTelemetry, Jaeger, AppInsights) - Knowledge of CI/CD pipelines (Azure DevOps, GitHub Actions) - Familiarity with integration patterns and enterprise integration architecture - Experience working in hybrid cloud or complex enterprise environments - Understanding of FinOps and cost optimisation in Azure.
  • Behavioural Capabilities - Strong analytical and systems-thinking mindset - Proactive and continuous improvement focus o Ability to balance operational stability with delivery demands - Effective communication and stakeholder engagement skills - An understanding of and commitment to UNSW’s aims, objectives and values in action, together with relevant policies and guidelines - Knowledge of health & safety (psychosocial and physical) responsibilities and commitment to attending relevant health and safety training.
Benefits and Culture
  • Flexible Working Options (work from home, flexible hours etc)
  • 17% Superannuation contributions and additional leave loading payments
  • Additional 3 days of leave over Christmas period
  • Discounts and entitlements (retail, education, fitness)

Please note: Sponsorship is not available for this role; valid Australian working rights are required on application.

Pre-Employment Checks

As part of our recruitment process candidates may be required to undergo pre-employment screening, which may include reference checks, qualification verification, right-to-work verification, and criminal history screening where relevant to the role.

Applications close

Thursday 17th of September at 11.30pm

UNSW is committed to equity diversity and inclusion. Applications from women, people of culturally and linguistically diverse backgrounds, those living with disabilities, members of the LGBTIQ+ community; and people of Aboriginal and Torres Strait Islander descent, are encouraged. UNSW provides workplace adjustments for people with disability, and access to flexible work options for eligible staff.

The University reserves the right not to proceed with any appointment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Azure)
Site Reliability Engineer (Azure)

University of New South Wales • Penrith City Council

Hybrid
AUD 119,000 - 146,000
17% Superannuation contributions
Leave loading
Additional 3 days of leave over the JS
+1
Integration Analyst
Integration Analyst

University of New South Wales • Sydney

Hybrid
AUD 107,000 - 130,000
Hybrid working
Competitive salary
Integration Analyst
Integration Analyst

University of New South Wales • Penrith City Council

Hybrid
AUD 101,000 - 136,000
Flexible working
17% Superannuation
Leave loading payments
+2
Senior Platform Engineer
Senior Platform Engineer

University of New South Wales • Penrith City Council

Hybrid
AUD 131,000 - 177,000
Flexible Working Options
17% Superannuation contributions and 3
additional leave over Christmas
Site Reliability Engineer
Site Reliability Engineer

usyd • Sydney

Hybrid
AUD 131,000 - 148,000
Superannuation
Parental leave
Salary packaging
+4
Product Manager - IT Services
Product Manager - IT Services

UNSW • Sydney

Hybrid
AUD 139,000 - 169,000
Flexible Working Options
Career development opportunities
17% Superannuation contributions and 3
Product Manager - IT Services
Product Manager - IT Services

University of New South Wales • Penrith City Council

Hybrid
AUD 131,000 - 177,000
Flexible Working Options (work from do
17% Superannuation contributions
Additional leave loading payments
+1
Project Manager – IT Focused
Project Manager – IT Focused

University of New South Wales • Penrith City Council

On-site
AUD 132,000 - 150,000
3 extra leave days during December
Up to 50% discount on UNSW courses
Flexible 17% superannuation and leave-
+1
Product Manager - IT Services
Product Manager - IT Services

University of New South Wales • Sydney

On-site
AUD 139,000 - 169,000
Flexible working
17% Superannuation
Leave loading
+1
Cyber Security Assurance Testing Lead
Cyber Security Assurance Testing Lead

University of New South Wales • Penrith City Council

Hybrid
AUD 140,000 - 170,000
Flexible Working Options
Career development opportunities
17% Superannuation
+2