Production Reliability Engineer

ASX

Sydney

Hybrid

AUD 150,000 - 190,000

Full time

22 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Flexible working
Hybrid work options

Job summary

ASX is seeking an experienced Site Reliability Engineer to ensure the reliability and performance of critical technology platforms. You will design and operate comprehensive observability, automation, and capacity planning in a hybrid Sydney-based environment.

Join a diverse team delivering secure, scalable services with on-call rotations and IaC-driven automation. Strong scripting, container orchestration, and observability experience are essential.

Qualifications

  • 5+ years of SRE/DevOps/Platform engineering experience.
  • Strong scripting in Python, PowerShell or alike.
  • Solid knowledge of Linux/Windows, networking and distributed systems.
  • Experience with CI/CD tooling and automated deployment pipelines.
  • Strong understanding of containerisation and orchestration (Docker, Kubernetes).
  • Experience with observability tools (Grafana, Splunk, OpenTelemetry).
  • Knowledge of databases (Oracle, SQL Server, PostgreSQL) and NoSQL.
  • Familiarity with APIs, real-time messaging and Kafka.
  • Experience with non-functional test planning and execution.

Responsibilities

  • Ensure availability, reliability, and performance of critical platforms.
  • Design and maintain monitoring, logging, alerting, and observability solutions.
  • Build dashboards and health indicators for real-time visibility.
  • Drive continual service improvement and incident reduction.
  • Lead root cause analyses and implement permanent fixes.
  • Maintain runbooks, playbooks, and recovery procedures.
  • Support 24x7 production environments and on-call rotations.
  • Implement IaC to automate infrastructure, deployment, and operations.
  • Collaborate with Engineering, Security, and Product teams to improve outcomes.
  • Conduct capacity planning and performance analyses to meet demand.

Skills

Site Reliability
Python/PowerShell
Linux/Windows
CI/CD
Kubernetes
Docker
Monitoring/Observability
SQL/NoSQL
Kafka

Education

Bachelor's degree in CS/ related

Tools

Grafana
Splunk
ITRS Geneos
OpenTelemetry
AWS CloudWatch

Job description

ASX: Powering Australia's financial markets
Why join the ASX?

When you join ASX, you’re joining a company with a strong purpose – to power a stronger economic future by enabling a fair and dynamic marketplace for all.

In your new role, you’ll be part of a leading global securities exchange with a strong brand. We are known for being a trusted market operator and an exciting data hub.

We are more than a securities exchange!

The ASX team brings together talented people from a diverse range of disciplines.

We run critical market infrastructure, with 1 in 3 people employed within technology. Yet we have a unique complexity of roles across a range of disciplines such as operations, program delivery, financial products, investor engagement, risk and compliance.

We’re proud to foster a workplace where diversity is celebrated and inclusion is part of our everyday culture. Our employee-led networks champion LGBTIQ+ inclusion, promote gender equality, accessibility and wellbeing, inspire giving and volunteering, and celebrate cultural and religious events, creating a sense of belonging for all. As an AWEI Bronze employer and member of the Champions of Change Coalition for gender equality, we’re committed to a fair and inclusive workplace where everyone can thrive.

Key purpose of the role

The role is accountable for ensuring the reliability, availability, performance, observability, and operational resilience of critical business platforms through proactive monitoring, automation, incident management, capacity planning, and continuous service improvement, enabling secure and stable technology services that meet business and customer expectations.

Your Responsibilities
  • Ensure the availability, reliability, scalability, and resilience of critical technology platforms and services.
  • Design, implement, and maintain comprehensive monitoring, logging, alerting, and observability solutions.
  • Develop meaningful dashboards, operational metrics, and health indicators to provide real-time visibility of platform performance.
  • Drive continuous improvement initiatives to reduce service disruptions and improve platform stability.
  • Proactively identify, assess, and mitigate reliability risks across production environments.
  • Conduct root cause analyses (RCA) and post-incident reviews, driving permanent resolutions and preventative actions.
  • Maintain operational runbooks, playbooks, and recovery procedures to support rapid incident resolution.
  • Support 24x7 production environments and participate in on-call rotations where required.
  • Develop and maintain infrastructure, deployment, and operational automation using Infrastructure as Code (IaC) principles.
  • Build automated recovery mechanisms to minimise manual intervention.
  • Champion engineering best practices across operational and development teams.
  • Conduct capacity planning, performance analysis, and workload forecasting to ensure services can meet current and future demand.
  • Identify and resolve performance bottlenecks across applications, infrastructure, databases, and integrations.
  • Partner with development teams to improve deployment reliability, release automation, and operational supportability.
  • Work closely with Engineering, Architecture, Security, Infrastructure, Product, and Operations teams to improve service outcomes.
Your Experience And Qualifications
Must have
  • 5+ years’ experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Engineering, Platform Engineering, or Production Support environments.
  • Strong scripting experience in Python, Java or PowerShell.
  • Strong knowledge of Linux, Windows, networking, and distributed systems.
  • Experience with CI/CD tooling and automated deployment pipelines.
  • Strong understanding of containerisation and orchestration technologies, including Docker and Kubernetes.
  • Experience with observability platforms such as Grafana, Splunk, ITRS Geneos, OpenTelemetry, AWS CloudWatch or similar.
  • Knowledge of database technologies, including Oracle, SQL Server, PostgreSQL or NoSQL platforms.
  • Familiarity with integration technologies, including APIs, real‑time messaging and streaming architectures such as Kafka.
  • Non-functional test planning and execution experience.
Nice to have
  • Experience working in the Capital Markets industry, with a solid understanding of product and transaction lifecycles.
  • Passionate about solving problems, troubleshooting software issues and triaging environment issues.
  • Excellent verbal and written communication skills, with the ability to work with internal and external stakeholders.
  • Lateral thinker who brings forward ideas that will automate repeatable tasks to reduce work effort and ensure quality deliverables.
  • Learns quickly and enjoys the challenge of learning new systems.
  • Compassionate, empathetic and self-motivated.
  • Able to take ownership and not be afraid to acknowledge failure.

We make hiring decisions based on your skills, capabilities and experience, and how you’ll help us to live our values.

At ASX Group, our diverse workforce is essential to build and maintain a fair and dynamic marketplace. We support flexible working and offer hybrid working options. Even if our roles are advertised as full‑time, we encourage you to apply if you are interested in part‑time or other flexible working arrangements.

We will arrange for successful candidates to have background checks, including reference and police checks, completed as part of the on-boarding process.

To be considered for this position, candidates must be legally authorised to work in Australia on a permanent basis without any restrictions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Reliability Engineer
Production Reliability Engineer

ASX • Sydney

On-site
AUD 90,000 - 120,000
Hybrid working options
Production Reliability Engineer
Production Reliability Engineer

ASX Limited • Australia

On-site
AUD 110,000 - 160,000
Hybrid working options
Flexible working arrangements
Employee benefits
Engineering Lead Data and DevOps - Equities Trading Technology
Engineering Lead Data and DevOps - Equities Trading Technology

ASX Limited • Australia

Hybrid
AUD 180,000 - 240,000
Flexible working
Employee benefits
Engineering Lead Data and DevOps - Equities Trading Technology
Engineering Lead Data and DevOps - Equities Trading Technology

ASX • Sydney

On-site
AUD 210,000 - 270,000
Hybrid working options
Flexible work arrangements
Senior Manager – Resilience Culture & Capability lead
Senior Manager – Resilience Culture & Capability lead

ASX • Sydney

Hybrid
AUD 180,000 - 240,000
Hybrid working options
Flexible working arrangements
Senior Manager – Resilience Culture & Capability lead
Senior Manager – Resilience Culture & Capability lead

ASX Limited • Australia

Hybrid
AUD 180,000 - 240,000
Hybrid working options
Operational Resilience Partner – Market Operations
Operational Resilience Partner – Market Operations

ASX Limited • Australia

On-site
AUD 120,000 - 150,000
Senior Quality & Test Engineer
Senior Quality & Test Engineer

ASX Limited • Australia

Hybrid
AUD 120,000 - 170,000
Hybrid working options
Technical Support Analyst
Technical Support Analyst

ASX Limited • Australia

On-site
AUD 120,000 - 150,000
Hybrid work options
Flexible working arrangements
Employee benefits
Senior Java Engineer - 12 Month Max Term Contract
Senior Java Engineer - 12 Month Max Term Contract

ASX Limited • Australia

Hybrid
AUD 180,000 - 240,000
Hybrid working options