Senior Production Reliability Engineer – Cloud Infrastructure & Automation

Westbury Partners

Sydney

On-site

AUD 140,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Westbury Partners in Sydney is seeking an experienced reliability/infrastructure engineer to design, operate, and evolve highly available cloud production systems.

You will automate provisioning, deployments, and monitoring, define SLOs/SLIs, lead on-call responses, and drive practices that improve resilience and deployment safety.

Qualifications

  • Linux-based production experience with cloud infrastructure.
  • Proficient in Python/Go/Bash scripting.
  • Hands-on with Kubernetes, Docker, and Terraform.

Responsibilities

  • Design, operate, and improve scalable production infrastructure.
  • Build automation for provisioning, deployments, configuration, and workflows.
  • Define SLOs, SLIs, reliability metrics, and error budgets.
  • Strengthen monitoring, logging, alerting, tracing, and observability.
  • Lead on-call incident response and post-incident reviews.
  • Improve deployment safety, change management, and release processes.
  • Identify bottlenecks and technical debt to optimise resilience.
  • Support capacity planning, disaster recovery, performance testing, and resilience initiatives.
  • Create infrastructure-as-code, runbooks, and documentation.

Skills

Linux
Cloud platforms
Python
Go
Bash
Kubernetes
Docker
Terraform
CI/CD

Tools

Kubernetes
Docker
Terraform

Job description

Westbury Partners – Sydney NSW

Full time

6d ago , from E-Financial Careers Australia

Drive reliability across critical production systems by engineering scalable cloud infrastructure, automating operations, strengthening observability, leading incident response, and enabling engineering teams to deliver safer, resilient services.

What You'll Do:
  • Design, operate, and continuously improve highly available, scalable, secure production infrastructure.
  • Build automation for provisioning, deployments, configuration, and operational workflows.
  • Establish meaningful SLOs, SLIs, reliability metrics, and error budgets.
  • Strengthen monitoring, logging, alerting, tracing, and overall observability.
  • Troubleshoot complex production issues and participate in on-call incident response.
  • Lead post-incident reviews and turn recurring problems into lasting improvements.
  • Improve deployment safety, rollback strategies, change management, and release processes.
  • Identify infrastructure bottlenecks, reliability risks, technical debt, and opportunities for optimisation.
  • Support capacity planning, disaster recovery, performance testing, and resilience initiatives.
  • Create infrastructure-as-code, runbooks, documentation, and operational best practices.
Your responsibilities will include:
  • You’ll take ownership of critical production environments while partnering with software, security, and infrastructure teams. You’ll automate repetitive operational work, improve system resilience, reduce toil, and help engineering teams adopt reliability-focused practices.
  • You’ll work across cloud platforms, Linux, containers, Kubernetes, infrastructure-as-code, CI/CD, observability, networking, databases, and distributed systems. Your work will directly contribute to improved availability, faster incident recovery, safer deployments, scalability, and infrastructure efficiency.
Why Join Us:
  • This is an opportunity to solve challenging production infrastructure problems while having a direct influence on how systems are designed, deployed, monitored, and operated.
  • You’ll be empowered to build meaningful automation, improve engineering practices, strengthen reliability, and create tools that make life easier for development teams while delivering a better experience for customers.
About You:
  • You’re an experienced reliability, infrastructure, DevOps, or production engineer with strong Linux and cloud expertise. You’re comfortable writing code or scripts in Python, Go, Bash, or similar languages and have hands‑on experience with technologies such as Kubernetes, Docker, Terraform, and modern CI/CD platforms.
  • You understand networking, DNS, HTTP/TLS, load balancing, databases, distributed systems, monitoring, and observability. You’re analytical, collaborative, calm during incidents, and naturally driven to automate problems rather than repeatedly work around them.
  • You take ownership, communicate clearly, enjoy solving complex technical challenges, and continuously look for ways to make production systems safer, simpler, and more reliable.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bell Financial Group (ASX:BFG) • City of Melbourne

Hybrid
AUD 140,000 - 190,000
Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

Intrinzic • Sydney

Hybrid
AUD 140,000 - 200,000
Blockchain Infrastructure & Reliability Engineer
Blockchain Infrastructure & Reliability Engineer

Westbury Partners • Sydney

On-site
AUD 180,000 - 240,000
DevOps Engineer
DevOps Engineer

TechForce Services • Sydney

On-site
AUD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Tribus • Sydney

On-site
AUD 150,000 - 190,000
Senior Trade Operations & Infrastructure Engineer
Senior Trade Operations & Infrastructure Engineer

Westbury Partners • Sydney

On-site
AUD 170,000 - 210,000
Blockchain Infrastructure & Production Engineer
Blockchain Infrastructure & Production Engineer

Westbury Partners • Sydney

On-site
AUD 120,000 - 180,000
Production Engineer - Trading Systems
Production Engineer - Trading Systems

Selby Jennings • Sydney

On-site
AUD 150,000 - 210,000
AWS Platform Engineer
AWS Platform Engineer

BGL Corporate Solutions • City of Melbourne

Hybrid
AUD 140,000 - 190,000
Hybrid working
Development budget
Paid study leave
+3
Senior DevOps Engineer - GCP or AWS
Senior DevOps Engineer - GCP or AWS

Ruby Talent • Sydney

Hybrid
AUD 140,000 - 190,000
Equity / share options
Amazing benefits