Site Reliability Engineer

Tribus

Sydney

On-site

AUD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Tribus in Sydney is seeking a hands-on Site Reliability Engineer (Trading Infrastructure) to design, build and maintain Linux-based production systems powering a global electronic trading platform.

You will implement IaC with Terraform and Ansible, build observability with Prometheus, Grafana and Splunk, and collaborate with software engineers across C++, Java, C# and Python components to improve latency, automation and incident response.

Qualifications

  • Strong Linux systems administration and troubleshooting experience.
  • Python software development skills, specifically building tools for internal teams
  • Experience with Infrastructure as Code using Terraform, Ansible or similar tools.
  • Hands-on experience with observability platforms such as Prometheus, Grafana or Splunk.
  • Experience designing monitoring strategies that focus on operational events, anomaly detection and actionable alerts.
  • Familiarity with PostgreSQL, InfluxDB or other production databases.
  • Strong networking fundamentals including TCP/IP, routing and firewalls.
  • Experience with Docker and modern development tooling.
  • Exposure to virtualisation platforms such as VMware, KVM, Proxmox or CEPH.
  • Experience supporting low-latency, distributed or mission-critical production environments is highly regarded.

Responsibilities

  • Design, build and maintain highly reliable Linux-based production infrastructure.
  • Develop Infrastructure as Code using Terraform and Ansible.
  • Build observability platforms using Prometheus, Grafana and Splunk.
  • Create event-driven monitoring, intelligent alerting and automated remediation workflows.
  • Improve operational tooling and incident response across business-critical trading systems.
  • Work with PostgreSQL and InfluxDB supporting production data platforms.
  • Support workflow orchestration using Prefect.
  • Administer virtualisation and storage platforms including CEPH, VMware, KVM and Proxmox.
  • Partner closely with software engineers developing C++, Java, C# and Python applications.
  • Improve CI/CD, deployment tooling and overall platform reliability through automation.

Skills

Linux
Python
Terraform
Ansible
Prometheus
Grafana
Splunk
PostgreSQL
InfluxDB
Docker
VMware
KVM
Proxmox
Networking fundamentals
Collaboration with C/C++/Java/C#

Tools

CEPH
NVIDIA GPUS

Job description

Site Reliability Engineer (Trading Infrastructure) Multiple organisations

Sponsorship Available

Sydney or Hong Kong | Onsite | Global Trading Environment

Join a high-performance engineering team responsible for the reliability, scalability and operational excellence of the infrastructure powering a global electronic trading platform.

This is a hands-on Site Reliability Engineering role sitting close to the trading stack, where you'll help build resilient production systems, improve automation, and work alongside software engineers to support latency-sensitive applications operating across global financial markets.

Unlike traditional SRE environments focused primarily on SLI/SLO metrics, this team takes an event-driven approach to reliability engineering. You'll design intelligent monitoring and automated operational workflows that identify abnormal system behaviour, infrastructure anomalies and production events before they impact trading. The focus is on actionable signals, rapid diagnosis and engineering-led remediation rather than simply measuring service health.

What You'll Be Doing
  • Design, build and maintain highly reliable Linux-based production infrastructure.
  • Develop Infrastructure as Code using Terraform and Ansible.
  • Build observability platforms using Prometheus, Grafana and Splunk.
  • Create event-driven monitoring, intelligent alerting and automated remediation workflows.
  • Improve operational tooling and incident response across business-critical trading systems.
  • Work with PostgreSQL and InfluxDB supporting production data platforms.
  • Support workflow orchestration using Prefect.
  • Administer virtualisation and storage platforms including CEPH, VMware, KVM and Proxmox.
  • Partner closely with software engineers developing C++, Java, C# and Python applications.
  • Improve CI/CD, deployment tooling and overall platform reliability through automation.
What We're Looking For
  • Strong Linux systems administration and troubleshooting experience.
  • Python software development skills, specifically building tools for internal teams
  • Experience with Infrastructure as Code using Terraform, Ansible or similar tools.
  • Hands-on experience with observability platforms such as Prometheus, Grafana or Splunk.
  • Experience designing monitoring strategies that focus on operational events, anomaly detection and actionable alerts rather than purely SLI/SLO-driven metrics.
  • Familiarity with PostgreSQL, InfluxDB or other production databases.
  • Strong networking fundamentals including TCP/IP, routing and firewalls.
  • Experience with Docker and modern development tooling.
  • Exposure to virtualisation platforms such as VMware, KVM, Proxmox or CEPH.
  • Experience supporting low-latency, distributed or mission-critical production environments is highly regarded.
Why This Role?
  • Work on infrastructure that directly supports real-time trading.
  • Solve complex reliability challenges where milliseconds matter.
  • Influence how monitoring, automation and operational engineering are built from the ground up.
  • Collaborate with experienced infrastructure and software engineers in a highly technical environment.
  • Work in a culture that values engineering ownership, continuous improvement and pragmatic problem solving.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Data Infrastructure
Site Reliability Engineer - Data Infrastructure

Tribus • Sydney

On-site
AUD 150,000 - 210,000
Real-Time Trading SRE — Resilient Infra & Automation
Real-Time Trading SRE — Resilient Infra & Automation

Tribus • Sydney

On-site
AUD 150,000 - 190,000
Site Reliability Engineer - Data Engineering
Site Reliability Engineer - Data Engineering

IMC Trading • Sydney

On-site
AUD 120,000 - 170,000
Site Reliability Engineer - Data Engineering
Site Reliability Engineer - Data Engineering

ittihad medical centre • Sydney

On-site
AUD 140,000 - 190,000
Data - Site Reliability Engineer
Data - Site Reliability Engineer

Optiver • Sydney

On-site
AUD 90,000 - 130,000
Performance-based bonus
Training & mentorship
Daily breakfast & in-house barista
+3
Data - Site Reliability Engineer
Data - Site Reliability Engineer

Optiver Private Jobs • Sydney

On-site
AUD 120,000 - 180,000
Performance-based bonus
Relocation package & visa sponsorship
Training & mentorship
+4
Site Reliability Engineer
Site Reliability Engineer

Firesoft People • Sydney

On-site
AUD 210,000 - 280,000
Production Reliability Engineer
Production Reliability Engineer

ASX • Sydney

Hybrid
AUD 150,000 - 190,000
Flexible working
Hybrid work options
Platform Reliability Engineer — Observability & Automation
Platform Reliability Engineer — Observability & Automation

ASX • Sydney

Hybrid
AUD 150,000 - 190,000
Flexible working
Hybrid work options
DevOps Engineer
DevOps Engineer

Cloud Raptor • Sydney

On-site
AUD 180,000 - 240,000