Site Reliability Engineer - Data Infrastructure

Tribus

Sydney

On-site

AUD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Tribus in Sydney is seeking a Site Reliability Engineer focused on data infrastructure. You will operate and improve a multi-petabyte environment powering millions of queries daily, spanning Linux, Kafka, HDFS, Kubernetes, and distributed query platforms.

The role emphasizes production reliability, monitoring, on-call rotation, and automation using Python. You will collaborate with traders, researchers, and developers to solve data infrastructure problems and to deploy robust, scalable systems.

Qualifications

  • Hands-on Linux systems administration and troubleshooting experience.
  • Experience owning production systems, including monitoring, incidents and on-call.
  • Operator-side experience with Kafka, HDFS or Kubernetes.
  • Python experience for infrastructure automation or operational tooling.
  • Understanding of networking, storage, processes, memory and system performance.
  • Experience with infrastructure automation, CI/CD or configuration management.

Responsibilities

  • Operate and improve large-scale distributed data infrastructure.
  • Work with technologies including Kafka, HDFS, Kubernetes and distributed query platforms.
  • Own monitoring, alerting, incident response and production reliability.
  • Automate infrastructure deployment, upgrades and operational processes using Python.
  • Troubleshoot Linux, networking, storage and distributed systems issues.
  • Improve capacity, resilience, failure handling and deployment processes.
  • Work directly with traders, researchers and developers to solve data infrastructure problems.
  • Participate in an on-call rotation and engineer out recurring issues.

Skills

Linux admin
Production systems
Kafka
HDFS
Kubernetes
Python
Networking
CI/CD

Tools

CI/CD tooling

Job description

Site Reliability Engineer – Data Infrastructure

Sydney | Quantitative Trading | Linux, Kafka, Python

We are working with a leading global quantitative trading firm that is growing its Data Engineering team in Sydney.

This is an SRE role focused on the infrastructure that powers large-scale data platforms. You will help operate a multi-petabyte environment supporting millions of queries each day, working across Linux, distributed data systems, automation and production reliability.

This is not a traditional data engineering role focused on writing pipelines, DAGs or analytics workloads. The team owns and operates the underlying platforms themselves.

What you'll be doing

  • Operate and improve large-scale distributed data infrastructure
  • Work with technologies including Kafka, HDFS, Kubernetes and distributed query platforms
  • Own monitoring, alerting, incident response and production reliability
  • Automate infrastructure deployment, upgrades and operational processes using Python
  • Troubleshoot Linux, networking, storage and distributed systems issues
  • Improve capacity, resilience, failure handling and deployment processes
  • Work directly with traders, researchers and developers to solve data infrastructure problems
  • Participate in an on-call rotation and engineer out recurring issues

What we're looking for

  • Hands-on Linux systems administration and troubleshooting experience
  • Experience owning production systems, including monitoring, incidents and on-call
  • Operator-side experience with at least one of Kafka, HDFS or Kubernetes
  • Python experience, ideally for infrastructure automation or operational tooling
  • An understanding of networking, storage, processes, memory and system performance
  • Experience with infrastructure automation, CI/CD or configuration management

Experience with Kafka or HDFS administration is particularly valuable, but you do not need to know the entire technology stack.

Engineers coming from SRE, infrastructure, platform engineering, systems engineering or production engineering backgrounds are encouraged to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Tribus • Sydney

On-site
AUD 150,000 - 190,000
SRE - Data Infrastructure: Large-Scale Linux/Kafka Automation
SRE - Data Infrastructure: Large-Scale Linux/Kafka Automation

Tribus • Sydney

On-site
AUD 150,000 - 210,000
Site Reliability Engineer - Data Engineering
Site Reliability Engineer - Data Engineering

IMC Trading • Sydney

On-site
AUD 120,000 - 170,000
Site Reliability Engineer - Data Engineering
Site Reliability Engineer - Data Engineering

Linuxconfig • Australia

On-site
AUD 120,000 - 160,000
Data - Site Reliability Engineer
Data - Site Reliability Engineer

Optiver • Sydney

On-site
AUD 90,000 - 130,000
Performance-based bonus
Training & mentorship
Daily breakfast & in-house barista
+3
Data Platform SRE: Scale & Automate Data Pipelines
Data Platform SRE: Scale & Automate Data Pipelines

IMC Trading • Sydney

On-site
AUD 120,000 - 170,000
Data - Site Reliability Engineer
Data - Site Reliability Engineer

Optiver Private Jobs • Sydney

On-site
AUD 120,000 - 180,000
Performance-based bonus
Relocation package & visa sponsorship
Training & mentorship
+4
Real-Time Trading SRE — Resilient Infra & Automation
Real-Time Trading SRE — Resilient Infra & Automation

Tribus • Sydney

On-site
AUD 150,000 - 190,000
Data Platform SRE (Big Data, Kafka, Kubernetes)
Data Platform SRE (Big Data, Kafka, Kubernetes)

Linuxconfig • Australia

On-site
AUD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

Cloud Raptor • Sydney

On-site
AUD 180,000 - 240,000