Site Reliability Engineer - Data Engineering

Talenza

Sydney

On-site

AUD 120,000 - 160,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Talenza in Sydney is seeking a Site Reliability Engineer to own and operate a large-scale data platform underpinning trader research and decision-making.

You will run and improve platforms like Kafka and HDFS, troubleshoot Linux-based issues, and build automation with CI/CD to accelerate deployments. This role emphasizes hands-on infrastructure depth over managed services and offers growth in a global trading context.

Qualifications

  • Approximately 2-3 years in an SRE/platform/infra/production engineering role.
  • Strong Linux troubleshooting across processes, filesystems, networking, memory pressure.
  • Hands-on operator experience with Kafka, HDFS or Kubernetes at platform level.
  • Experience operating self-managed infrastructure (bare metal, datacentre, self-run VMs or self-managed Kubernetes).
  • Python for systems automation, health checks or deployment workflows.
  • Exposure to Docker, Kubernetes, Helm and infrastructure-focused CI/CD.
  • Curious, pragmatic mindset with interest in how complex systems behave under pressure.
  • Experience with Dremio, Presto, Airflow, Prefect, Ansible, Puppet, Terraform or cloud platforms is beneficial but non-essential.

Responsibilities

  • Run, monitor and improve large-scale data platforms including Kafka, HDFS and internally built pipelines.
  • Troubleshoot production issues across Linux, storage, networking and distributed infrastructure.
  • Build automation and CI/CD to speed up, improve safety and repeatability of deployments.
  • Support upgrades, capacity planning, incident response and long-term reliability improvements.
  • Collaborate with systems and network engineers, developers, researchers and end users to solve complex data-platform problems.
  • Evaluate and introduce new technology as the environment evolves.

Skills

Linux troubleshooting
Python automation
Docker/Kubernetes/Helm
CI/CD
Curious mindset

Tools

Kafka
HDFS
Kubernetes

Job description

Site Reliability Engineer - Data Platform

Sydney | Global Trading Firm

Join a global trading firm's high-performing Data Engineering team and help operate the large-scale platform that underpins trader research, simulation, reporting and decision-making.

This is a proper infrastructure role-not one for someone who has only built data pipelines on top of managed services. You'll get hands-on with the underlying platform: Kafka, HDFS, Dremio, Linux and in-house data tooling across a multi-petabyte environment processing around two million queries each day.

You'll join a small, experienced Sydney team with international engineering counterparts, giving you genuine ownership, strong mentoring and exposure to complex distributed systems at meaningful scale.

The role

Run, monitor and improve large-scale data platforms including Kafka, HDFS, Dremio and internally built pipelines.

Troubleshoot real production issues across Linux, storage, networking and distributed infrastructure.

Build automation and CI/CD capability to make deployments faster, safer and more repeatable.

Support upgrades, capacity planning, incident response and long-term reliability improvements.

Work closely with systems and network engineers, developers, researchers and end users to solve complex data-platform problems.

Help evaluate and introduce new technology as the environment continues to evolve.

What we're looking for

Around 2-3 years' experience in an SRE, platform, infrastructure, systems or production engineering role.

Strong Linux troubleshooting skills across processes, filesystems, networking, disk and memory pressure.

Hands-on operator experience with at least one of Kafka, HDFS or Kubernetes-you have deployed, configured, upgraded, tuned or supported the platform itself.

Experience operating self-managed infrastructure, whether bare metal, datacentre, self-run VMs or self-managed Kubernetes.

Python experience for systems automation, operational tooling, health checks or deployment workflows.

Exposure to Docker, Kubernetes, Helm and infrastructure-focused CI/CD.

A curious, pragmatic mindset and a genuine interest in understanding how complex systems behave under pressure.

Experience with Dremio, Presto, Airflow, Prefect, Ansible, Puppet, Terraform or cloud platforms would be beneficial, but it is the operational mindset and underlying Linux/infrastructure depth that matter most.

This is a standout opportunity for an engineer who wants to move beyond managed services, get close to the underlying technology and build a career operating high-scale, business-critical systems.

Site Reliability Engineer - Data Engineering Sydney, NSW, AU

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Data Engineering
Site Reliability Engineer - Data Engineering

IMC Trading • Sydney

On-site
AUD 120,000 - 170,000
Site Reliability Engineer - Data Engineering
Site Reliability Engineer - Data Engineering

IMC B.V. • Sydney

On-site
AUD 120,000 - 180,000
Site Reliability Engineer, Data Platform - Scale & Tech
Site Reliability Engineer, Data Platform - Scale & Tech

Talenza • Sydney

On-site
AUD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Tribus • Sydney

On-site
AUD 150,000 - 190,000
Data - Site Reliability Engineer
Data - Site Reliability Engineer

Optiver • Sydney

On-site
AUD 90,000 - 130,000
Performance-based bonus
Training & mentorship
Daily breakfast & in-house barista
+3
SRE - Data Infrastructure: Large-Scale Linux/Kafka Automation
SRE - Data Infrastructure: Large-Scale Linux/Kafka Automation

Tribus • Sydney

On-site
AUD 150,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Firesoft People • Sydney

On-site
AUD 230,000 - 320,000
Data Platform SRE: Scale & Automate Data Pipelines
Data Platform SRE: Scale & Automate Data Pipelines

IMC Trading • Sydney

On-site
AUD 120,000 - 170,000
Lead Platform Engineer
Lead Platform Engineer

Halcyon Knights • City of Melbourne

Hybrid
AUD 180,000 - 240,000
Hybrid working environment
Global footprint
Career growth into tech leadership
Practice Manager - Platform Reliability, Operations Hub & Automation
Practice Manager - Platform Reliability, Operations Hub & Automation

Datacom • City of Brisbane

On-site
AUD 180,000 - 240,000
Social events
Chill-out spaces
Remote working
+2