Site Reliability Engineer

Precisely

Sydney

On-site

AUD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Precisely is seeking a Site Reliability Engineer to ensure reliability, performance, and scalability across our cloud and on‑prem environments. You will build automation, observability tooling, and reliability standards, guiding engineering teams to meet production readiness and incident controls.

You will work with Terraform, Ansible, Datadog, Python, and Bash to implement IaC, monitoring, and deployments, while aligning with security and data protection requirements.

Qualifications

  • Strong Linux skills in multi-site environments.
  • Hands-on IaC with Terraform and/or Ansible.
  • Experience in AWS cloud deployments (EC2, ECS, S3, VPC, IAM).
  • Scripting with Python or Bash for automation.
  • Monitoring/observability using Datadog or similar.
  • Understanding TCP/IP networking and load balancing.
  • Experience with CI/CD pipelines and deployment automation.
  • Ability to perform root cause analysis and post-incident reviews.

Responsibilities

  • Set and maintain reliability standards across services and environments.
  • Design and implement infrastructure-as-code and deployment automation.
  • Guide CI/CD practices and embed reliability into service design.
  • Lead incident response and post-incident analyses.
  • Produce root cause analyses and preventive automation.
  • Maintain runbooks and monitoring gaps, incidents, and risks.
  • Coach engineers on reliability and observability practices.

Skills

Linux proficiency
Terraform
Ansible
AWS (EC2, ECS, S3, VPC, IAM)
Python
Bash
Datadog
Networking fundamentals
CI/CD automation
SRE/ORR/SLO
Cross-functional collaboration

Education

Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent

Tools

Docker
Kubernetes
Git / GitOps

Job description

At Precisely, we're not just building software — we're shaping the future of data integrity. As a global leader in data quality, data enrichment, and location intelligence, Precisely helps thousands of the world's most trusted brands make confident decisions with data they can rely on. We're an AI-first organization, which means artificial intelligence isn't a buzzword here — it's woven into how we build products, how we work, and how we think about solving complex problems for our customers. When you join Precisely, you join a team of curious, driven innovators who believe that better data makes the world run better.

Overview:

The Site Reliability Engineer (SRE) is responsible for the reliability, performance, and scalability of Precisely's infrastructure platforms across CEDAR (CCX) — an on-premises, private cloud managed services environment; RapidCX (RCX) — an AWS cloud environment for SaaS-delivered customer communications management; and Hosted Managed Services (HMS) — an AWS cloud environment supporting managed client deployments.

This role bridges software engineering and systems operations, building automation, observability tooling, and reliability standards to ensure platform availability and operational excellence. SREs are enabling partners: they set reliability standards, define what 'reliable' looks like for each service, validate production readiness, and coach engineering teams on operational best practices. Engineering teams own the reliability outcomes of the services they build; the SRE ensures they have the standards, tooling, and guidance to meet them.

What you will do:
  • Set and maintain reliability standards across CCX, RCX, and HMS, including SLOs, SLIs, error budgets, alerting, logging, tracing, and MTTR improvements.
  • Build and maintain infrastructure-as-code, deployment automation, monitoring, and observability tooling using Terraform, Ansible, Datadog, Python, and Bash.
  • Guide CI/CD and deployment standards while partnering with engineering teams to embed reliability, scalability, backup, recovery, and failure‑mode planning into service design.
  • Review designs and lead Operational Readiness Reviews to validate production and disaster recovery readiness.
  • Lead major incident response, including incident command and clear stakeholder communication.
  • Produce root cause analyses, identify recurring issues, and implement preventive automation to reduce manual effort and improve reliability.
  • Maintain runbooks, reliability backlogs, and shared knowledge on monitoring gaps, incidents, lessons learned, and operational risks.
  • Use Precisely‑provided AI tools for infrastructure code, incident analysis, troubleshooting, testing runbooks, and documentation.
  • Ensure infrastructure meets security, compliance, vulnerability‑remediation, and data‑protection requirements, including applicable SOC 2 and FedRAMP standards.
  • Coach engineers, contribute to cross‑team reviews, stay current with SRE practices, and participate in the rotating on‑call schedule for critical escalations and changes.
What we are looking for:
Required:
  • Educational requirements (equivalent work experience will be accepted in place of the education requirement): Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
  • Years of experience: 3+ years of systems or infrastructure engineering experience in an enterprise production environment.
  • Years of experience needed with specific skills: Not separately specified beyond the overall experience requirement above.
  • Specific technical or software skills required:
  • Strong proficiency with Linux (RHEL/Oracle Linux) in a multi‑site, multi‑environment context.
  • Hands‑on experience with infrastructure‑as‑code tools: Terraform and/or Ansible.
  • Experience deploying and managing workloads in AWS (EC2, ECS, S3, VPC, IAM).
  • Proficiency with at least one scripting language (Python, Bash) for automation development.
  • Demonstrated experience building and maintaining monitoring and alerting systems (Datadog preferred).
  • Solid understanding of TCP/IP networking, DNS, load balancing, and distributed systems.
  • Experience with CI/CD pipeline design and deployment automation standards.
  • Strong analytical skills; ability to perform structured root cause analysis and post‑incident review.
  • Ability to define SLOs and lead Operational Readiness Reviews (ORRs); comfortable partnering with engineering teams on production readiness.
  • Demonstrated ability to work cross‑functionally with engineering teams on reliability standards and observability requirements.
  • Necessary certifications: None required (see Preferred Skills below).
  • Travel is required: No — approximately 0%.
AI Skills/Knowledge:

Active use of Precisely-provided AI tools (GitHub Copilot, Claude, or equivalent) for code development, troubleshooting, and documentation is a required baseline for this role, not a differentiator.

  • Apply AI tools for infrastructure‑as‑code development, incident analysis, and runbook authoring.
  • Use AI tools to accelerate troubleshooting and solution testing.
  • Maintain working fluency with Precisely‑approved AI coding assistants as part of daily practice.
Preferred Skills (a plus but not required):
  • Experience with containerization and orchestration (Docker, ECS, Kubernetes).
  • Familiarity with GitOps workflows and source control best practices (Git, GitLab).
  • Knowledge of enterprise virtualization platforms in a hybrid cloud context.
  • Understanding of change management and ITIL operational practices.
  • Experience with enterprise security tooling (Qualys, CrowdStrike, Rapid7).
  • AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification.
Application and Interview Impersonation Notice

Impersonating another individual when applying for employment, and/or participating in an interview process to assist another individual in obtaining employment, with Precisely Software Incorporated (“Precisely”) is unlawful. If Precisely identifies such fraudulent conduct, then as applicable and to the extent permitted by law, the application will be rejected, an offer (if made) will be rescinded, or the employment will be terminated, and legal action may be taken against the impersonators.

The personal data that you provide as a part of this job application will be handled in accordance with relevant laws. For more information about how Precisely handles the personal data of job applicants, please see thePrecisely Candidate Privacy Notice

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Precisely • City of Melbourne

On-site
AUD 120,000 - 180,000
AI-Driven SRE: On-Prem & Cloud Reliability Lead
AI-Driven SRE: On-Prem & Cloud Reliability Lead

Precisely • Australia

On-site
AUD 140,000 - 190,000
Support Engineer - B2B Integration
Support Engineer - B2B Integration

Precisely • Council of the City of Sydney

On-site
AUD 70,000 - 110,000
Site Reliability Engineer - AI-Driven Infra & Cloud
Site Reliability Engineer - AI-Driven Infra & Cloud

preciselyinternationaljobs • Australia

On-site
AUD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

HCLTech • City of Melbourne

On-site
AUD 120,000 - 160,000
Infrastructure Engineer, APAC
Infrastructure Engineer, APAC

DTEX Systems Inc. • Adelaide

On-site
AUD 110,000 - 170,000
Company computer hardware
Virtual events
Monthly Internet & Phone Reimbursement
+1
Software Engineering Manager, SRE, AI Foundations, Data Intelligence
Software Engineering Manager, SRE, AI Foundations, Data Intelligence

Google • Sydney

On-site
AUD 180,000 - 260,000
Infrastructure Engineer, APAC
Infrastructure Engineer, APAC

DTEX Systems • Adelaide

Hybrid
AUD 140,000 - 190,000
Company hardware of your choice
Virtual events & learning
Senior Solution Consultant
Senior Solution Consultant

New Relic, Inc. • City of Melbourne

On-site
AUD 140,000 - 210,000
Hybrid work model
Senior Site Reliability Engineer – AUS Region
Senior Site Reliability Engineer – AUS Region

Nucleus Security • City of Melbourne

On-site
AUD 150,000 - 210,000
Health insurance
Equity in startup
Flexible PTO
+1