Site Reliability Engineer

Cover Genius

Sydney

On-site

AUD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cover Genius is seeking a Site Reliability Engineer to lead reliability and infrastructure across multiple teams. You will shape system design, tooling, and processes for scalable production systems and ensure rapid, safe delivery of features.

You will bring strong cloud, observability, and automation skills, with a track record in Kubernetes, IaC, AWS/GCP, and AI-enabled development. The role emphasizes collaboration, security, and proactive risk reduction.

Qualifications

  • 3+ years of experience in SRE, Platform Engineering, DevOps or related roles.
  • Strong understanding of SRE and platform engineering principles with cross-team project leadership.
  • Experience with observability tools such as Datadog, Elasticsearch, Prometheus, Grafana.
  • Experience with cloud native tech: Docker, Kubernetes; infrastructure as code with Terraform.
  • Scripting and internal tooling with Bash and at least one language (Python/Go).
  • Fluency with AI-assisted development environments (Cursor, Claude Code, Codex) and applying them to production workflows.
  • Experience with Linux, networking, distributed systems and high-availability deployments.
  • Bachelor's degree in CS/Engineering or equivalent desirable.

Responsibilities

  • Lead reliability and infrastructure projects across teams and domains.
  • Architect and build reliable, highly-available cloud infrastructure across products.
  • Develop observability standards, tooling, and dashboards for multiple teams.
  • Champion SLOs and incident response; act as incident commander for major outages.
  • Reduce toil through automation and self-service tooling.
  • Develop and maintain runbooks and incident playbooks.
  • Apply AI-assisted development to infrastructure problems and guide teams.
  • Contribute to capacity planning and cost optimization.
  • Mentor engineers on systems thinking and production ownership.
  • Implement security best practices in pipelines and infra.

Skills

SRE
Platform Eng
DevOps
Observability
Kubernetes
Terraform
Python/Go
AI tools
Linux
Networking
AWS/GCP

Education

Bachelor's degree in Computer Science/Engineering

Tools

Datadog
Elasticsearch
Prometheus
Grafana
Docker
Kubernetes
Terraform
Bash
Python
Go
Cursor
Claude Code
Codex

Job description

Cover Genius is the global infrastructure for embedded protection. Active in over 60 countries and all 50 US States, we protect the customers of the world’s largest digital companies, including Klarna, Revolut, Stripe, Priceline, Agoda, Booking.com, Turkish Airlines, Tongcheng Travel, eBay, and Uber, with seamless, end-to-end experiences. Cover Genius has protected more than 73M customers globally across 240M policies with USD $3.2BN in gross written sales.

Coming off a stellar year with 40% YoY revenue growth and a recent $100M capital raise, putting our valuation at $1.9BN, we are accelerating into our next phase of growth. As part of our team, you’ll help drive our AI-first roadmap, developing hyper-personalization engines, agentic distribution, and automated claims infrastructure, while building the scalable technology powering the fast-growing $70B embedded protection market.

Our people are: Accountable, customer-obsessed, collaborative, driven
Our people are not: Passive, defensive, siloed, hesitant
About the Company

Cover Genius is the global infrastructure for embedded protection. Active in over 60 countries and all 50 US States, we protect the customers of the world’s largest digital companies, including Klarna, Revolut, Stripe, Priceline, Agoda, Booking.com, Turkish Airlines, Tongcheng Travel, eBay, and Uber, with seamless, end-to-end experiences. Cover Genius has protected more than 73M customers globally across 240M policies with USD $3.2BN in gross written sales.

Coming off a stellar year with 40% YoY revenue growth and a recent $100M capital raise, putting our valuation at $1.9BN, we are accelerating into our next phase of growth. As part of our team, you’ll help drive our AI-first roadmap, developing hyper-personalization engines, agentic distribution, and automated claims infrastructure, while building the scalable technology powering the fast-growing $70B embedded protection market.

Our people are: Accountable, customer-obsessed, collaborative, driven
Our people are not: Passive, defensive, siloed, hesitant
About the Role

As a Site Reliability Engineer, you'll lead reliability and infrastructure projects that span multiple teams and business domains. Decisions you make on system design, tooling, and process will directly shape how the teams you work with build and operate at scale.

To drive success in this role, you will have a strong background in cloud infrastructure and platform engineering, with experience across infrastructure-as-code, CI/CD and release automation, observability, security, and disaster recovery. You should possess strong technical judgement, the ability to independently lead complex projects from design through delivery, and a proactive approach to eliminating operational risk before it becomes a problem.

Regular collaboration with software engineering teams, security teams, and other relevant stakeholders will be key in ensuring the reliability and efficiency of our production systems are achieved.

Key Responsibilities
  • Analyze, test, and evolve systems to improve reliability and performance at an architectural/infrastructure level, leading medium-to-large projects from design through delivery

  • Apply AWS and GCP expertise to architect and build reliable, highly-available cloud infrastructure across multiple products and projects

  • Implement observability strategy and standards, developing the tooling and dashboards other teams build on

  • Champion SLOs and error budgets for services in your area, using them to prioritise reliability work against feature velocity

  • Act as incident commander for significant production incidents, lead troubleshooting on complex issues, using blameless post-mortems to drive continuous improvement

  • Reduce operational toil by building automation and self-service tooling, that other engineers can adopt, rather than absorbing repetitive work yourself

  • Develop and maintain design, troubleshooting, and runbook standards that other engineers can follow

  • Apply AI-assisted development to infrastructure problems, and guide other engineers on using AI tools effectively within your team's workflows

  • Contribute to capacity planning and cost optimisation

  • Mentor other engineers on systems thinking, incident management, and production ownership, raising the bar within your team

  • Implement security best practices in infrastructure development and maintenance - such as least-privilege access, secrets management, policy-as-code gates in your pipelines

Skills & Experience
What you will bring:
  • 3+ years of experience in SRE, Platform Engineering, DevOps or other related roles

  • Strong understanding of SRE and platform engineering principles, with experience applying them to lead cross-team projects

  • Experience using and configuring modern observability tools such as Datadog, Elasticsearch, Prometheus, Grafana

  • Experience with cloud native and container technology such as Docker, and hands-on experience using and managing Kubernetes clusters

  • Experience authoring reusable, parameterised infrastructure-as-code modules (e.g. Terraform)

  • Comfortable scripting and developing internal tooling with Bash and at least one programming language (e.g. Python, Go)

  • Fluent with AI-driven development environments like Cursor, Claude Code, or Codex, with a proven ability to leverage these tools within production engineering workflows

  • Experience working with Linux

  • Strong understanding of networking, distributed systems, and system architecture at scale

  • Proven experience deploying, scaling, and monitoring web applications and databases in high-availability environments

  • Deep expertise in AWS and/or GCP platforms

  • Bachelor's degree in Computer Science/Engineering, a postgraduate degree and/or record of academic achievement is also desirable

What you will have:
Ownership & Delivery
  • Takes ambiguous problems and drives them to shipped outcomes - not just code, but results, and takes accountability even without a clear owner

  • Balances speed with quality — knows when to iterate fast and when to invest in durability

  • Manages risk proactively — identifies failure modes and mitigates before they bite

Communication & Influence
  • Creates clarity from ambiguity; documents decisions so others can build on your work

  • Influences through evidence and collaboration, not authority — mentors and unblocks teammates

  • Communicates technical concepts clearly to engineers, product, and business stakeholders

AI-First Mindset
  • Treats AI tools as essential infrastructure, not optional add-ons — continuously experiments with new capabilities

  • Understands LLM strengths and limitations — knows when to prompt and when to build differently

  • Thinks in leverage: automates the repetitive, focuses human attention on judgement calls

  • Helps others adopt AI workflows, shares what works, and raises the floor for the whole team

Why Cover Genius?

At Cover Genius, we create magic by turning the archaic into the extraordinary. We take one of the world’s oldest and most complex industries and reinvent it with world-class technology. Cover Genius doesn’t just disrupt legacy insurance; we make the impossible feel effortless.Our operating principles guide our mission:

  • Make today matter: We deliver with urgency and excellence. We act decisively, move with intention, and hold a high bar.

  • Act with accountability: We own our commitments and take pride in delivering results that move us forward.

  • Grow together: We are a collective of curious minds. We learn from wins and setbacks, share knowledge generously, and elevate each other every day.

  • Inspire Each Other: We push each other to think bigger and pull together to go further.

  • Champion Our Customers: We lead with empathy and center the customer, turning complex disruptions into seamless moments of trust.

  • Flexible Work Environment - Our teams are hybrid. We work from home on a Wednesday and Thursday and attend the office on Monday, Tuesday and Friday with flexibility around start/finish times.

  • Don't just take our word for it - hear about our culture straight from our people here: story time.

Ready to make an impact? If you’re looking for a place where you’ll be challenged, trusted, and empowered, we’d love to meet you.

The Legal & Privacy Stuff

Cover Genius promotes diversity and inclusivity. We don't tolerate discrimination, demeaning treatment of anyone, or harassment due to race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or any other legally protected status.

By submitting your application, you acknowledge that we may collect, store, and process your personal data for recruitment purposes. To ensure a fair evaluation, we may use AI to assist in sorting applications, but all final decisions are made by our hiring team and no candidate dispositions are automated. We will keep your information on file for three years from the date of your application. For detailed information about how we handle your data and our use of AI, please review our full Privacy Policy.

Be careful - Don’t provide your bank or credit card details when applying for jobs. Don't transfer any money or complete suspicious online surveys. If you see something suspicious, report this job ad .

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Manager — Customer Operations & AI Tools
Engineering Manager — Customer Operations & AI Tools

Cover Genius • Sydney

Hybrid
AUD 150,000 - 190,000
Hybrid work environment
Employee stock options
Global office access
Lead, Developer Experience Engineering
Lead, Developer Experience Engineering

Neara • Sydney

Hybrid
AUD 120,000 - 150,000
Flexible Work Environment
Employee Stock Options
Wellness day
+1
Engineering Manager, Customer Operations (Hybrid)
Engineering Manager, Customer Operations (Hybrid)

King River Capital Group • Sydney

Hybrid
AUD 180,000 - 240,000
Hybrid work
Employee stock options
Global offices
+1
Senior Software Engineer (Backend focussed)
Senior Software Engineer (Backend focussed)

King River Capital Group • Sydney

On-site
AUD 150,000 - 200,000
Senior Software Engineer (Backend focussed)
Senior Software Engineer (Backend focussed)

Cover Genius • Sydney

Hybrid
AUD 120,000 - 180,000
Hybrid work model
Engineering Manager, Direct-to-Consumer
Engineering Manager, Direct-to-Consumer

King River Capital Group • Sydney

On-site
AUD 180,000 - 240,000
Hybrid work arrangement
Flexible work schedule
Software Engineer III - Logistics
Software Engineer III - Logistics

King River Capital Group • City of Brisbane

On-site
CAD 115,000 - 145,000
Hybrid work model
Flexible work from home days
Office on-site days
Enterprise Voice and Multi-Channel Support Specialist
Enterprise Voice and Multi-Channel Support Specialist

Cover Genius • Sydney

Hybrid
AUD 60,000 - 85,000
Hybrid work environment
Employee stock options
Office flexibility
+1
Engineering Manager, Customer Operations
Engineering Manager, Customer Operations

King River Capital Group • Sydney

On-site
AUD 180,000 - 240,000
Hybrid work
Employee stock options
Global offices
+1
Senior Software Engineer (TypeScript, React, NestJS)
Senior Software Engineer (TypeScript, React, NestJS)

Cover Genius • Sydney

On-site
AUD 180,000 - 240,000
Flexible work environment
Global company with offices worldwide
Employee stock options
+1