Site Reliability Engineer II

Zuora Community

United Kingdom

Hybrid

AUD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Medical, dental, and vision insurance
Generous time off and paid holidays

Job summary

Zuora Community is looking for a dynamic operations specialist to enhance automated solutions for our SaaS platform. Your role involves implementing intelligent operations, ensuring system reliability, and leveraging AI/ML for performance optimization. You'll work in a hybrid model - 3 days in the office, 2 days remote.

The ideal candidate has 2–4 years of experience in Linux systems administration, solid Python development skills, and hands-on Docker experience. Join us to drive innovation and excellence in operations.

Qualifications

  • 2–4 years of experience in Linux systems administration and/or Python development in production environments.
  • Strong Linux administration skills, including troubleshooting and performance tuning.
  • Hands-on experience with Docker and familiarity with Kubernetes concepts.

Responsibilities

  • Design and implement intelligent automation for infrastructure lifecycle management.
  • Apply AI/ML techniques for predictive monitoring and proactive performance optimization.
  • Lead complex incident response efforts and root cause analyses.

Skills

Linux systems administration
Python development
Docker
Kubernetes
CI/CD pipelines
Monitoring platforms

Job description

About Zuora

At Zuora, we help businesses grow smarter and adapt faster. Our platform powers modern business models—from subscriptions and usage‑based pricing to AI‑driven and outcome‑based offerings—helping companies launch new products, automate complex billing, and unlock predictable, recurring revenue.

Location

This position is located in the Zuora Costa Rica office (Heredia) and requires a hybrid schedule: 3 days in office and 2 days remote.

Opportunity

Join Zuora’s high‑impact Operations team to power the backbone of our industry‑leading SaaS platform. You will ensure the reliability, scalability, and performance of Zuora’s global production environment while building the next generation of intelligent operations.

Responsibilities
  • Design and implement intelligent automation for infrastructure lifecycle management, including self‑healing, anomaly detection, and automated remediation using Infrastructure as Code (IaC) and AI‑driven tooling.
  • Apply AI/ML techniques for predictive monitoring and proactive performance optimization to identify issues before they impact customers.
  • Lead complex incident response efforts and root cause analyses, embedding automation and continuous learning into operational processes.
  • Improve system reliability through dynamic scaling, telemetry instrumentation, and automated performance tuning.
  • Enhance operational runbooks and playbooks by eliminating manual processes through automation.
  • Evaluate and adopt emerging AIOps, cloud‑native, and distributed systems technologies to continuously improve our platform.
  • Partner cross‑functionally with Product Engineering, Customer Support, Global Services, Deal Desk, and Sales to deliver exceptional customer experiences.
Qualifications
  • 2–4 years of experience in Linux systems administration and/or Python development in production environments.
  • Strong Linux administration skills, including troubleshooting, service management, performance tuning, and networking fundamentals.
  • Experience developing Python scripts or lightweight applications to automate operational workflows and system management.
  • Hands‑on experience with Docker and familiarity with Kubernetes concepts, including deployments, services, and scaling.
  • At least one year of experience supporting SaaS or cloud‑native production environments.
  • Working knowledge of messaging platforms and databases such as Kafka, Redis, MySQL, or similar technologies.
  • Experience contributing to CI/CD pipelines and deployment automation.
  • Hands‑on experience with monitoring and observability platforms such as Prometheus, Grafana, or similar tools.
  • Experience participating in incident response, post‑incident reviews, and root cause analysis.
  • A demonstrated passion for automation and improving operational efficiency.
Nice to Have
  • Experience with Jenkins, Terraform, GitOps, or advanced Infrastructure as Code practices.
  • Exposure to AI/ML technologies for anomaly detection, predictive operations, or intelligent automation.
  • Relevant certifications such as RHCSA, AWS/Azure/GCP certifications, PCAP (Python), Docker Certified Associate (DCA), Certified Kubernetes Administrator (CKA), or SRE‑related certifications.
Benefits
  • Competitive compensation, variable bonus and performance‑based reward opportunities, and retirement programs.
  • Medical, dental, and vision insurance.
  • Generous, flexible time off, plus paid holidays, wellness days, and a company‑wide year‑end break.
  • Paid parental leave (including fully paid leave for eligible employees, subject to local policy).
  • Learning & development stipend to support ongoing growth.
  • Opportunities to volunteer and give back, including charitable donation matching where available.
  • Mental wellbeing resources and support.
  • Benefits may vary by location; details will be shared during the interview process.
About The Team

Zuora’s Operations team is responsible for keeping our global SaaS platform running reliably, securely, and at scale. We combine operational excellence with engineering best practices to build resilient systems that enable our customers to succeed.

Commitment to an Inclusive Workplace

We do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, disability status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply. Applicants needing special assistance during the interview process or to access our website may contact assistance@zuora.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Customer Solution Engineering
Manager, Customer Solution Engineering

Zuora • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive compensation
Medical, dental and vision insurance
Generous time off
+4
Senior Software Engineer London, Greater London, England, United Kingdom
Senior Software Engineer London, Greater London, England, United Kingdom

Zuora Inc • Greater London

On-site
GBP 70,000 - 90,000
Competitive compensation
Medical, dental, and vision insurance
Generous flexible time off
+2
Senior Software Engineer
Senior Software Engineer

Zuora • Greater London

On-site
GBP 60,000 - 80,000
Competitive compensation
Medical, dental, and vision insurance
Generous flexible time off
+2
SRE II: AI‑Driven Automation & Reliability
SRE II: AI‑Driven Automation & Reliability

Zuora Community • United Kingdom

Hybrid
AUD 90,000 - 120,000
Competitive compensation
Medical, dental, and vision insurance
Generous time off and paid holidays
Site Reliability Engineer (SRE) - Application Support
Site Reliability Engineer (SRE) - Application Support

ZILO™ • City Of London

Hybrid
GBP 60,000 - 90,000
Enhanced leave – 38 days including public holidays
Private Health Care including family cover
Life Assurance – 5x salary
+6
Sales Development Representative
Sales Development Representative

Zuora • Greater London

On-site
GBP 30,000 - 42,000
Company equity
Retirement programs
Medical, dental and vision insurance
+2
Support Engineer
Support Engineer

Zone- • United Kingdom

Remote
USD 100,000 - 140,000
Remote-first globally distributed team
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
DevOps Engineer
DevOps Engineer

Zuno Tech Group • Greater London

Hybrid
GBP 70,000 - 100,000
30 days annual leave + bank holidays
Private Medical Cover with Aviva *
4x salary Death in Service cover with 
+4
Sales Operations Manager
Sales Operations Manager

Zen Internet Limited • Rochdale

On-site
GBP 60,000 - 90,000
Life Assurance (2x)
Annual leave 25–30 days
Private Medical Healthcare
+7