Senior Site Reliability Engineer

Rabbit

Jakarta Pusat

On-site

IDR 500,000,000 - 750,000,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Rabbit seeks an experienced Site Reliability Engineer to advance AI-driven reliability across enterprise Google Cloud environments. You will own automated validation, progressive rollout, and safe releases while reducing operational toil.

You will define SLOs/SLIs, implement error budgets, and extend Terraform-driven infrastructure. Strong Go or Python coding, and AGI-aware tooling are essential for fast, dependable delivery.

Qualifications

  • 6+ years in SRE or production engineering or infrastructure-heavy backend roles.
  • Strong production GCP experience, including Cloud Run, networking and IAM.
  • Infrastructure-as-code fluency with Terraform, plus solid experience in CI/CD and deployment safety.
  • Strong observability and troubleshooting skills.
  • Coding ability in Go or Python for operational tooling.
  • Experience leading production incident investigation and driving follow-up improvements.
  • Fluency with AI coding agents and ability to review, test and validate outputs.
  • Strong written English and ability to collaborate asynchronously with a distributed team.

Responsibilities

  • Make AI-assisted delivery faster and safer with automated validation and progressive rollout.
  • Make reliability measurable with SLOs, SLIs and error budgets.
  • Automate workflows with AI agents for alert triage, incident investigation and maintenance.
  • Improve observability with better logging, metrics, tracing and alerting.
  • Turn incidents into lasting improvements via postmortems and tests.
  • Keep infrastructure reproducible by extending Terraform and delivery tooling.
  • Improve GCP reliability and efficiency, balancing performance and cost.
  • Use AI to accelerate reliability engineering with maintainable tooling.

Skills

Go
Python
GCP
Terraform
CI/CD
Observability
Incident response
English

Tools

Kubernetes
Datadog

Job description

Help Rabbit build and ship faster with AI — safely, securely and reliably.

Rabbit runs automated cost optimization across enterprise Google Cloud environments. When we change a customer's BigQuery reservations or rightsize their GKE clusters, those changes need to be correct and dependable. Reliability is central to the trust customers place in our product.

Our foundation is already in place: logging, alerting, automated deployment and Terraform-managed infrastructure. Your mission is to evolve that foundation for an AI-accelerated engineering team: turn faster implementation into faster, dependable delivery through automated validation, safe releases and rapid feedback.

You'll apply proven SRE practices — SLOs, observability, incident response and deployment safety — to AI-assisted development and agent-driven workflows. The goal is to increase how quickly the team can deliver verified improvements, while controlling production risk and reducing manual operational work.

What You'll Do
  • Make AI-assisted delivery faster and safer. Build automated validation, progressive rollout and recovery mechanisms that let engineers and agents move quickly with clear checks before and after changes reach production.
  • Make reliability measurable. Define and operationalize SLOs, SLIs and error budgets, and use them to guide practical decisions about delivery speed, stability and reliability work.
  • Automate workflows with AI agents. Identify repetitive operational work and build reusable agent-driven workflows for alert triage, incident investigation, routine maintenance and reporting. Add verification and human approval where needed, and measure the reduction in manual effort.
  • Improve observability and feedback. Evolve logging, metrics, tracing and alerting so failures are detected early and changes can be traced, investigated and verified.
  • Turn incidents into lasting improvements. Improve runbooks, investigation and blameless postmortems, and translate recurring problems into tests, safeguards and automation.
  • Keep infrastructure reproducible. Extend our Terraform and delivery tooling so environments remain consistent and changes stay reviewable as the platform grows.
  • Improve GCP reliability and efficiency. Strengthen our cloud infrastructure, networking, access controls and capacity management, balancing performance, reliability and cost.
  • Use AI to accelerate reliability engineering itself. Build maintainable tooling in Go, Python or a comparable language, and use agents to accelerate investigation, implementation, testing and documentation while verifying their outputs.
How We Work — AI-First, Agentic by Default

AI-assisted engineering is an expectation of this role, not an optional experiment. Claude Code, Cursor and agent-driven workflows are part of how we work, including infrastructure and reliability engineering.

We want someone who actively looks for ways to increase engineering speed with AI and makes those improvements safe to repeat. That means shorter feedback loops, automated checks, traceable changes and recovery paths — not simply generating more code.

You remain accountable for engineering judgment: what to automate, how to verify it, when human approval is needed and when a change should be stopped or rolled back. Success means faster delivery of reliable improvements, less repetitive work and a platform the team can trust.

What You'll Bring
Must-have
  • 6+ years in SRE, production engineering or infrastructure-heavy backend roles, with hands-on ownership of production systems.
  • Strong production GCP experience, including Cloud Run, networking and IAM. Hands-on Google Cloud experience is required and will be assessed during the interview process.
  • Infrastructure-as-code fluency with Terraform, plus solid experience in CI/CD and deployment safety.
  • Strong observability and troubleshooting skills: you can make systems debuggable, identify root causes and verify that a fix works.
  • Coding ability in Go, Python or a comparable language, with experience building maintainable operational tooling.
  • Experience leading production incident investigation and driving follow-up improvements that prevent recurrence.
  • Practical fluency with AI coding agents and the ability to critically review, test and validate their work. You are motivated to make AI-assisted engineering faster and more dependable.
  • Strong written English and the discipline to collaborate asynchronously with a distributed team.
Nice to have
  • Kubernetes / GKE experience, including deploying, operating, debugging and scaling containerized services.
  • Hands-on Datadog experience, including dashboards, monitors, logs, APM and distributed tracing.
  • Experience improving delivery speed and safety through progressive delivery, policy-as-code and automated rollback.
  • GCP cost management or FinOps experience.
  • Security experience is a plus.
Why This Role
  • Shape how an AI-first engineering team scales. Build the reliability practices and automation that let Rabbit turn faster development into dependable customer outcomes.
  • Work on systems with real customer impact. Rabbit operates inside enterprise GCP environments, where reliability and security directly affect customer trust.
  • Own meaningful improvements. Work in a small team with short decision paths and end-to-end ownership, supported by appropriate review and production safeguards.
  • Use AI as an engineering multiplier. Apply agentic tools to infrastructure, delivery and operations, and help define how we measure and improve their impact.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

amIT Global Solutions Sdn Bhd • Indonesia

On-site
IDR 420,000,000 - 660,000,000
AI-Driven SRE: Accelerate Reliable Delivery on GCP
AI-Driven SRE: Accelerate Reliable Delivery on GCP

Rabbit • Jakarta Pusat

On-site
IDR 500,000,000 - 750,000,000
AI Application Full Stack Engineer
AI Application Full Stack Engineer

Mekari • Jakarta Pusat

On-site
IDR 400,000,000 - 700,000,000
AI Systems Architect
AI Systems Architect

UltaHost • Jakarta Pusat

On-site
IDR 600,000,000 - 900,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Jakarta Pusat

On-site
IDR 500,000,000 - 900,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
IDR 167,400,000 - 279,000,000
Infra Engineer (AI Agents) | IDR 30-45m p/m
Infra Engineer (AI Agents) | IDR 30-45m p/m

Kulu • Denpasar

On-site
IDR 279,000,000 - 446,400,000
Infra Engineer (AI Agents) | IDR 25-45m p/m
Infra Engineer (AI Agents) | IDR 25-45m p/m

frontierinteractionscom • Denpasar

On-site
IDR 279,000,000 - 446,400,000
AI-First SRE: Build Fast, Safe, Reliable Deployments
AI-First SRE: Build Fast, Safe, Reliable Deployments

amIT Global Solutions Sdn Bhd • Indonesia

On-site
IDR 420,000,000 - 660,000,000
Infra Engineer (AI Agents) | IDR 25-40m p/m
Infra Engineer (AI Agents) | IDR 25-40m p/m

Kulu • Denpasar

On-site
Confidential