Senior Site Reliability Engineer

Gradle Inc.

United States

Remote

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salaries
Equity grants
Remote work
In-person meetings
Ownership of production systems

Job summary

A leading software development company is seeking a Senior Site Reliability Engineer to ensure the availability and performance of their Develocity tool. You will be responsible for operating their cloud application platform and troubleshooting issues. The ideal candidate has strong Kubernetes experience and proficiency in observability tools. This is a full-time remote position offering a competitive salary range of $150-200K, based on location and qualifications.

Qualifications

  • 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
  • Strong experience in Kubernetes in production environments.
  • Expertise in cloud infrastructure, preferably AWS.

Responsibilities

  • Operate and maintain all Develocity instances and supporting services.
  • Participate in a follow-the-sun on-call rotation.
  • Drive automation across application deployment and upgrades.

Skills

SRE experience
Strong Kubernetes
Cloud infrastructure expertise
Observability tools
Incident management
Scripting proficiency
24/7 on-call experience
English communication

Tools

Terraform
Prometheus
Grafana

Job description

Join to apply for the Senior Site Reliability Engineer role at Gradle Inc.

Develocity is a first‑of‑its‑kind toolchain observability and acceleration platform that helps software teams adopt and improve DORA capabilities (including continuous delivery) in order to achieve software delivery excellence. It combines build and test acceleration with deep observability for builds and tests with Gradle Build Tool, Apache Maven™, sbt, npm, and Python, and applies to both CI and local builds and tests. Ultimately, Develocity provides an operational layer across an organization’s toolchains to speed up, troubleshoot, and optimize local developer and remote CI feedback loops.

Our software is used by some of the world’s leading software organizations, such as Netflix, Airbnb, SAP, several top ten banks, and many other major customers across all verticals. We regularly collaborate with these and other users to make our products continuously better.

We have partnered with the Apache Software Foundation, the Commonhaus Foundation, the Scala Center, the Micronaut Foundation, and other OSS projects like Spring, Quarkus, Kotlin, JUnit, AndroidX, and many more to bring the values of Develocity also to the OSS Community.

Our Values

Seek to Understand: Everything starts with listening and understanding, and we strive to understand different viewpoints, problems, and motivations. Before we take action, we ensure we truly grasp the challenges, perspectives, and goals.

Know the Why: We approach our work with a clear sense of purpose, ensuring every step is deliberate and focused. We take meaningful action with urgency, but never at the expense of thoughtful consideration.

Innovate & Iter. We embrace challenges and are not afraid to try new things, even if they might fail. With deep understanding and a clear purpose, we can develop creative and bold solutions to tackle challenges.

Own the Outcome: We are empowered to take initiative and we maintain transparency in our work and its outcomes. When we execute, we take responsibility for our decisions, measure the success of our innovations, and learn from the results.

Who You Are

We’re building a new SRE team and looking for founding members to help shape how we operate. You’ll be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open‑source projects, and public‑facing services, plus supporting infrastructure like artifact registries.

You’ll work on our internally‑built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it. When incidents happen, you’ll troubleshoot issues across the stack, from application to infrastructure. You’ll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the same task twice, you’ll fit in well.

You’ll be part of a distributed, remote‑first team that values asynchronous communication and written documentation. Strong self‑direction and clear communication across time zones are essential.

Responsibilities
  • Operate and maintain all Develocity instances and supporting services.
  • Participate in a follow‑the‑sun on‑call rotation, owning incident response and troubleshooting issues across the stack.
  • Drive automation across application deployment, upgrades, monitoring, self‑healing, and recovery.
  • Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
  • Work with engineering teams to build reliability into features from the start.
  • Run incident response and retrospectives, and make sure we learn from them.
  • Own disaster recovery, backups, and business continuity.
  • Communicate with customers during incidents and maintenance windows.
  • Optimize performance, resource usage, and costs.
  • Help evolve our SaaS operations as we grow.
Qualifications
  • 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
  • Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
  • Track record of incident management and response.
  • Knowledge of SRE best practices (SLAs, SLOs).
  • Scripting proficiency (Python, Bash) for automation.
  • Experience with 24/7 on‑call rotations.
  • Strong written and verbal English communication.
Preferred qualifications
  • Experience operating SaaS platforms at scale.
  • Familiarity with Develocity.
  • JVM language experience (Java, Kotlin).
  • Disaster recovery planning and execution experience.
  • Customer‑facing incident communication skills.
  • Experience establishing SRE practices in new or growing teams.
What We Offer
  • A ground‑floor role in a new SRE team—you’ll shape how we do things, not inherit someone else’s decisions.
  • Real ownership of production systems used by engineers at companies you’ve heard of.
  • Direct interaction with customers when things go wrong (and when they go right).
  • A culture that values automation over heroics.
  • In‑person meetings, such as our annual company off‑site and team meetings.
  • Work from home in a remote‑first environment.
  • Competitive salaries and equity grants.

Compensation

The U.S. salary range for this position is $150-200k which reflects the target ranges for all U.S. locations. Within this range, individual pay is determined by geographic location and additional factors including but not limited to experience, relevant skills, qualifications, seniority, performance, and travel requirements. Our recruiting team can share more information about the specific salary range for your location during the hiring process.

Location
  • Remote from anywhere in EST (Lead Engineer) or PST timezone (Senior Engineer).
  • While our team works remotely and is spread across the globe, we deeply value daily interactions and collaboration.
Seniority level

Mid‑Senior level

Employment type

Full‑time

Job function

General Business

Industries

Software Development

Referrals increase your chances of interviewing at Gradle Inc. by 2x

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding SRE for AI-Native Cloud Platform — Remote
Founding SRE for AI-Native Cloud Platform — Remote

Gradle Technologies • North Township (IN)

On-site
USD 150,000 - 190,000
Remote-first culture
Competitive salary + equity
Annual company offsite
Senior SRE - Remote-First, Reliability & Automation
Senior SRE - Remote-First, Reliability & Automation

Gradle Inc. • United States

Remote
USD 150,000 - 200,000
Senior DevOps Engineer
Senior DevOps Engineer

Rentvine • Estero (FL)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

Calabrio • United States

On-site
USD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • Town of Italy (NY)

On-site
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) with 4% match (US Only)
Health, Dental, Vision, Life Insurance
+8
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

Verint Systems, Inc. • Frankfort (KY)

Hybrid
USD 65,000 - 90,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Madrona Venture Labs • United States

Remote
USD 120,000 - 150,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1