Weekend Site Reliability Engineer

Sporty Group

United States

Remote

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote first
Bonuses (quarterly)
28 days leave
Core hours 10am-3pm
Referral bonuses
Equipment provided
Annual retreats

Job summary

Sporty Group is seeking a Weekend Site Reliability Engineer to join a DevOps-led team. The role covers Saturday, Sunday and Monday with flexible days off, focusing on stabilizing Kubernetes, cloud infra, and GitOps-driven deployments across multiple countries.

You will own weekend on-call, monitor dashboards (Grafana), implement alerting pipelines, define SLIs/SLOs, mentor colleagues, and coordinate audits with security teams.

Qualifications

  • 3+ years DevOps / platform engineering experience.
  • Based in Europe or Asia or LatAM
  • Experience leading the planning and deployment of a project
  • Strong AWS knowledge and ability to provision resources to meet demand
  • Kubernetes experience with EKS and GitOps tools (ArgoCD, Helm)
  • IaC experience with Terraform
  • Scripting in Bash/Python/Go; Rust a plus
  • Observability stacks: Prometheus, Grafana, Loki, Tempo, OpenTelemetry
  • RUM experience a plus; Grafana Faro or OpenTelemetry SDK instrumentation a plus
  • On-call and incident response experience; post-mortems and follow-ups

Responsibilities

  • Weekend SRE covering Saturday, Sunday and Monday with flexible days off.
  • Improve infrastructure and deployments across deployed countries.
  • Improve Kubernetes platform stability, cost efficiency, and provisioning via GitOps.
  • Monitor and maintain cloud infra with autoscaling, alerting pipelines, and Grafana dashboards.
  • Own weekend on-call, triage incidents, perform root cause analysis, and drive post-incident reviews.
  • Design alert pipelines to minimize noise and avoid waterfall alerting.
  • Define and maintain SLIs/SLOs to drive reliability improvements.
  • Mentor junior team members.

Skills

DevOps
Cloud platforms
Kubernetes
GitOps
Scripting

Tools

AWS
Terraform
ArgoCD
Helm

Job description

What You’ll Be Doing
  • Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
  • Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
  • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
  • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
  • Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
  • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
  • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
  • Take ownership and responsibility for our cloud operation activities
  • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
  • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
  • Mentoring less experienced team members
What You’ll Bring
  • 3+ years DevOps / platform engineering experience
  • Must be based in Europe or Asia or LatAM
  • Experience independently leading the planning and deployment of a project
  • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
  • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
  • Experience with Infrastructure-as-Code, particularly Terraform
  • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
  • Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
  • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
  • Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
  • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
  • Experience defining SLIs and SLOs and using them to inform reliability work
  • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
  • Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
  • Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
  • A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
  • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
Our stack
  • Languages: Java / Spring Boot, Node.js, Python, JavaScript
  • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
  • Cache: ElastiCache, Redis, Valkey
  • Messaging: Apache RocketMQ, AutoMQ, Kafka
  • Networking & Proxy: Nginx, Kong, Cilium, eBPF
  • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
  • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
  • CI/CD: Jenkins, GitHub Actions
  • Metrics: Prometheus, Mimir, Grafana, Alertmanager
  • Logs: Loki, Vector
  • Traces: Tempo, OpenTelemetry, Alloy
  • Profiling: Pyroscope
  • RUM: Grafana Faro, OpenTelemetry SDK
  • Infrastructure as Code: Terraform
  • CDN & Edge: Cloudflare, AWS CloudFront
  • AWS CloudWatch
What’s In It For You
  • Sporty is a remote first company in pursuit of sustainability
  • A competitive salary + individual performance based bonuses every quarter
  • 28 days paid annual leave
  • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
  • Referral bonuses & flash bonuses
  • Top of the line equipment
  • Annual company retreats to provide great internal networking opportunities
Interview process
  • Remote video screening with our Talent Acquisition Team
  • Online assessment via Hackerrank
  • Remote video interview with 3 x Team Members (45 mins each, not separate days)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE / DevOps Engineer
SRE / DevOps Engineer

HeadHR • Town of Poland (NY)

On-site
USD 120,000 - 150,000
Site Reliability / Infrastructure Engineer
Site Reliability / Infrastructure Engineer

General Intuition & Medal • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and meaningful equity
Comprehensive medical, dental, and vision coverage
401(k)
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Juniper Square • United States

Remote
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+4
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Site Reliability Engineer (Mid-Senior)
Site Reliability Engineer (Mid-Senior)

Oxylabs • United States

On-site
USD 46,000 - 95,000
Private health insurance
Psychotherapy
On-site well-being consultants
+13
Senior DevOps Engineer, Infrastructure & Reliability
Senior DevOps Engineer, Infrastructure & Reliability

Worth AI, Inc. • Orlando (FL), Tampa (FL), Miami (FL), Atlanta (GA)

Hybrid
USD 140,000 - 170,000
Health insurance
401k
Life Insurance
+7
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Senior DevOps Engineer, Infrastructure & Reliability
Senior DevOps Engineer, Infrastructure & Reliability

Worth AI • Atlanta (GA)

Hybrid
USD 140,000 - 190,000
Health insurance
401k
Life Insurance
+7
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Ironclad • New York (NY), Chicago (IL), San Francisco (CA)

Hybrid
USD 220,000 - 235,000
Health coverage
Parental leave
Wellbeing stipends