Weekend Site Reliability Engineer

Sporty Group

United States

À distance

USD 120 000 - 180 000

Plein temps

14 jours+
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.

Passez les filtres ATS

Avantages offerts par ce poste

Remote first
Bonuses (quarterly)
28 days leave
Core hours 10am-3pm
Referral bonuses
Equipment provided
Annual retreats

Résumé du poste

Sporty Group is seeking a Weekend Site Reliability Engineer to join a DevOps-led team. The role covers Saturday, Sunday and Monday with flexible days off, focusing on stabilizing Kubernetes, cloud infra, and GitOps-driven deployments across multiple countries.

You will own weekend on-call, monitor dashboards (Grafana), implement alerting pipelines, define SLIs/SLOs, mentor colleagues, and coordinate audits with security teams.

Qualifications

  • 3+ years DevOps / platform engineering experience.
  • Based in Europe or Asia or LatAM
  • Experience leading the planning and deployment of a project
  • Strong AWS knowledge and ability to provision resources to meet demand
  • Kubernetes experience with EKS and GitOps tools (ArgoCD, Helm)
  • IaC experience with Terraform
  • Scripting in Bash/Python/Go; Rust a plus
  • Observability stacks: Prometheus, Grafana, Loki, Tempo, OpenTelemetry
  • RUM experience a plus; Grafana Faro or OpenTelemetry SDK instrumentation a plus
  • On-call and incident response experience; post-mortems and follow-ups

Responsabilités

  • Weekend SRE covering Saturday, Sunday and Monday with flexible days off.
  • Improve infrastructure and deployments across deployed countries.
  • Improve Kubernetes platform stability, cost efficiency, and provisioning via GitOps.
  • Monitor and maintain cloud infra with autoscaling, alerting pipelines, and Grafana dashboards.
  • Own weekend on-call, triage incidents, perform root cause analysis, and drive post-incident reviews.
  • Design alert pipelines to minimize noise and avoid waterfall alerting.
  • Define and maintain SLIs/SLOs to drive reliability improvements.
  • Mentor junior team members.

Connaissances

DevOps
Cloud platforms
Kubernetes
GitOps
Scripting

Outils

AWS
Terraform
ArgoCD
Helm

Description du poste

What You’ll Be Doing
  • Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
  • Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
  • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
  • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
  • Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
  • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
  • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
  • Take ownership and responsibility for our cloud operation activities
  • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
  • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
  • Mentoring less experienced team members
What You’ll Bring
  • 3+ years DevOps / platform engineering experience
  • Must be based in Europe or Asia or LatAM
  • Experience independently leading the planning and deployment of a project
  • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
  • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
  • Experience with Infrastructure-as-Code, particularly Terraform
  • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
  • Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
  • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
  • Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
  • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
  • Experience defining SLIs and SLOs and using them to inform reliability work
  • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
  • Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
  • Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
  • A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
  • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
Our stack
  • Languages: Java / Spring Boot, Node.js, Python, JavaScript
  • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
  • Cache: ElastiCache, Redis, Valkey
  • Messaging: Apache RocketMQ, AutoMQ, Kafka
  • Networking & Proxy: Nginx, Kong, Cilium, eBPF
  • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
  • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
  • CI/CD: Jenkins, GitHub Actions
  • Metrics: Prometheus, Mimir, Grafana, Alertmanager
  • Logs: Loki, Vector
  • Traces: Tempo, OpenTelemetry, Alloy
  • Profiling: Pyroscope
  • RUM: Grafana Faro, OpenTelemetry SDK
  • Infrastructure as Code: Terraform
  • CDN & Edge: Cloudflare, AWS CloudFront
  • AWS CloudWatch
What’s In It For You
  • Sporty is a remote first company in pursuit of sustainability
  • A competitive salary + individual performance based bonuses every quarter
  • 28 days paid annual leave
  • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
  • Referral bonuses & flash bonuses
  • Top of the line equipment
  • Annual company retreats to provide great internal networking opportunities
Interview process
  • Remote video screening with our Talent Acquisition Team
  • Online assessment via Hackerrank
  • Remote video interview with 3 x Team Members (45 mins each, not separate days)
Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

sportygroup • États-Unis

À distance
USD 140 000 - 170 000
Remote-first company
Competitive salary with quarterlyBonu
28 days paid annual leave
+4
Site Reliability Engineer
Site Reliability Engineer

sportygroup • États-Unis

À distance
USD 150 000 - 210 000
Remote-first company
Competitive salary
Site Reliability Engineer
Site Reliability Engineer

Sporty Group • États-Unis

À distance
USD 140 000 - 190 000
Remote-first company
Quarterly performance bonuses
28 days paid annual leave
+4
Database Reliability Engineer
Database Reliability Engineer

sportygroup • États-Unis

À distance
USD 130 000 - 180 000
Remote-first culture
Competitive salary + bonuses
28 days paid annual leave
+3
Senior Devops Engineer
Senior Devops Engineer

Physitrack Limited • États-Unis

À distance
USD 77 000 - 103 000
SRE Platform Engineer
SRE Platform Engineer

Buxton • États-Unis

À distance
USD 120 000 - 160 000
Remote-first work
Flexible working hours
Paid time off
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Runware • États-Unis

À distance
USD 140 000 - 190 000
Generous paid time off
Stock options
Remote-first setup
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Social Discovery Ventures • États-Unis

À distance
USD 120 000 - 180 000
REMOTE OPPORTUNITY
Vacation 28 days/year
Wellness days 7/year
+4
Senior DevOps Engineer
Senior DevOps Engineer

Linuxconfig • Northern (KY)

Sur place
USD 120 000 - 160 000
Fully remote work
Paid vacation
Private medical insurance
+1
Senior Software Engineer, Platform & Developer Experience
Senior Software Engineer, Platform & Developer Experience

Worth AI • Atlanta (GA)

Hybride
USD 140 000 - 180 000
Health Care Plan
Retirement Plan
Life Insurance
+5