Site Reliability Engineer

Sporty Group

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote-first company
Quarterly performance bonuses
28 days paid annual leave
Core hours 10am–3pm in local time zone
Referral bonuses
Top-of-the-line equipment
Annual company retreats

Job summary

Sporty Group is seeking an experienced DevOps/Platform Engineer to join a remote-first team. You will drive Kubernetes/EKS operations, implement GitOps workflows (ArgoCD, Helm), and manage cloud infrastructure with Terraform on AWS.

You will own on-call rotations, incident response, and post-incident follow-ups, while mentoring junior engineers and guiding rapid deployments to new countries. The role emphasizes reliability, observability, and cost optimization across multi-country deployments,

Qualifications

  • 3+ years DevOps / SRE / platform engineering experience.
  • Must be based in Europe.
  • Experience independently leading the planning and deployment of a project.
  • Strong Kubernetes knowledge including AWS EKS and GitOps with ArgoCD and Helm.
  • IaC with Terraform.
  • Scripting with Bash, Python, or Golang; Rust a plus.
  • Observability stacks: Prometheus, Loki, Tempo, Pyroscope, OpenTelemetry.
  • RUM: Grafana Faro or OpenTelemetry SDK instrumentation.
  • On-call and incident response experience with post-mortems.
  • Design alert frameworks to minimize noise and alert fatigue.
  • Define SLIs and SLOs for reliability.
  • Familiarity with service mesh concepts (Cilium).
  • Networking knowledge: TCP/IP, HTTP.
  • High availability for high-traffic environments.
  • Caching: CDN, Redis, Memcached.
  • Troubleshooting Linux and JVM optimization.

Responsibilities

  • Work with a team of DevOps and DBA professionals.
  • Improve infrastructure and processes across deployed countries and prepare for future deployments.
  • Improve Kubernetes platform stability and efficiency, optimize resources, reduce costs, and GitOps-first provisioning.
  • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards.
  • Own weekend on-call operations, triage production incidents, perform root cause analysis, and drive post-incident reviews.
  • Design and manage alert pipelines to ensure actionable signals and prevent alert fatigue.
  • Define and maintain SLIs and SLOs for critical services to drive reliability improvements.
  • Take ownership of cloud operation activities.
  • Liaise with external security agencies for annual audits and perform internal security sweeps.
  • Aid in reconfiguring architecture for rapid deployments to new countries.
  • Mentor less experienced team members.

Skills

DevOps / SRE / platform engineering
Kubernetes
EKS
GitOps
ArgoCD
Helm
Terraform
Bash
Python
Golang
Rust
Prometheus
Loki
Tempo
Pyroscope
OpenTelemetry
Grafana Faro
Redis
Memcached
Nginx
Kong
Cilium
HTTP
Networking
Incident response
On-call
Troubleshooting Linux

Tools

AWS
Docker
Kubernetes (EKS)
Grafana
Prometheus
OpenTelemetry
Jenkins
GitHub Actions

Job description

What you’ll be doing
  • Work with a team of DevOps and DBA professionals
  • Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
  • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
  • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
  • Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
  • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
  • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
  • Take ownership and responsibility for our cloud operation activities
  • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
  • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
  • Mentoring less experienced team members
What you’ll bring
  • 3+ years DevOps / SRE / platform engineering experience
  • Must be based in Europe
  • Experience independently leading the planning and deployment of a project
  • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
  • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
  • Experience with Infrastructure-as-Code, particularly Terraform
  • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
  • Hands‑on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
  • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
  • Proven on‑call and incident response experience, comfortable triaging production issues under pressure, leading post‑mortems, and driving follow‑up actions
  • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
  • Experience defining SLIs and SLOs and using them to inform reliability work
  • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium‑based service mesh in non‑production environments
  • Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
  • Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
  • A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
  • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
Our stack
  • Languages: Java / Spring Boot, Node.js, Python, JavaScript
  • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
  • Cache: ElastiCache, Redis, Valkey
  • Messaging: Apache RocketMQ, AutoMQ, Kafka
  • Networking & Proxy: Nginx, Kong, Cilium, eBPF
  • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
  • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
  • CI/CD: Jenkins, GitHub Actions
  • Metrics: Prometheus, Mimir, Grafana, Alertmanager
  • Logs: Loki, Vector
  • Traces: Tempo, OpenTelemetry, Alloy
  • Profiling: Pyroscope
  • RUM: Grafana Faro, OpenTelemetry SDK
  • Infrastructure as Code: Terraform
  • CDN & Edge: Cloudflare, AWS CloudFront
  • AWS CloudWatch
What’s in it for you
  • Sporty is a remote first company in pursuit of sustainability
  • A competitive salary + individual performance based bonuses every quarter
  • 28 days paid annual leave
  • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
  • Referral bonuses & flash bonuses
  • Top of the line equipment
  • Annual company retreats to provide great internal networking opportunities
Interview process
  • Remote video screening with our Talent Acquisition Team
  • Online assessment via Hackerrank
  • Remote video interview with 3 x Team Members (45 mins each, not separate days)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

sportygroup • United States

Remote
USD 150,000 - 210,000
Remote-first company
Competitive salary
Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

sportygroup • United States

Remote
USD 140,000 - 170,000
Remote-first company
Competitive salary with quarterlyBonu
28 days paid annual leave
+4
Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

Sporty Group • United States

On-site
USD 120,000 - 180,000
Remote first
Bonuses (quarterly)
28 days leave
+4
Senior DevOps Engineer
Senior DevOps Engineer

Linuxconfig • Northern (KY)

On-site
USD 120,000 - 160,000
Fully remote work
Paid vacation
Private medical insurance
+1
Senior QA Engineer
Senior QA Engineer

sportygroup • United States

Remote
USD 90,000 - 130,000
Remote-First
Bonuses/Quarterly rewards
28 days paid annual leave
+4
Senior Devops Engineer
Senior Devops Engineer

Physitrack Limited • United States

Remote
USD 77,000 - 103,000
Senior DevOps Engineer, Infrastructure & Reliability
Senior DevOps Engineer, Infrastructure & Reliability

Worth AI, Inc. • Orlando (FL), Northern (KY)

Hybrid
USD 140,000 - 190,000
Health Care Plan (Medical, Dental &amp
Retirement Plan (401k)
Life Insurance
+7
Devops Engineer
Devops Engineer

PhysicsX • New York (NY)

Hybrid
USD 160,000 - 230,000
Equity options
401(k) contribution
Free team lunch 1x/week
+6
Senior Software Engineer | Kubernetes Automation | EU
Senior Software Engineer | Kubernetes Automation | EU

Cast AI • Union (NJ)

Remote
USD 81,000 - 113,000
Equity options
Remote-first environment
Learning budget and hackathons
+1
DevOps Engineer (Remote)
DevOps Engineer (Remote)

C Teleport • United States

Remote
USD 120,000 - 180,000
Compensated on-call rotation