Site Reliability Engineer

Amelco Limited

Greater London

On-site

GBP 85,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Pension Scheme: Amelco matches up to 7
Staff Benefits Portal after probation
Yearly discretionary bonus scheme
Travel to Poland and Hungary

Job summary

Amelco Limited is seeking a hands-on Site Reliability Engineer to own reliability, observability, and cost efficiency for a high‑throughput betting platform. You’ll work in production operations with development teams to improve resilience and incident response.

The role is 60% infrastructure automation and 40% application‑focused reliability, covering Java Spring Boot services, Kubernetes patterns, and scalable CI/CD with Flux and Terraform.

Qualifications

  • 3+ years in SRE/Platform Engineering/DevOps with hands-on production experience.
  • Strong Kubernetes operational expertise (AWS EKS, GitOps with Flux, Helm packaging).
  • Proven track record building observability for Java Spring Boot apps using Prometheus, Grafana, Loki.

Responsibilities

  • Reliability Engineering (40%): define and manage SLOs/SLIs for betting platform services; enhance observability; design and run chaos experiments.
  • Infrastructure & Platform (40%): own Kubernetes reliability; automate ops with Python/Bash in GitHub Actions; optimize costs and implement Flux/Terraform controls.
  • Incident & Operations (20%): contribute to major incident response; develop automated remediation; build runbooks and cross-env automation.

Skills

SRE
Kubernetes
Observability
Terraform
GitOps Flux
Python Bash
Automation
Communication
Java Spring Boot

Tools

GitHub Actions
Prometheus
Grafana
Loki
FluxCD
AWS EKS

Job description

About Us

Amelco Ltd are a leading gaming and gambling solution software provider with a strong presence in the USA, UK, and Europe. Through partnerships with global gaming companies, we build cutting-edge technical platforms across sportsbooks, lottery, casino, virtual gaming, and financial trading. Our vision is to shape the future of gaming by transforming operations into intelligent, data-driven solutions that deliver exceptional customer experiences and create sustainable value for all stakeholders. We believe in teamwork, knowledge sharing, and transparency with accountability.

The Role

We’re looking for a hands-on Site Reliability Engineer (SRE) to own the reliability, observability, and cost efficiency of our high-throughput betting platform. You’ll be embedded in our production operations, working directly with development teams to build resilient systems, implement actionable observability, and drive incident response from detection to remediation.

This role is 60% infrastructure automation and 40% application-focused reliability work — you’ll need to dive deep into both our Java Spring Boot services and Kubernetes deployment patterns to build lasting improvements.

What You’ll Work With
Infrastructure & Platform

Kubernetes: AWS EKS clusters and on-prem deployments
GitOps: FluxCD for declarative Kubernetes management across 20+ environments
Infrastructure as Code: Terraform for AWS resources and Kubernetes configurations
CI/CD: GitHub Actions pipelines with automated AI code review workflows

Observability Stack

Monitoring: Prometheus metrics with custom Java Micrometer instrumentation
Logging: Loki for distributed log aggregation
Tracing: Tempo for distributed tracing
AWS: Cloudwatch

Application Environment

Core Platform: Java Spring Boot microservices for betting, trading, and customer management
Event Processing: High-throughput event queues with latency monitoring (Kakfa, JMS)
Data Pipeline: Avro-based S3 buffer systems with fault-tolerant write-ahead logs (StatefulS3Buffer))
DB: Postgres DB on RDS/Aurora. Key Responsibilities

Key Responsibilities

Reliability Engineering (40%):
Partner with development teams to define and manage SLOs/SLIs specific to betting platform services (latency, queue depth, bet processing success rates) Enhance observability of our Java Spring Boot services — ensure metrics, logs, and tracing are actionable for detecting and fixing production issues Implement chaos engineering experiments targeting our event queue systems and high-availability betting services Design and execute pre-deployment readiness checks and post-release validation for risk‑critical services

Infrastructure & Platform (40%):
Own Kubernetes cluster reliability across development, QA, and production environments Automate operational processes using Python and Bash scripting within our GitHub Actions ecosystem Optimize infrastructure costs through rightsizing EKS workloads, tuning autoscaling policies, and implementing efficient resource utilization patterns Strengthen platform guardrails through Flux GitOps policies and Terraform module validation

Incident & Operations (20%):
Contribute to major incident response for betting platform outages, providing engineering expertise on Java service behavior and infrastructure dependencies Design and implement automated remediation patterns for common failure modes (event queue backpressure, connection pool exhaustion, database connectivity) Build hourly observability agent rules and thresholds for early‑warning detection of production anomalies Develop runbooks and automation for common operational tasks across our multi‑environment platform

Required Skills & Experience

3+ years in SRE, Platform Engineering, or DevOps roles with hands‑on production experience Strong Kubernetes operational expertise (AWS EKS, GitOps with Flux, Helm packaging) Proven track record building observability for Java Spring Boot applications using Prometheus, Grafana, and Loki Infrastructure as Code proficiency with Terraform and GitOps workflows Python and Bash scripting skills for automation and tooling development Experience designing and operating monitoring for high-throughput event processing systems Demonstrated ability to balance infrastructure cost efficiency with betting platform reliability requirements Excellent communication skills and ability to work across development, platform, and incident management teams

Nice-to-Have Skills

Experience with betting/gaming platform architecture and reliability patterns Familiarity with Avro-based data pipelines and S3 storage optimization AWS Certifications (Solutions Architect, DevOps Engineer, or AWS Certified Kubernetes Administrator) Background in chaos engineering for stateful event processing systems Knowledge of payment processing and risk calculation system reliability patterns Experience with automated incident detection and response workflows

  • Pension Scheme: Amelco matches up to 7% contributions of your base salary for all staff. You will automatically be entered at 4%.
  • Staff Benefits Scheme: access to the staff benefits portal after successful completion of a 4-month probation period.
  • Yearly discretionary bonus scheme and pay reviews
  • Opportunity to travel and visit our office in Poland and Hungary
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

Amelco Limited • Greater London

On-site
GBP 75,000 - 110,000
Pension Scheme - company matched up to
Staff Benefits Portal access
Yearly discretionary bonus
+1
SRE: High-Throughput Betting Platform Reliability & Observability
SRE: High-Throughput Betting Platform Reliability & Observability

Amelco Limited • Greater London

On-site
GBP 85,000 - 110,000
Pension Scheme: Amelco matches up to 7
Staff Benefits Portal after probation
Yearly discretionary bonus scheme
+1
Database Administrator
Database Administrator

Amelco Limited • Greater London

Hybrid
GBP 70,000 - 100,000
Pension scheme
Staff benefits portal
Yearly discretionary bonus
+1
Site Reliability Engineer
Site Reliability Engineer

bet365 Group • Manchester

On-site
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Software Engineer, Site Reliability Engineering
Software Engineer, Site Reliability Engineering

bet365 Group • Stoke-on-Trent

Hybrid
GBP 60,000 - 80,000
Eye care
Flu Vaccinations
Life Assurance
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Manchester

Hybrid
GBP 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Stoke-on-Trent

Hybrid
GBP 60,000 - 80,000
Senior Software Engineer New London, England, United Kingdom
Senior Software Engineer New London, England, United Kingdom

Genius Sports Group • Greater London

On-site
GBP 60,000 - 90,000
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

Spectrum IT Recruitment • City Of London

Hybrid
GBP 85,000 - 110,000
Life Insurance - 4 x Annual Salary
Private Medical Insurance
Bonus Scheme
+3
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

SPECTRUM IT • Greater London

Hybrid
GBP 70,000 - 100,000
Life Insurance
Private Medical Insurance
Bonus Scheme
+3