Get more replies from employers
Send a job-specific resume in minutes.
Amelco Limited is seeking a hands-on Site Reliability Engineer to own reliability, observability, and cost efficiency for a high‑throughput betting platform. You’ll work in production operations with development teams to improve resilience and incident response.
The role is 60% infrastructure automation and 40% application‑focused reliability, covering Java Spring Boot services, Kubernetes patterns, and scalable CI/CD with Flux and Terraform.
Amelco Ltd are a leading gaming and gambling solution software provider with a strong presence in the USA, UK, and Europe. Through partnerships with global gaming companies, we build cutting-edge technical platforms across sportsbooks, lottery, casino, virtual gaming, and financial trading. Our vision is to shape the future of gaming by transforming operations into intelligent, data-driven solutions that deliver exceptional customer experiences and create sustainable value for all stakeholders. We believe in teamwork, knowledge sharing, and transparency with accountability.
We’re looking for a hands-on Site Reliability Engineer (SRE) to own the reliability, observability, and cost efficiency of our high-throughput betting platform. You’ll be embedded in our production operations, working directly with development teams to build resilient systems, implement actionable observability, and drive incident response from detection to remediation.
This role is 60% infrastructure automation and 40% application-focused reliability work — you’ll need to dive deep into both our Java Spring Boot services and Kubernetes deployment patterns to build lasting improvements.
Kubernetes: AWS EKS clusters and on-prem deployments
GitOps: FluxCD for declarative Kubernetes management across 20+ environments
Infrastructure as Code: Terraform for AWS resources and Kubernetes configurations
CI/CD: GitHub Actions pipelines with automated AI code review workflows
Monitoring: Prometheus metrics with custom Java Micrometer instrumentation
Logging: Loki for distributed log aggregation
Tracing: Tempo for distributed tracing
AWS: Cloudwatch
Core Platform: Java Spring Boot microservices for betting, trading, and customer management
Event Processing: High-throughput event queues with latency monitoring (Kakfa, JMS)
Data Pipeline: Avro-based S3 buffer systems with fault-tolerant write-ahead logs (StatefulS3Buffer))
DB: Postgres DB on RDS/Aurora. Key Responsibilities
Reliability Engineering (40%):
Partner with development teams to define and manage SLOs/SLIs specific to betting platform services (latency, queue depth, bet processing success rates) Enhance observability of our Java Spring Boot services — ensure metrics, logs, and tracing are actionable for detecting and fixing production issues Implement chaos engineering experiments targeting our event queue systems and high-availability betting services Design and execute pre-deployment readiness checks and post-release validation for risk‑critical services
Infrastructure & Platform (40%):
Own Kubernetes cluster reliability across development, QA, and production environments Automate operational processes using Python and Bash scripting within our GitHub Actions ecosystem Optimize infrastructure costs through rightsizing EKS workloads, tuning autoscaling policies, and implementing efficient resource utilization patterns Strengthen platform guardrails through Flux GitOps policies and Terraform module validation
Incident & Operations (20%):
Contribute to major incident response for betting platform outages, providing engineering expertise on Java service behavior and infrastructure dependencies Design and implement automated remediation patterns for common failure modes (event queue backpressure, connection pool exhaustion, database connectivity) Build hourly observability agent rules and thresholds for early‑warning detection of production anomalies Develop runbooks and automation for common operational tasks across our multi‑environment platform
3+ years in SRE, Platform Engineering, or DevOps roles with hands‑on production experience Strong Kubernetes operational expertise (AWS EKS, GitOps with Flux, Helm packaging) Proven track record building observability for Java Spring Boot applications using Prometheus, Grafana, and Loki Infrastructure as Code proficiency with Terraform and GitOps workflows Python and Bash scripting skills for automation and tooling development Experience designing and operating monitoring for high-throughput event processing systems Demonstrated ability to balance infrastructure cost efficiency with betting platform reliability requirements Excellent communication skills and ability to work across development, platform, and incident management teams
Experience with betting/gaming platform architecture and reliability patterns Familiarity with Avro-based data pipelines and S3 storage optimization AWS Certifications (Solutions Architect, DevOps Engineer, or AWS Certified Kubernetes Administrator) Background in chaos engineering for stateful event processing systems Knowledge of payment processing and risk calculation system reliability patterns Experience with automated incident detection and response workflows