Senior Site Reliability Engineer

Flowcode

York and North Yorkshire

On-site

GBP 90,000 - 120,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Unlimited Vacation
Health benefits
Stock options
Custom swag
Lunch stipend & office snacks
Happy hours & team dinner
Team offsites
Parental Leave
Diverse Managers

Job summary

Flowcode is seeking a Senior Site Reliability Engineer to enhance reliability, scalability and observability across our platforms. You will own EKS-based infrastructure,end-to-end responsibilities include incident response, postmortems, and driving IaC via Terraform.

Join a team building scalable cloud infrastructure, expanding GitOps with ArgoCD and GitHub Actions, and partnering with engineering to deliver robust, highly available systems in a fast-growing environment in the United Kingdom.

Qualifications

  • 4+ years of professional experience across SRE, DevOps, or Platform Engineering
  • Hands-on experience with GitOps workflows via ArgoCD and Helm deployments
  • Mastery of core AWS services (EKS, Networking/VPC, RDS, IAM)
  • Experience maintaining and scaling CI/CD automation using GitHub Actions
  • Background in supporting large-scale distributed systems in high-availability production environments
  • Ability to author production-grade code in Go or Python alongside robust shell scripting

Responsibilities

  • Own key pieces of our EKS-based infrastructure end-to-end
  • Improve system availability, scalability, and resilience across Flowcode’s platforms
  • Develop and operate scalable cloud infrastructure for Flowcode
  • Design and scale deployment pipelines using GitHub Actions
  • Expand GitOps practices and tooling through ArgoCD
  • Collaborate with product engineering to optimize internal developer experience
  • Manage and scale our core AWS footprint (EKS, VPC, RDS) via Terraform
  • Develop high-signal metrics, tracing, and dashboards with observability tooling
  • Contribute to incident response and postmortems; implement durable fixes

Skills

SRE/DevOps experience
Go or Python coding
Shell scripting
Cloud architecture
Incident response
System reliability
Collaboration with engineering teams
DevOps best practices

Tools

GitHub Actions
ArgoCD
Helm
Terraform
OpenTofu
Kubernetes/EKS
AWS (EKS, VPC, RDS, IAM)
Karpenter / Cluster Autoscaler

Job description

  • Flowcode is seeking a Senior Site Reliability Engineer (SRE) to work on reliability and infrastructure efforts across our platforms. This role will help grow and drive our infrastructure strategy, operational rigor and observability while building and supporting the systems and tooling required to support Flowcode’s continued growth
  • As an individual contributor within our engineering organization, you will develop and operate scalable cloud infrastructure, establish best practices around deployment and reliability, and partner closely with engineering teams to ensure systems are scalable, resilient and observable
  • Reliability & Infrastructure
  • Improve system availability, scalability, and resilience across Flowcode’s platforms
  • Own key pieces of our EKS-based infrastructure end-to-end
  • Contribute to incident response and postmortems, turning findings into durable fixes
  • Support engineering teams with infrastructure questions, escalations, and day-to-day unblocking
  • Cloud & Platform Engineering
  • Manage and scale our core AWS footprint (EKS, VPC, RDS) through Infrastructure as Code (Terraform)
  • Enhance disaster recovery and failover mechanisms to protect mission-critical workloads
  • Collaborate with product engineering to streamline and optimize internal developer experience
  • CI/CD & Deployment Automation
  • Design and scale deployment pipelines using GitHub Actions
  • Expand GitOps practices and tooling through ArgoCD
  • Facilitate secure delivery with automated validation and progressive rollout strategies
  • Observability & Monitoring
  • Oversee and optimize the organization’s monitoring, logging, and alerting infrastructure
  • Develop high-signal metrics, tracing, and visualization dashboards while minimizing operational noise
  • Establish and monitor Service Level Objectives for managed platform components
Benefits
  • Unlimited Vacation
  • Paid Health Benefits
  • Employee Stock Options
  • Custom swag
  • Lunch stipend & office snacks
  • Happy hours & team dinner
  • Team offsites
  • Parental Leave
  • Diverse Managers
  • Experience maintaining and scaling CI/CD automation using GitHub Actions within collaborative environments
  • Advanced Terraform or OpenTofu expertise, encompassing module architecture and production state management
  • 4+ years of professional experience across SRE, DevOps, or Platform Engineering domains
  • Hands-on operational experience with GitOps workflows via ArgoCD and Helm-based deployments
  • Mastery of core AWS services, specifically EKS, Networking/VPC, RDS, and IAM
  • Proven track record of leading infrastructure initiatives from initial design through to long-term operation
  • Background in supporting large-scale distributed systems within high-availability production environments
  • Ability to author production-grade code in Go or Python alongside robust shell scripting
  • Adept at navigating interrupt-driven workflows, balancing strategic project delivery with day-to-day operational support
  • Technical proficiency in Kubernetes, including cluster troubleshooting and managing controllers or CRDs
  • Exposure to Crossplane or alternative Kubernetes-native solutions for infrastructure provisioning
  • Deep observability experience utilizing Datadog or Prometheus to engineer SLOs, high-signal dashboards, and intelligent alerting
  • Practical knowledge of modern secrets management frameworks and implementation
  • Experience optimizing cluster efficiency through autoscaling technologies such as Karpenter or Cluster Autoscaler
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability & Observability Expert
Senior SRE: Cloud Reliability & Observability Expert

Flowcode • York and North Yorkshire

On-site
GBP 90,000 - 120,000
Unlimited Vacation
Health benefits
Stock options
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

P2P • Greater London

On-site
GBP 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brevan Howard • Greater London

On-site
GBP 90,000 - 130,000
Director of DevOps & SRE
Director of DevOps & SRE

InvestorFlow • Greater London

Hybrid
GBP 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Senior Software Engineer (Infrastructure Engineering)
Senior Software Engineer (Infrastructure Engineering)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 130,000
Senior Solutions Engineer
Senior Solutions Engineer

Kroll • United Kingdom

On-site
GBP 80,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Platform Engineer
Senior Platform Engineer

Jobtailor • Greater London

Hybrid
GBP 90,000 - 120,000
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

Spectrum IT Recruitment • City Of London

Hybrid
GBP 65,000 - 120,000
Life Insurance 4x Annual Salary
Private Medical Insurance
Bonus Scheme
+3