AWS SRE: Observability, Automation & Resilient Infra

Recorded Future

Göteborgs kommun

On-site

SEK 900,000 - 1,300,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Recorded Future is seeking a Site Reliability Engineer to ensure reliability, scalability, and performance of critical systems. You will collaborate with development teams to build robust infrastructure and automate operations, with a focus on AWS-based design and observability.

The role requires 3+ years in SRE/DevOps, deep AWS networking knowledge, and strong Linux skills, plus hands-on experience with Grafana, ELK, Prometheus, Terraform, and Chef.

Qualifications

  • 3+ years of experience in a Site Reliability Engineer, DevOps Engineer, or similar role.
  • Extensive hands-on experience with Amazon Web Services (AWS), including a deep understanding of networking concepts within AWS.
  • Expert-level troubleshooting and diagnostic skills
  • Proven track record of reducing system downtime
  • Ability to grasp complex architectures.
  • Advanced Linux skills (engineering fundamentals, networking, storage, operating systems)
  • Exposure managing and optimizing observability suites (e.g., Grafana, ELK Stack).
  • Strong proficiency in Terraform and Chef.
  • A strong preference for automating tasks and implementing solutions via Infrastructure as Code rather than manual changes.
  • Skilled in creating clear, concise incident reports and technical documentation
  • Ability to stay calm under pressure during an outage.
  • Fantastic collaboration skills.
  • Spectacular collaborator and communicator.
  • A team player but self motivated.

Responsibilities

  • Ensure the performance, capacity, scalability, reliability, resiliency, security, compliance, support, cost efficiency, SLA, SLOs, RPOs and RTOs for the platform, either directly or in collaboration with other teams.
  • Make systemic improvements both proactively and for recurring issues.
  • Perform comprehensive Root Cause Analysis for outages.
  • Design, implement, and maintain scalable and reliable infrastructure on AWS.
  • Develop and manage observability solutions using Grafana, ELK (Elasticsearch, Logstash, Kibana), and Prometheus to monitor system health and performance.
  • Automate infrastructure provisioning and configuration using Terraform and Chef.
  • Participate in a 24/7 on-call rotation to respond to and resolve production incidents.
  • Collaborate with engineering teams to ensure applications are designed for high availability and resilience.
  • Proactively identify and address performance bottlenecks and potential issues.
  • Drive continuous improvement through automation, process optimization, and post-incident reviews.

Skills

AWS
Linux
Observability
Terraform
Chef
Grafana
ELK
Prometheus
Kubernetes
On-call
Incident response

Tools

Grafana
ELK Stack
Prometheus
Terraform
Chef

Job description

Recorded Future is seeking a Site Reliability Engineer to ensure reliability, scalability, and performance of critical systems. You will collaborate with development teams to build robust infrastructure and automate operations, with a focus on AWS-based design and observability.

The role requires 3+ years in SRE/DevOps, deep AWS networking knowledge, and strong Linux skills, plus hands-on experience with Grafana, ELK, Prometheus, Terraform, and Chef.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Recordedfuture • Göteborgs kommun

Hybrid
SEK 700,000 - 1,000,000
Senior SRE - Cloud Reliability & Observability Lead
Senior SRE - Cloud Reliability & Observability Lead

Recordedfuture • Göteborgs kommun

Hybrid
SEK 700,000 - 1,000,000
Site Reliability Engineer
Site Reliability Engineer

Segment (Twilio) • Göteborgs kommun

On-site
SEK 900,000 - 1,200,000
Senior SRE: Scale, Automate & Own Resilient Cloud
Senior SRE: Scale, Automate & Own Resilient Cloud

Recorded Future • Göteborgs kommun

On-site
SEK 700,000 - 1,100,000
Site Reliability Engineer - Observability & Automation
Site Reliability Engineer - Observability & Automation

Kindred People AB • Stockholms kommun

On-site
SEK 650,000 - 950,000
Senior SRE: Scale, Reliability & Observability
Senior SRE: Scale, Reliability & Observability

Fountain • Stockholms kommun

On-site
SEK 900,000 - 1,300,000
Flexible vacation policy
Paid holidays
Monthly lunch stipends
+4
Senior SRE: Scale, Resilience & Cloud Infra (Remote)
Senior SRE: Scale, Resilience & Cloud Infra (Remote)

SCALIS • Stockholms kommun

On-site
SEK 900,000 - 1,200,000
Health Insurance
Paid Time Off
Paid Holidays
+2
Platform SRE: Scalable Infra, CI/CD & Observability
Platform SRE: Scalable Infra, CI/CD & Observability

Linuxcareers • Stockholms kommun

On-site
SEK 600,000 - 800,000
SRE Engineer: Build Reliable, Scalable Systems
SRE Engineer: Build Reliable, Scalable Systems

Neo4j • Malmö kommun

Hybrid
SEK 900,000 - 1,300,000
Hybrid work model
Competitive compensation
Site Reliability Engineer
Site Reliability Engineer

Recorded Future • Göteborgs kommun

On-site
SEK 900,000 - 1,300,000