Site Reliability Engineer

Segment (Twilio)

Göteborgs kommun

On-site

SEK 900,000 - 1,200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Recorded Future in Gothenburg, Sweden, is seeking a Site Reliability Engineer to ensure reliability, scalability, and performance of our critical systems. You will work with development teams to build and maintain robust infrastructure, implement automation, and promote operational excellence in cloud environments.

The role emphasizes observability, IaC, and on-call responsibility, with collaboration across teams to design for high availability and resilience.

Qualifications

  • 3+ years in SRE/DevOps or similar role.
  • Strong AWS knowledge, including networking.
  • Proficient in Linux and automation with IaC.
  • Experience with observability, metrics, and incident response.

Responsibilities

  • Ensure performance, capacity, scalability, reliability, and SRE SLAs across the platform.
  • Design and maintain scalable infrastructure on AWS.
  • Develop observability solutions with Grafana, ELK, and Prometheus.
  • Automate provisioning with Terraform and Chef.
  • Participate in a 24/7 on-call rotation and post-incident reviews.
  • Collaborate with engineering to build high-availability applications.
  • Identify bottlenecks and drive automation improvements.

Skills

AWS expertise
Linux
Observability tooling
Terraform
Chef
Kubernetes
Incident response
Team collaboration

Tools

Grafana
ELK Stack
Prometheus

Job description

With 1,000+ intelligence professionals serving over 1,900 clients worldwide, Recorded Future is the world’s most advanced, and largest, intelligence company!

We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) to join our growing team. In this role, you will be instrumental in ensuring the reliability, scalability, and performance of our critical systems. You will work closely with development teams to build and maintain robust infrastructure, implement automation, and foster a culture of operational excellence. This position requires a strong understanding of cloud environments, observability, and infrastructure as code principles.

What You’ll Do:
  • Ensure the performance, capacity, scalability, reliability, resiliency, security, compliance, support, cost efficiency, SLA, SLOs, RPOs and RTOs for the platform, either directly or in collaboration with other teams.
  • Make systemic improvements both proactively and for recurring issues.
  • Perform comprehensive Root Cause Analysis for outages.
  • Design, implement, and maintain scalable and reliable infrastructure on AWS.
  • Develop and manage observability solutions using tools such as Grafana, ELK (Elasticsearch, Logstash, Kibana), and Prometheus to monitor system health and performance.
  • Automate infrastructure provisioning and configuration using Terraform and Chef.
  • Participate in a 24/7 on-call rotation to respond to and resolve production incidents.
  • Collaborate with engineering teams to ensure applications are designed for high availability and resilience.
  • Proactively identify and address performance bottlenecks and potential issues.
  • Drive continuous improvement through automation, process optimization, and post-incident reviews.
What You’ll Bring:
  • 3+ years of experience in a Site Reliability Engineer, DevOps Engineer, or similar role.
  • Extensive hands‑on experience with Amazon Web Services (AWS), including a deep understanding of networking concepts within AWS.
  • Expert‑level troubleshooting and diagnostic skills
  • Proven track record of reducing system downtime
  • Ability to grasp complex architectures.
  • Advanced Linux skills (engineering fundamentals, networking, storage, operating systems)
  • Exposure managing and optimizing observability suites (e.g., Grafana, ELK Stack).
  • Strong proficiency in Terraform and Chef.
  • A strong preference for automating tasks and implementing solutions via Infrastructure as Code rather than manual changes.
  • Skilled in creating clear, concise incident reports and technical documentation
  • Ability to stay calm under pressure during an outage.Fantastic collaboration skills.
  • Spectacular collaborator and communicator.
  • A team player but self‑motivated.
Preferred Qualifications:
  • Knowledge and experience with Kubernetes.
  • Familiarity with message brokers such as RabbitMQ and Apache Kafka.
  • Experience with NoSQL databases, particularly MongoDB and Elasticsearch.
  • Familiarity with OpenTelemetry
  • Experience with large distributed systems and microservices architecture
  • Experience with CI/CD pipelines.

We are committed to maintaining an environment that attracts and retains talent from a diverse range of experiences, backgrounds and lifestyles. By ensuring all feel included and respected for being unique and bringing their whole selves to work, Recorded Future is made a better place every day.

If you need any accommodation or special assistance to navigate our website or to complete your application, please send an e‑mail with your request to our recruiting team at careers@recordedfuture.com.

Recorded Future is an equal opportunity and affirmative action employer and we encourage candidates from all backgrounds to apply. Recorded Future does not discriminate based on race, religion, color, national origin, gender including pregnancy, sexual orientation, gender identity, age, marital status, veteran status, disability or any other characteristic protected by law.

Recorded Future will not discharge, discipline or in any other manner discriminate against any employee or applicant for employment because such employee or applicant has inquired about, discussed, or disclosed the compensation of the employee or applicant or another employee or applicant.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Recordedfuture • Göteborgs kommun

Hybrid
SEK 700,000 - 1,000,000
Site Reliability Engineer
Site Reliability Engineer

Recorded Future • Göteborgs kommun

On-site
SEK 900,000 - 1,300,000
Site Reliability Engineer: AWS, Observability & Automation
Site Reliability Engineer: AWS, Observability & Automation

Recorded Future • Göteborgs kommun

On-site
SEK 900,000 - 1,200,000
Software Engineer
Software Engineer

Recordedfuture • Göteborgs kommun

On-site
SEK 386,000 - 580,000
AWS SRE: Observability, Automation & Resilient Infra
AWS SRE: Observability, Automation & Resilient Infra

Recorded Future • Göteborgs kommun

On-site
SEK 900,000 - 1,300,000
Software Engineer
Software Engineer

Recorded Future • Göteborgs kommun

On-site
SEK 620,000 - 880,000
Senior SRE - Cloud Reliability & Observability Lead
Senior SRE - Cloud Reliability & Observability Lead

Recordedfuture • Göteborgs kommun

Hybrid
SEK 700,000 - 1,000,000
Software Engineer - Site Reliability Engineering
Software Engineer - Site Reliability Engineering

Neo4j Inc • Malmö kommun

On-site
SEK 486,000 - 704,000
Site Reliability Engineer
Site Reliability Engineer

Kindred People AB • Stockholms kommun

On-site
SEK 650,000 - 950,000
Platform Engineer: Encryption & Data Security
Platform Engineer: Encryption & Data Security

Recorded Future • Göteborgs kommun

On-site