Senior Site Reliability Engineer I

RX Brasil

Philadelphia (Philadelphia County)

On-site

USD 95,000 - 159,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Annual incentive bonus

Job summary

Elsevier in the United States is seeking a Senior Site Reliability Engineer to lead reliability, scalability, and performance across critical platforms. You will drive automation, improve observability, and partner with engineering to design resilient systems.

The role emphasizes incident response, post-mortems, and production readiness, with a focus on AI tooling deployment and secure self-service for multiple teams.

Qualifications

  • Advanced Terraform: modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state.
  • AWS Operations: production multi-account/multi-region environments across ECS, RDS, ALB, VPC, IAM, Route53, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, CloudWatch.
  • GitHub Actions CI/CD: reusable workflows, OIDC, approval gates, runners, Terraform deployments, app deployments, migrations.
  • ECS Fargate & Containers: Docker, ECR, ECS task definitions/services, IAM roles, health checks, autoscaling, ALB, deployment rollbacks.
  • AWS Networking & Security: VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, cloud security best practices.
  • Incident Response & Observability: troubleshooting with logs, metrics, alarms, RCAs, rollback decisions, runbooks.
  • Linux & Automation: Linux and Git basics with Bash/Python scripting for AWS CLI automation, CI/CD, tooling.
  • AI Tooling Deployment: deploying AI services in production with monitoring, reliability, security for AI-powered features.
  • Developer Enablement: support multiple teams, troubleshoot infra/app layers, document solutions, enable self-service.

Responsibilities

  • Creating monitoring queries and establishing service level baselines.
  • Support senior engineers during incidents.
  • Contribute to post-mortems and RCAs.
  • Participate in disaster recovery tests.
  • Implement automation and execute code in production.
  • Contribute to SRE knowledge documentation.
  • Support deployment, monitoring, and reliability of AI-integrated services.
  • Assist architecture and engineers in infrastructure topology drawings and deployment workflows.
  • Test availability, reliability, and recoverability in non-production environments.

Skills

Terraform
AWS
GitHub Actions
Docker
ECS
ECR
CI/CD
Networking
Observability
Incident Response
Linux
Automation
Security
AI Tooling
Developer Enablement

Tools

Terraform
AWS
GitHub Actions
Docker
ECS
ECR
Route53
VPC
CloudWatch
IAM
KMS
Secrets Manager

Job description

Senior Site Reliability Engineer

About the team

Embedded Innovation Teams are cross-functional squads embedded within our segments to rapidly turn internal AI experimentation into validated, reusable solutions, building the capabilities we need to deliver customer value and growth. We work problem-first rather than tool-first, directly inside segment and function teams, improving the internal workflows that help our people deliver better outcomes for customers, faster.

About the role

As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve service availability, streamline operations, and enhance system recovery capabilities. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host-function stakeholders to make sure what you build fits real workflows, not assumed ones. You will also support handover and capability-building so the solution is owned and operable after the squad moves on.

Key Responsibilities
  • Creating monitoring queries and establishes service level baselines.
  • Supporting senior engineers during incidents.
  • Making contributions during post-mortems and RCAs.
  • Participating in disaster recovery tests.
  • Implementing automation and executes code in production environments.
  • Contributing to SRE knowledge documentation.
  • Supporting the deployment, monitoring, and reliability of services integrating AI tools.
  • Supporting architecture and senior engineers in the creation of infrastructure topology drawings and deployment workflows.
  • Carrying out the testing of availability, reliability, and recoverability in non-production environments.
Requirements
  • Advanced Terraform: Expertise in modules, providers, state management, lifecycle controls, drift detection, safe refactoring, and remote state (S3, locking, cross-stack dependencies).
  • AWS Operations: Hands-on experience managing production, multi-account, multi-region AWS environments across ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
  • GitHub Actions CI/CD: Experience building and troubleshooting reusable workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
  • ECS Fargate & Containers: Knowledge of Docker, ECR, ECS task definitions/services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
  • AWS Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices.
  • Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.
  • Linux & Automation: Strong Linux and Git fundamentals with Bash/Python scripting for AWS CLI automation, CI/CD, and operational tooling.
  • AI Tooling Deployment: Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices for AI-powered features.
  • Developer Enablement: Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.
Benefits

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. This job is eligible for an annual incentive bonus.

Compensation

U.S. National Base Pay Range: $95,300 - $158,800. Geographic differentials may apply in some locations to better reflect local market rates.

Equal Opportunity

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law. USA Job Seekers: EEO Know Your Rights.

Company Overview

Elsevier is a renowned global information analytics company that primarily focuses on providing scientific, technical, and medical (STM) research content, tools, and services. It is one of the largest publishers of academic journals and scholarly literature in the world. Elsevier operates in various domains, including science, technology, medicine, social sciences, and more. They publish a vast number of peer-reviewed journals covering a wide range of disciplines. These journals act as platforms for researchers and academics to share their findings and contribute to the advancement of knowledge in their respective fields. In addition to publishing, Elsevier offers a suite of digital solutions and services to support researchers, scientists, and professionals in their work. They provide online platforms like ScienceDirect, Scopus, and Mendeley, which offer access to a vast repository of scholarly articles, research papers, and other scientific content. These platforms often serve as essential resources for software developers seeking to stay updated with the latest scientific advancements. In addition to publishing, Elsevier offers a suite of digital solutions and services to support researchers, scientists, and professionals in their work. They provide online platforms like ScienceDirect, Scopus, and Mendeley, which offer access to a vast repository of scholarly articles, research papers, and other scientific content. These platforms may serve as essential resources for software developers seeking to stay updated with the latest scientific advancements.

Our Mission & Vision

Elsevier is a global leader in advanced information and decision support for science and healthcare. We believe that by working together with the communities we serve, we can shape human progress to go further, happen faster, and benefit all. We support continuous discovery and uphold the highest standards of content integrity, reliability, and reproducibility so the communities we serve can advance their field of science, healthcare or innovation with confidence. By combining high-quality content with powerful analytics, we transform complexity into clarity and deliver mission-critical insights that help professionals make better decisions when it matters most. We deliver insights that help research institutions, governments, and funders achieve their goals. We help researchers discover and share knowledge, collaborate, and accelerate innovation. We help librarians provide verified, quality information to universities. We help innovators turn knowledge into new products. We help health professionals improve patient care and educators train the next generation of doctors and nurses. Connecting quality content and innovative technologies, we make progress go further and happen faster. And by championing inclusion and sustainability, we ensure progress benefits all. With 9,500 employees, over 2,300 technologists in 5 major tech hubs, and more than 60 locations across the globe, we are committed to supporting the scientific and healthcare communities around the world. We offer a diverse range of opportunities across technology, commercial, business, and early career jobs. If you are looking for a career that inspires progress in science, innovation and health, and allows you to grow every day, find your team at Elsevier. Elsevier is part of RELX Group. Let's shape progress together. Join us. elsevier.com/about/careers

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer I
Senior Site Reliability Engineer I

RXinsider LTD. • Philadelphia

On-site
USD 95,000 - 159,000
Annual incentive bonus
Country-specific benefits
Senior Data Scientist I
Senior Data Scientist I

Elsevier • Philadelphia

On-site
MXN 1,959,000 - 3,264,000
Senior Software Engineer I - Ruby on Rails with Sidekiq
Senior Software Engineer I - Ruby on Rails with Sidekiq

Elsevier • Philadelphia

On-site
USD 87,000 - 144,000
Sr Product Manager II, ClinicalKey Nursing
Sr Product Manager II, ClinicalKey Nursing

Elsevier • Northern (KY)

On-site
USD 115,000 - 192,000
Annual incentive bonus
Country-specific benefits
Senior Site Reliability Engineer I
Senior Site Reliability Engineer I

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Sales Development Representative
Sales Development Representative

Elsevier • New York (NY), Northern (KY)

On-site
USD 54,000 - 152,000
Health plan benefits
Employee Assistance Program
Retirement Benefits
+3
Nursing Education Program Implementation Specialist (Remote)
Nursing Education Program Implementation Specialist (Remote)

RX Brasil • Missouri

Remote
USD 79,000 - 131,000
Sales Engineer – ClinicalKey
Sales Engineer – ClinicalKey

Elsevier • California (MO)

On-site
USD 81,000 - 151,000
React Node Senior Software Engineer I
React Node Senior Software Engineer I

Elsevier Inc. Company • Philadelphia

On-site
USD 87,000 - 144,000
Principal Software Engineer/Principal AI Engineer
Principal Software Engineer/Principal AI Engineer

Elsevier • Philadelphia

On-site
USD 115,000 - 192,000
Generous vacation entitlement
Comprehensive Pension Plan
Family leave and sabbatical options
+3