Senior Site Reliability Engineer I

RXinsider LTD.

Philadelphia (Philadelphia County)

On-site

USD 95,000 - 159,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Annual incentive bonus
Country-specific benefits

Job summary

Elsevier is seeking a Senior Site Reliability Engineer to lead reliability initiatives for our critical platforms and services in the United States. You will drive automation, improve observability, and reduce operational toil while ensuring high availability and performance of AI-enabled applications.

You will collaborate with engineering teams to design resilient systems, participate in incident response, and build robust handover capabilities to ensure operability after squads move on.

Qualifications

  • Advanced Terraform including modules, state management, drift detection and remote state handling.
  • Hands-on AWS operations across multi-account, multi-region environments including ECS, RDS, S3, DynamoDB, Lambda, SQS and KMS.
  • Experience with GitHub Actions CI/CD workflows, OIDC auth, and deployment pipelines.
  • Experience with ECS Fargate, Docker, ECR, IAM roles, health checks, autoscaling and deployment rollbacks.
  • Proficiency in AWS Networking & Security: VPCs, ALBs, Route53, TLS, IAM, Secrets Manager, KMS, cloud security best practices.
  • Strong Linux skills and scripting (Bash/Python) for automation and tooling.
  • Experience integrating AI services in production with monitoring and security considerations.
  • Ability to support multiple teams, document solutions, and enable self-service across infra and apps.

Responsibilities

  • Create monitoring queries and establish service level baselines.
  • Support senior engineers during incidents and post-mortems/RCA analysis.
  • Contribute to disaster recovery tests and reliability improvements.
  • Implement automation and execute code in production environments.
  • Document SRE knowledge and runbooks for operations handover.
  • Support deployment, monitoring, and reliability of AI-enabled services.
  • Assist in creating infrastructure topology drawings and deployment workflows.
  • Test availability, reliability, and recoverability in non-production environments.

Skills

Terraform
AWS
CI/CD
Linux
Observability
Incident Response
Python Bash
AI Tooling

Tools

GitHub Actions
ECS Fargate
Docker
ECR
VPC
Route53
CloudWatch

Job description

Senior Site Reliability Engineer

Are you passionate about building resilient, scalable systems that power mission-critical applications?

Do you thrive on automating operations, improving reliability, and ensuring exceptional system performance?

About The Team

Embedded Innovation Teams are cross-functional squads embedded within our segments to rapidly turn internal AI experimentation into validated, reusable solutions, building the capabilities we need to deliver customer value and growth. We work problem-first rather than tool-first, directly inside segment and function teams, improving the internal workflows that help our people deliver better outcomes for customers, faster.

About The Role

As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences.

You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve service availability, streamline operations, and enhance system recovery capabilities.

You will hold a high bar on code quality, flag risks and blockers early, and work alongside host-function stakeholders to make sure what you build fits real workflows, not assumed ones. You will also support handover and capability-building so the solution is owned and operable after the squad moves on.

Key Responsibilities
  • Creating monitoring queries and establishes service level baselines.
  • Supporting senior engineers during incidents.
  • Making contributions during post-mortems and RCAs.
  • Participating in disaster recovery tests.
  • Implementing automation and executes code in production environments.
  • Contributing to SRE knowledge documentation.
  • Supporting the deployment, monitoring, and reliability of services integrating AI tools.
  • Supporting architecture and senior engineers in the creation of infrastructure topology drawings and deployment workflows.
  • Carrying out the testing of availability, reliability, and recoverability in non-production environments.
Requirements
  • Advanced Terraform: Expertise in modules, providers, state management, lifecycle controls, drift detection, safe refactoring, and remote state (S3, locking, cross-stack dependencies).
  • AWS Operations: Hands-on experience managing production, multi-account, multi-region AWS environments across ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
  • GitHub Actions CI/CD: Experience building and troubleshooting reusable workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
  • ECS Fargate & Containers: Knowledge of Docker, ECR, ECS task definitions/services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
  • AWS Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices.
  • Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.
  • Linux & Automation: Strong Linux and Git fundamentals with Bash/Python scripting for AWS CLI automation, CI/CD, and operational tooling.
  • AI Tooling Deployment: Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices for AI-powered features.
  • Developer Enablement: Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.

Elsevier is a renowned global information analytics company that primarily focuses on providing scientific, technical, and medical (STM) research content, tools, and services. It is one of the largest publishers of academic journals and scholarly literature in the world.

Elsevier operates in various domains, including science, technology, medicine, social sciences, and more. They publish a vast number of peer-reviewed journals covering a wide range of disciplines. These journals act as platforms for researchers and academics to share their findings and contribute to the advancement of knowledge in their respective fields.

In addition to publishing, Elsevier offers a suite of digital solutions and services to support researchers, scientists, and professionals in their work. They provide online platforms like ScienceDirect, Scopus, and Mendeley, which offer access to a vast repository of scholarly articles, research papers, and other scientific content. These platforms often serve as essential resources for software developers seeking to stay updated with the latest scientific advancements.

U.S. National Base Pay Range: $95,300 - $158,800. Geographic differentials may apply in some locations to better reflect local market rates. This job is eligible for an annual incentive bonus.

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer I
Senior Site Reliability Engineer I

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Senior Site Reliability Engineer I
Senior Site Reliability Engineer I

RX Brasil • Philadelphia

On-site
USD 95,000 - 159,000
Annual incentive bonus
AI Software Engineering Lead
AI Software Engineering Lead

RXinsider LTD. • Philadelphia

On-site
USD 115,000 - 192,000
Annual incentive bonus
Country-specific benefits
Principal Software Engineer/Principal AI Engineer
Principal Software Engineer/Principal AI Engineer

Elsevier • Philadelphia

On-site
USD 115,000 - 192,000
Generous vacation entitlement
Comprehensive Pension Plan
Family leave and sabbatical options
+3
Senior Software Engineer I - Ruby on Rails with Sidekiq
Senior Software Engineer I - Ruby on Rails with Sidekiq

Elsevier • Philadelphia

On-site
USD 87,000 - 144,000
Systems Engineer
Systems Engineer

RELX • Gainesville (FL)

On-site
USD 71,600 - 119,400
React Node Senior Software Engineer I
React Node Senior Software Engineer I

Elsevier Inc. Company • Philadelphia

On-site
USD 87,000 - 144,000
Senior Security Engineer - Sec Ops
Senior Security Engineer - Sec Ops

RXinsider LTD. • New Jersey

On-site
USD 93,000 - 149,000
Senior ML Ops Engineer
Senior ML Ops Engineer

RELX • Philadelphia

On-site
USD 95,000 - 159,000
Systems Engineer
Systems Engineer

RXinsider LTD. • Gainesville (FL)

On-site
USD 72,000 - 119,000