Senior SRE: Reliability, Automation & AI Platforms

RX Brasil

Philadelphia (Philadelphia County)

On-site

USD 95,000 - 159,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Annual incentive bonus

Job summary

Elsevier in the United States is seeking a Senior Site Reliability Engineer to lead reliability, scalability, and performance across critical platforms. You will drive automation, improve observability, and partner with engineering to design resilient systems.

The role emphasizes incident response, post-mortems, and production readiness, with a focus on AI tooling deployment and secure self-service for multiple teams.

Qualifications

  • Advanced Terraform: modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state.
  • AWS Operations: production multi-account/multi-region environments across ECS, RDS, ALB, VPC, IAM, Route53, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, CloudWatch.
  • GitHub Actions CI/CD: reusable workflows, OIDC, approval gates, runners, Terraform deployments, app deployments, migrations.
  • ECS Fargate & Containers: Docker, ECR, ECS task definitions/services, IAM roles, health checks, autoscaling, ALB, deployment rollbacks.
  • AWS Networking & Security: VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, cloud security best practices.
  • Incident Response & Observability: troubleshooting with logs, metrics, alarms, RCAs, rollback decisions, runbooks.
  • Linux & Automation: Linux and Git basics with Bash/Python scripting for AWS CLI automation, CI/CD, tooling.
  • AI Tooling Deployment: deploying AI services in production with monitoring, reliability, security for AI-powered features.
  • Developer Enablement: support multiple teams, troubleshoot infra/app layers, document solutions, enable self-service.

Responsibilities

  • Creating monitoring queries and establishing service level baselines.
  • Support senior engineers during incidents.
  • Contribute to post-mortems and RCAs.
  • Participate in disaster recovery tests.
  • Implement automation and execute code in production.
  • Contribute to SRE knowledge documentation.
  • Support deployment, monitoring, and reliability of AI-integrated services.
  • Assist architecture and engineers in infrastructure topology drawings and deployment workflows.
  • Test availability, reliability, and recoverability in non-production environments.

Skills

Terraform
AWS
GitHub Actions
Docker
ECS
ECR
CI/CD
Networking
Observability
Incident Response
Linux
Automation
Security
AI Tooling
Developer Enablement

Tools

Terraform
AWS
GitHub Actions
Docker
ECS
ECR
Route53
VPC
CloudWatch
IAM
KMS
Secrets Manager

Job description

Elsevier in the United States is seeking a Senior Site Reliability Engineer to lead reliability, scalability, and performance across critical platforms. You will drive automation, improve observability, and partner with engineering to design resilient systems.

The role emphasizes incident response, post-mortems, and production readiness, with a focus on AI tooling deployment and secure self-service for multiple teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Automation & Reliability Champion
Senior SRE: Automation & Reliability Champion

RXinsider LTD. • Philadelphia

On-site
USD 95,000 - 159,000
Annual incentive bonus
Country-specific benefits
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,100 - 283,600
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior SRE Platform Engineer – AI-Powered Reliability
Senior SRE Platform Engineer – AI-Powered Reliability

UiPath • Denver (CO)

Hybrid
USD 160,000 - 210,000
Senior SRE: AI-Driven Cloud Reliability & Automation
Senior SRE: AI-Driven Cloud Reliability & Automation

Hidden Jobs • United States

Remote
USD 191,000 - 226,000
Equity incentive
Flexible PTO
Health insurance
+2
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Senior SRE: AI-Driven Reliability & Automation (Hybrid)
Senior SRE: AI-Driven Reliability & Automation (Hybrid)

Namely • United States

Hybrid
USD 120,000 - 150,000
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1