Sr. Site Reliability Engineer

Staffing Science

Arizona

On-site

USD 180,000 - 240,000

Full time

39 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Staffing Science is seeking a Senior Site Reliability Engineer (Cloud Communications / High-Compliance Enterprise SaaS) to join a global, always-on platform. You will own AWS infrastructure at scale, implement Terraform/Terragrunt modules, and drive CI/CD across infrastructure and apps.

The role emphasizes deep OS/networking, production experience with Tomcat/Java, and advanced observability using Prometheus, Grafana, and ELK. Expect on-call rotations and collaboration with security leadership.

Qualifications

  • 15+ years of Linux/UNIX systems engineering, with meaningful time in SRE/DevOps roles at enterprise scale.
  • Expert-level AWS (EC2, S3, RDS, VPC, IAM, Lambda, ECS/EKS, CloudWatch) and infrastructure design best practices.
  • Expert-level Terraform and Terragrunt (module authorship required).
  • Strong production experience in Python and Go.
  • Deep enterprise networking fundamentals — DNS, TCP/IP, mail routing and delivery troubleshooting.

Responsibilities

  • Design, write, and maintain Terraform and Terragrunt modules managing AWS infrastructure at scale.
  • Own and evolve CI/CD pipelines across GitHub Actions and AWS CodePipeline for infrastructure and deployments.
  • Troubleshoot deep networking and mail deliverability issues at global scale.
  • Maintain Apache and Postfix-based mail infrastructure in production.
  • Provide SRE support into Tomcat/Java-based environments and diagnose memory/GC issues.
  • Build and maintain observability and alerting with Prometheus, Grafana, OpenTelemetry, and ELK; define SLOs/SLIs.

Skills

Linux/UNIX
AWS expert
SRE/DevOps
Python/Go
Networking fundamentals
Java/Tomcat
Agile/Scrum
On-call experience

Tools

Terraform
Terragrunt
GitHub Actions
AWS CodePipeline
Prometheus
Grafana
OpenTelemetry
ELK stack
New Relic
Jaeger
Zipkin
Docker
ECS/EKS

Job description

Senior Site Reliability Engineer — Cloud Communications / High-Compliance Enterprise SaaS

Our client is a large-scale, high-compliance cloud communications and healthcare data company that does business with the federal government and requires all employees to be U.S. citizens. For over 25 years, they've been a leader in their space and they're continuing to invest heavily in infrastructure modernization, automation, and platform reliability at global scale.

They are adding a 4th member to a tight-knit SRE team supporting mission-critical infrastructure that processes and delivers data across a global, high-volume, always-on platform. This is a highly visible role — you'll partner directly with Engineering and Information Security leadership on infrastructure strategy, not just execute tickets.

This role is built for a systems-first engineer — ideally a ~15–20 year Linux/UNIX systems engineer who has evolved into an SRE, rather than a developer who transitioned into infrastructure. This is not a role for someone purely cloud-native with no deep OS/networking background. The team needs someone who has genuinely run enterprise-scale platforms and lived through the operational fires that come with it — not someone who has only worked greenfield, cloud-native environments.

What you'll own:
  • Design, write, and maintain Terraform and Terragrunt modules (not just consumption) managing AWS infrastructure at scale, including EC2, S3, RDS, VPC, IAM, Lambda, and both ECS and EKS environments
  • Own and evolve CI/CD pipelines across GitHub Actions and AWS CodePipeline (GitLab experience a plus) — both for infrastructure and application deployments
  • Troubleshoot deep networking and mail delivery issues — DNS, IP/sender reputation, deliverability at global scale, and the class of problems that come with high-volume mail transport
  • Support and maintain Apache and Postfix-based mail infrastructure in production
  • Provide SRE support into Tomcat/Java-based application environments — diagnosing the class of problems Java/Tomcat teams run into (memory/GC issues, thread pool exhaustion, connection pooling, API contract issues) even without being a full-time Java developer
  • Build and maintain observability and alerting across the stack using tools like Prometheus, Grafana, OpenTelemetry, and the ELK stack; define and manage SLOs/SLIs for critical services
  • Support and extend configuration automation using tools such as AWS Config, SSM, Ansible, Puppet, or Chef
  • Work hands-on with containerization (Docker) and container orchestration, with particular depth in Amazon ECS
  • Use APM tooling (New Relic, Jaeger, Zipkin, or OpenTelemetry) to identify and resolve performance bottlenecks across distributed systems
  • Operate in a heavy on-call rotation supporting large-scale, high-availability systems (once fully staffed: 1 week on, 3 weeks off); lead and contribute to blameless postmortems following incidents
  • Draft and contribute to RFCs and internal standards for IaC, automation, and infrastructure design patterns
  • Mentor other engineers — including junior/lower-level ops staff — on Linux systems, IaC, CI/CD, and operational best practices
  • Ramp expectation: contribute to the codebase within 30 days, understand the environment deeply and take on smaller scoped projects by 60 days, own projects independently by 90 days
What you bring:
  • 15+ years of Linux/UNIX systems engineering, with meaningful time in SRE/DevOps roles at enterprise scale
  • Expert-level AWS (EC2, S3, RDS, VPC, IAM, Lambda, ECS/EKS, CloudWatch) and infrastructure design best practices
  • Expert-level Terraform and Terragrunt (module authorship required — this is not a "consume existing modules" role)
  • Strong production experience in Python and Go
  • Deep enterprise networking fundamentals — DNS, TCP/IP, mail routing and delivery troubleshooting, sender/IP reputation management
  • Hands-on Apache and Postfix administration experience
  • Prior exposure to Tomcat/Java-based production environments and the operational issues specific to that stack
  • Experience with configuration automation tooling (Ansible, Puppet, Chef, or AWS Config/SSM)
  • Solid containerization background (Docker), with strong familiarity in the ECS ecosystem specifically
  • Experience with observability/monitoring frameworks at scale (Prometheus, Grafana, OpenTelemetry, ELK, Thanos)
  • Experience operating in large-scale, high-compliance enterprise environments (healthcare, fintech, or similarly regulated industries — PCI/HITRUST/FedRamp not required, but that class of environment strongly preferred)
  • Customer-facing operational support background (UNIX/Windows) is a strong plus
  • Comfortable working within Agile/Scrum and Waterfall environments, using Jira/Confluence
  • Comfortable digesting large amounts of information quickly and moving from context to independent contribution fast
  • A self-starter mentality — able to work independently with minimal oversight while staying aligned with team goals
You’ll stand out if you also have:
  • Experience mentoring engineers into stronger DevOps/SRE practices
  • Background operating in PCI, HITRUST, or FedRamp/GovCloud-adjacent environments
  • Experience leading a team or platform through an SDLC/DevOps maturity transition
  • A GitHub profile or code samples that showcase personal or professional projects
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Town of Florida (NY)

On-site
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Calance • United States

Hybrid
USD 150,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000