Senior SRE Engineer - Cloud Reliability & Automation

SFE

Englewood Cliffs (NJ)

On-site

USD 130,000 - 170,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

SFE in New Jersey is seeking an experienced Site Reliability Engineer to design and operate highly available production systems. You will implement automation, monitor health, and ensure service reliability across cloud and on-prem components.

Responsibilities include incident management, RCA, capacity planning, and collaboration with development, QA, cloud, and support teams to improve deployment processes and reliability.

Qualifications

  • 6-7 years of experience in Site Reliability Engineering, Production Support, DevOps, or Infrastructure Operations.
  • Strong understanding of Linux administration and troubleshooting.
  • Hands-on experience with AWS cloud services (EC2, RDS, IAM, VPC, CloudWatch, S3).
  • Experience with monitoring, alerting, and observability tools.
  • Knowledge of incident management, problem management, and RCA processes.
  • Experience with automation and scripting using Shell and/or Python.
  • Working knowledge of PostgreSQL and MySQL databases.
  • Experience with Git version control.
  • Understanding of CI/CD concepts and tools such as Jenkins.

Responsibilities

  • Design, implement, and maintain highly available and reliable production systems.
  • Automate operational tasks and infrastructure management using Shell, Python, Ansible, or Terraform.
  • Manage and support AWS services including EC2, RDS, S3, IAM, VPC, CloudWatch, and related cloud services.
  • Perform Linux server administration, troubleshooting, patching, and performance tuning.
  • Monitor application and infrastructure health using tools such as Grafana, Prometheus, CloudWatch, Datadog, Splunk.
  • Participate in incident management, root cause analysis (RCA), and problem management activities.
  • Define and maintain SLIs, SLOs, and SLAs to ensure service reliability.
  • Support PostgreSQL and MySQL databases for operational and basic administration tasks.
  • Collaborate with development, QA, cloud, and support teams to improve system reliability and deployment processes.
  • Drive automation, observability, capacity planning, security, and operational best practices.
  • Participate in on-call support and production issue resolution.

Skills

Linux administration
AWS
Observability
Automation
Scripting
PostgreSQL
MySQL
Git
CI/CD concepts

Tools

Jenkins
Terraform
Ansible
Shell
Python
Grafana
Prometheus
CloudWatch
Datadog
Splunk

Job description

SFE in New Jersey is seeking an experienced Site Reliability Engineer to design and operate highly available production systems. You will implement automation, monitor health, and ensure service reliability across cloud and on-prem components.

Responsibilities include incident management, RCA, capacity planning, and collaboration with development, QA, cloud, and support teams to improve deployment processes and reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Senior SRE: Scale Systems with Automation & Observability
Senior SRE: Scale Systems with Automation & Observability

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000
Senior Cloud SRE — AWS, Automation & Security
Senior Cloud SRE — AWS, Automation & Security

Okta • Bellevue (CA)

On-site
USD 180,000 - 260,000
Amazing Benefits
Making Social Impact
Diversity, Equity & Inclusion
Senior SRE II: Automate & Scale Cloud Reliability
Senior SRE II: Automate & Scale Cloud Reliability

Genuine Parts • Birmingham (AL)

On-site
USD 90,000 - 120,000
Senior SRE: Architect Scalable, Reliable Cloud Infra
Senior SRE: Architect Scalable, Reliable Cloud Infra

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer (SRE) – Automation & Cloud Ops
Senior Site Reliability Engineer (SRE) – Automation & Cloud Ops

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Senior SRE (FedRAMP) – Cloud Reliability & Automation
Senior SRE (FedRAMP) – Cloud Reliability & Automation

Empleora • Northern (KY)

Hybrid
USD 160,000 - 210,000
Senior Cloud SRE: Reliability, Automation & Remote Work
Senior Cloud SRE: Reliability, Automation & Remote Work

IO Connect Services • United States

Remote
USD 120,000 - 180,000
Base salary and permanent contract
Paid certifications
Career plan
+9
Senior SRE: Reliability, CI/CD & Secure Cloud
Senior SRE: Reliability, CI/CD & Secure Cloud

NeuroFlow • Philadelphia

On-site
USD 140,000 - 210,000
Flexible work schedule
Unlimited PTO
Medical coverage
+6
Senior SRE Engineer: Cloud Reliability & Platform
Senior SRE Engineer: Cloud Reliability & Platform

flowcode • New York (NY)

Hybrid
USD 140,000 - 190,000