Senior Site Reliability Engineer

Holistic Partners, Inc

Deerfield (IL)

Hybrid

USD 140,000 - 200,000

Full time

18 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Holistic Partners, Inc. in Riverwoods, IL is seeking an experienced Site Reliability Engineer for a hybrid role (2-3 days in the office).

You will partner with Application Development teams to build resiliency, implement SLOs, and drive observability across payment applications. Responsibilities include monitoring, alerting, dashboards, automation, capacity planning, disaster recovery planning, and participating in on-call.

Qualifications

  • 12+ years of experience as an SRE or in a similar role.
  • Hands-on software development with SDLC
  • Experience with DevOps and CI/CD tools
  • Strong Linux, AWS Cloud, and on-prem deployments
  • Experience with observability tools (APM, Datadog)
  • Knowledge of Grafana, Kibana, and release management

Responsibilities

  • Partner with Application Development teams to build resiliency for payment applications.
  • Partner with Application Development teams to implement SLOs.
  • Collaborate with SREs to enable end-to-end observability across systems.
  • Implement monitoring, alerting, and dashboards for applications.
  • Automate operational processes and drive capacity/performance tools.
  • Contribute to DR planning for critical apps and participate in on-call rotations.

Skills

AWS
Python
Go
Shell scripting
Java
Datadog
Grafana
Kibana
Jenkins
Kubernetes
OpenShift
Ansible
SNOW
Linux
SQL
JIRA
Release Management

Tools

OpenShift
Kubernetes
Jenkins
Ansible
Datadog
ELK
Grafana
Kibana
RabbitMQ
Kafka
Hadoop
Spark
SNOW
Jira

Job description

Location: Riverwoods, IL (hybrid role requiring 2-3 days per week in the office)
Job Description
Responsibilities:
  • Partner with Application Development teams to build resiliency for payment applications.
  • Partner with Application Development teams to implement Service Level Objectives (SLOs).
  • Work with Application Development teams and other SREs to build end-to-end observability.
  • Implement monitoring, alerting, and dashboards for applications.
  • Automate operational processes.
  • Help develop capacity and performance management tools.
  • Help define the Disaster Recovery (DR) plan for critical applications.
  • Participate in an on-call rotation and support production incidents.
SRE Skillsets - Pricing & Settlements Team:
  • Strong understanding of hybrid infrastructure.
  • Expertise in AWS.
  • Expertise in one or more general-purpose programming languages: Python, Go, Shell scripting (Unix/Linux), or Java.
  • Experience with CI/CD pipelines, preferably Jenkins.
  • Experience with container technologies such as OpenShift or Kubernetes.
  • Experience with automation tools, preferably Ansible.
  • Expertise in observability tools, including APM (Datadog), synthetic monitoring, and log aggregation (ELK).
  • Experience with dashboarding tools such as Grafana and Kibana.
  • Understanding Agile concepts and experience with JIRA.
  • Basic understanding of Release Management.
  • Hands-on experience with ServiceNow (SNOW).
SRE Skillsets - Data Platform Team:
  • Expertise with message brokers, preferably RabbitMQ or Kafka.
  • Experience with Hadoop and Spark commands, JSON formatting, etc.
Qualifications:
  • 12-15 years of overall experience.
  • Professional experience as a Site Reliability Engineer (SRE).
  • Hands-on software development experience with a strong understanding of SDLC and application delivery.
  • Ability to translate functional and non-functional requirements into appropriate NFT automation tests.
  • Experience with DevOps and CI/CD tools.
  • Strong experience with Linux, AWS Cloud, and on-premises deployments.
  • Strong experience with systems observability and APM tools, preferably Datadog.
  • Ability to contribute to technical discussions around application integration, high availability, resiliency, and observability.
  • Strong knowledge of JIRA.
Key Skills:
  • AWS Lambda Services - 4/5
  • Strong database concepts, SQL, and MySQL - 4/5
  • Hands-on experience with ServiceNow (SNOW)
Must Have:
  • Professional experience as an SRE.
  • Experience in performance testing and the ability to translate functional and non-functional requirements into appropriate NFT automation tests.
  • Strong experience with AWS Cloud applications.
  • Strong experience with Linux, AWS Cloud, and on-premises deployments.
  • Strong experience with systems observability and APM tools, preferably Datadog.
  • Experience with Grafana and Kibana.
  • Strong understanding of application integration, high availability, resiliency, and observability.
  • Expertise in Python, Shell scripting (Unix/Linux), or Java.
Nice to Have:
  • Hands-on experience with ServiceNow (SNOW).
  • Experience with OpenShift or Kubernetes.
  • Strong JIRA knowledge.
  • Basic understanding of Release Management.
  • Experience with CI/CD pipelines, preferably Jenkins.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

On-site
USD 120,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
SRE Engineer
SRE Engineer

Programmers.io • Austin (TX)

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

On-site
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Madison-Davis, LLC • United States

On-site
USD 120,000 - 150,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

On-site
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)