Senior SRE - Real-Time Batch Platform Reliability (Remote)

BIP US

Chicago (IL)

Hybrid

USD 130,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Discretionary bonus
Remote/hybrid work environment

Job summary

BIP US is seeking an experienced operations-focused engineer to own the real-time reliability, observability, and stability of the SFRC batch platform. This role acts as the first line of defense for platform services, combining monitoring, incident response, root cause analysis, and automation to ensure critical risk and regulatory reporting processes run on time.

You will drive continuous service improvement, reduce manual support, and champion SRE/DevOps best practices across teams, with a

Qualifications

  • 7+ years’ experience supporting production batch-processing environments.
  • Experience with scheduling/orchestration platforms such as Quartz, AutoSys, Control-M, or equivalent tools.
  • Experience with monitoring and logging tools such as Splunk, AppDynamics, Grafana, Bob Monitor, or equivalent platforms.
  • Experience with automation and scripting tools such as Python, Shell, and Ansible.
  • Strong incident management and operational support experience.
  • Ability to work effectively during critical reporting windows.
  • Strong communication and stakeholder-management skills.
  • Ability to analyze issues and coordinate resolutions across multiple teams.

Responsibilities

  • Own real-time operational visibility of the SFRC batch platform, ensuring that critical business processes execute reliably and within agreed SLAs.
  • Monitor batch health, workflow orchestration, infrastructure dependencies, application performance, and data-processing pipelines across the end-to-end service landscape.
  • Understand the business context to proactively identify performance degradation, bottlenecks, capacity constraints, and failure patterns before they impact downstream consumers.
  • Perform first-line operational triage, log analysis, and root cause investigations using observability and monitoring platforms.
  • Coordinate technical recovery activities during incidents, including reruns, restarts, failover procedures, and troubleshooting across applications, infrastructure, and data domains.
  • Analyse operational metrics, trends, and recurring issues to identify opportunities for automation, optimization, and stability improvements.
  • Champion SRE and DevOps best practices, including automation, observability, continuous improvement, and operational excellence.

Skills

Incident management
SRE/DevOps
Observability
Stakeholder management
Communication

Tools

Quartz
AutoSys
Control-M
Splunk
AppDynamics
Grafana
Python
Shell
Ansible

Job description

BIP US is seeking an experienced operations-focused engineer to own the real-time reliability, observability, and stability of the SFRC batch platform. This role acts as the first line of defense for platform services, combining monitoring, incident response, root cause analysis, and automation to ensure critical risk and regulatory reporting processes run on time.

You will drive continuous service improvement, reduce manual support, and champion SRE/DevOps best practices across teams, with a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote SRE for Batch Platform Reliability & Automation
Remote SRE for Batch Platform Reliability & Automation

BIP US • New York (NY)

Hybrid
USD 130,000 - 190,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior SRE Lead — Real-Time Healthcare Platform
Senior SRE Lead — Real-Time Healthcare Platform

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Senior SRE — Real-Time Systems & Incident Leadership
Senior SRE — Real-Time Systems & Incident Leadership

LSEG • St. Louis (MO)

On-site
USD 100,000 - 130,000
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Senior SRE — AI-Driven Reliability & Oncall Leadership
Senior SRE — AI-Driven Reliability & Oncall Leadership

Block • San Francisco (CA)

On-site
USD 160,700 - 283,600
Healthcare coverage
Health Savings Account
Retirement Plans
+5
Senior SRE - Automation, Cloud & Reliability
Senior SRE - Automation, Cloud & Reliability

Shrive Technologies • Schaumburg (IL)

On-site
USD 120,000 - 180,000
Senior SRE: Build Scalable, Reliable Platforms — Remote
Senior SRE: Build Scalable, Reliable Platforms — Remote

Nord Security • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Premium healthcare
Work from anywhere
Mentorship programs
+3
Senior SRE Lead: Reliability, Observability & Automation
Senior SRE Lead: Reliability, Observability & Automation

Bank of America • Chandler (AZ)

On-site
USD 180,000 - 240,000