Senior SRE: Production Reliability & Automation

Bank of America

Jersey City (NJ)

On-site

USD 108,000 - 162,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Bank of America in Jersey City, NJ is seeking an experienced Production Support/SRE professional to partner with engineering teams to implement instrumentation, tooling, and on-call routines for key services, focusing on reliability and rapid incident resolution. The role encompasses real-time monitoring, incident triage, root-cause analysis, and collaboration with cross-functional teams across regions.

A strong background in Java/J2EE, Linux, and monitoring tools is required, with on-call

Qualifications

  • 7+ years of production support experience.
  • Experience supporting Java/J2EE apps in an enterprise environment, including WebLogic, web services, Spring Boot, and strong SQL/PL/SQL skills for troubleshooting.
  • Strong working knowledge of Linux/Unix environments and scripting languages such as Shell/Python, including applications deployed on JBoss.
  • Experience using monitoring and observability tools (e.g., Splunk, Dynatrace, Nastel, SiteScope) in a production support environment.
  • Hands on experience troubleshooting network related production incidents, including load balancing, traffic routing, and DNS issues, across on prem and cloud environments, with a focus on rapid service restoration.
  • Experience and understanding database concepts (SQL / Oracle) and writing basic queries.
  • Banking or Capital Markets domain knowledge preferred.
  • Ability to manage multiple tasks simultaneously and adapt quickly to changing priorities and production demands.
  • Self-starter with the ability to work independently as well as collaboratively within cross-functional teams.
  • Excellent analytical skills with the ability to identify root causes of complex production issues.
  • Familiarity with SRE principles including SLO/SLI definition, error budgets, and toil reduction strategies.
  • Experience identifying and automating repetitive operational tasks to reduce toil and improve team efficiency.
  • Understanding of CI/CD pipelines and ability to support post-deployment validation in an automated release environment.
  • Experience defining and owning application availability targets, contributing to reliability improvement plans, and driving proactive measures to prevent recurrence of production degradation.
  • Provide on-call rotational support, including off-hours support, during weeknights, Saturdays, and Sundays

Responsibilities

  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring Site Reliability Engineer (SRE) resources on reliability practices and established tools/capabilities
  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the SRE Lead
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and ‘noise’ in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations
  • Participates regularly in an on-call rotation with Production Support teammates to learn more about reliability issues affecting their portfolio
  • Provide front line production support and monitoring to ensure application stability and availability
  • Triage, troubleshoot, and resolve production incidents, including business impacting issues
  • Lead incident response and bridge calls, coordinating troubleshooting and escalation as needed
  • Perform root cause analysis and drive remediation and preventative actions
  • Monitor system alerts, logs, dashboards, and performance metrics to assess impact and restore service
  • Support batch jobs and data feeds, including time sensitive failures
  • Conduct post release validation and routine application health checks
  • Resolve user requests related to access, technical issues, and data discrepancies
  • Maintain accurate incident documentation, runbooks, and knowledge articles
  • Partner with technology, operations, vendors, and business teams to improve reliability
  • Participate in a rotational weekend and after hours support schedule as required

Skills

Production Support
Java/J2EE
Linux/Unix
Monitoring Tools
SQL/PLSQL
WebLogic
Spring Boot
CI/CD
On-Call
SRE Principles

Tools

Splunk
Dynatrace
Nastel
SiteScope
JBoss
WebLogic

Job description

Bank of America in Jersey City, NJ is seeking an experienced Production Support/SRE professional to partner with engineering teams to implement instrumentation, tooling, and on-call routines for key services, focusing on reliability and rapid incident resolution. The role encompasses real-time monitoring, incident triage, root-cause analysis, and collaboration with cross-functional teams across regions.

A strong background in Java/J2EE, Linux, and monitoring tools is required, with on-call

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — Automation, Observability & Reliability
Senior SRE — Automation, Observability & Reliability

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Senior SRE: Production Reliability & Automation
Senior SRE: Production Reliability & Automation

PNC Financial Services Group, Inc. • Denver (CO)

On-site
USD 86,000 - 158,000
Medical and prescription drug coverage
Dental and vision
401(k) with company match
+4
Site Reliability Engineer II
Site Reliability Engineer II

Bank of America • Jersey City (NJ)

On-site
USD 108,000 - 162,000
Senior Production SRE Engineer for Trading Platforms
Senior Production SRE Engineer for Trading Platforms

Deutsche Bank AG • Cary (NC)

Hybrid
USD 100,000 - 153,000
Senior SRE & Production Ops — Onsite (US Banks)
Senior SRE & Production Ops — Onsite (US Banks)

Federal Reserve Bank of Boston • Boston (MA)

On-site
USD 90,000 - 140,000
SRE Lead: Reliability, Incident Command & Automation
SRE Lead: Reliability, Incident Command & Automation

Relha LLC • Atlanta (GA), Northern (KY)

Hybrid
USD 112,000 - 131,000
Life insurance
Disability
Parental leave
+5
Production Services Lead: Incident & Change Expert
Production Services Lead: Incident & Change Expert

Bank of America • United States

On-site
CAD 65,000 - 95,000
Production Reliability Engineer (SRE & Automation)
Production Reliability Engineer (SRE & Automation)

NRnP Technology • Northern (KY)

Hybrid
USD 90,000 - 140,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Jersey City (NJ)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Dallas (TX)

Hybrid
USD 140,000 - 180,000