Site Reliability Engineer II

Bank of America

Jersey City (NJ)

On-site

USD 108,000 - 162,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bank of America in Jersey City, NJ is seeking an experienced Production Support/SRE professional to partner with engineering teams to implement instrumentation, tooling, and on-call routines for key services, focusing on reliability and rapid incident resolution. The role encompasses real-time monitoring, incident triage, root-cause analysis, and collaboration with cross-functional teams across regions.

A strong background in Java/J2EE, Linux, and monitoring tools is required, with on-call

Qualifications

  • 7+ years of production support experience.
  • Experience supporting Java/J2EE apps in an enterprise environment, including WebLogic, web services, Spring Boot, and strong SQL/PL/SQL skills for troubleshooting.
  • Strong working knowledge of Linux/Unix environments and scripting languages such as Shell/Python, including applications deployed on JBoss.
  • Experience using monitoring and observability tools (e.g., Splunk, Dynatrace, Nastel, SiteScope) in a production support environment.
  • Hands on experience troubleshooting network related production incidents, including load balancing, traffic routing, and DNS issues, across on prem and cloud environments, with a focus on rapid service restoration.
  • Experience and understanding database concepts (SQL / Oracle) and writing basic queries.
  • Banking or Capital Markets domain knowledge preferred.
  • Ability to manage multiple tasks simultaneously and adapt quickly to changing priorities and production demands.
  • Self-starter with the ability to work independently as well as collaboratively within cross-functional teams.
  • Excellent analytical skills with the ability to identify root causes of complex production issues.
  • Familiarity with SRE principles including SLO/SLI definition, error budgets, and toil reduction strategies.
  • Experience identifying and automating repetitive operational tasks to reduce toil and improve team efficiency.
  • Understanding of CI/CD pipelines and ability to support post-deployment validation in an automated release environment.
  • Experience defining and owning application availability targets, contributing to reliability improvement plans, and driving proactive measures to prevent recurrence of production degradation.
  • Provide on-call rotational support, including off-hours support, during weeknights, Saturdays, and Sundays

Responsibilities

  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring Site Reliability Engineer (SRE) resources on reliability practices and established tools/capabilities
  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the SRE Lead
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and ‘noise’ in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations
  • Participates regularly in an on-call rotation with Production Support teammates to learn more about reliability issues affecting their portfolio
  • Provide front line production support and monitoring to ensure application stability and availability
  • Triage, troubleshoot, and resolve production incidents, including business impacting issues
  • Lead incident response and bridge calls, coordinating troubleshooting and escalation as needed
  • Perform root cause analysis and drive remediation and preventative actions
  • Monitor system alerts, logs, dashboards, and performance metrics to assess impact and restore service
  • Support batch jobs and data feeds, including time sensitive failures
  • Conduct post release validation and routine application health checks
  • Resolve user requests related to access, technical issues, and data discrepancies
  • Maintain accurate incident documentation, runbooks, and knowledge articles
  • Partner with technology, operations, vendors, and business teams to improve reliability
  • Participate in a rotational weekend and after hours support schedule as required

Skills

Production Support
Java/J2EE
Linux/Unix
Monitoring Tools
SQL/PLSQL
WebLogic
Spring Boot
CI/CD
On-Call
SRE Principles

Tools

Splunk
Dynatrace
Nastel
SiteScope
JBoss
WebLogic

Job description

Job Description:

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.

Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.

We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.

Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.

At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

Job Description:

This job is responsible for partnering with engineering and technology teams to implement measures as prescribed by lead/senior SRE engineers. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, identifying root causes of issues through production triage efforts, and suggesting code enhancements to technology teams to automate services and improve reliability and efficiency. Job expectations include using software development skills to improve efficiency and to address gaps in reliability.

Overview:

This role provides critical technical production and end‑user support for Capital Markets and Investment Banking platforms that enable deal execution, reporting, client mandate processing, and compliance clearance. The position supports multiple mission critical, time sensitive applications and partners closely with internal and external clients, technology teams, and business stakeholders across regions and time zones.

Key responsibilities include real time production monitoring, incident triage and resolution, root cause analysis, batch and data feed support, user service requests, post release validation, application health checks, and compliance related requests. The role leads business impacting incident response, ensures adherence to incident and problem management standards, restores service within strict SLAs, and drives root cause remediation and follow up actions.

Responsibilities:
  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring Site Reliability Engineer (SRE) resources on reliability practices and established tools/capabilities
  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the SRE Lead
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and ‘noise’ in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations
  • Participates regularly in an on-call rotation with Production Support teammates to learn more about reliability issues affecting their portfolio
  • Provide front line production support and monitoring to ensure application stability and availability
  • Triage, troubleshoot, and resolve production incidents, including business impacting issues
  • Lead incident response and bridge calls, coordinating troubleshooting and escalation as needed
  • Perform root cause analysis and drive remediation and preventative actions
  • Monitor system alerts, logs, dashboards, and performance metrics to assess impact and restore service
  • Support batch jobs and data feeds, including time sensitive failures
  • Conduct post release validation and routine application health checks
  • Resolve user requests related to access, technical issues, and data discrepancies
  • Maintain accurate incident documentation, runbooks, and knowledge articles
  • Partner with technology, operations, vendors, and business teams to improve reliability
  • Participate in a rotational weekend and after hours support schedule as required
Required Qualifications:
  • Minimum 7+ years of experience in application production support role

  • Production support experience supporting Java/J2EE applications in an enterprise environment, including WebLogic, web services, Spring Boot, and strong SQL/PL/SQL skills for troubleshooting

  • Strong working knowledge of Linux/Unix environments and scripting languages such as Shell/Python, including applications deployed on JBoss

  • Experience using monitoring and observability tools (e.g., Splunk, Dynatrace, Nastel, SiteScope) in a production support environment

  • Hands on experience troubleshooting network related production incidents, including load balancing, traffic routing, and DNS issues, across on prem and cloud environments, with a focus on rapid service restoration

  • Experience and understanding database concepts (SQL / Oracle) and writing basic queries

  • Banking or Capital Markets domain knowledge preferred

  • Ability to manage multiple tasks simultaneously and adapt quickly to changing priorities and production demands

  • Self-starter with the ability to work independently as well as collaboratively within cross-functional teams

  • Excellent analytical skills with the ability to identify root causes of complex production issues

  • Familiarity with SRE principles including SLO/SLI definition, error budgets, and toil reduction strategies

  • Experience identifying and automating repetitive operational tasks to reduce toil and improve team efficiency

  • Understanding of CI/CD pipelines and ability to support post-deployment validation in an automated release environment

  • Experience defining and owningapplication availability targets, contributing to reliability improvement plans, and driving proactive measures to prevent recurrence of production degradation

  • Provide on-call rotational support, including off-hours support, during weeknights, Saturdays, and Sundays

Desired Qualifications:
  • Proven experience as a proactive problem solver in a production support environment

  • Strong verbal and written communication skills, with the ability to convey technical concepts to both technical and non technical audiences

  • Ability to review system logs and monitor data to identify subtle performance anomalies

  • Strong prioritization skills with the ability to manage multiple incidents under tight deadlines

  • Collaborative mindset with experience working across cross functional teams

Skills:
  • Analytical Thinking
  • Automation
  • Collaboration
  • Production Support
  • Result Orientation
  • Application Development
  • Architecture
  • Influence
  • Project Management
  • Solution Design
  • Adaptability
  • DevOps Practices
  • Risk Management
  • Solution Delivery Process
  • Stakeholder Management
Shift:

1st shift (United States of America)

Hours Per Week:

40

Pay Transparency details

US - NJ - Jersey City - 525 Washington Blvd (NJ2525)

Pay and benefits information

Pay range

$108,000.00 - $161,900.00 annualized salary, offers to be determined based on experience, education and skill set.

Discretionary incentive eligible

This role is eligible to participate in the annual discretionary plan. Employees are eligible for an annual discretionary award based on their overall individual performance results and behaviors, the performance and contributions of their line of business and/or group; and the overall success of the Company.

Benefits

This role is currently benefits eligible. We provide industry-leading benefits, access to paid time off, resources and support to our employees so they can make a genuine impact and contribute to the sustainable growth of our business and the communities we serve.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 125,000 - 168,000
Industry-leading benefits
Discretionary incentive plan
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Plano (TX)

On-site
USD 125,000 - 168,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 125,000 - 168,000
Technology Services Lead - Global Markets Equities Production Services
Technology Services Lead - Global Markets Equities Production Services

Bank of America • New York (NY)

On-site
USD 92,000 - 161,000
Benefits eligible
In‑office role in New York
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Charlotte (NC)

On-site
USD 125,000 - 168,000
Discretionary incentive eligible
Annual discretionary plan
Senior Production Support Engineer - Assistant Vice President
Senior Production Support Engineer - Assistant Vice President

Deutsche Bank • Cary (NC)

Hybrid
USD 100,000 - 153,000
Hybrid work model
Flexible vacation & personal days
ERGs and community engagement
+1
Production Site Reliability Engineer
Production Site Reliability Engineer

Charles Schwab Corporation • Austin (TX)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Koitecc Solutions • Plano (TX), Northern (KY)

Hybrid
USD 153,000 - 192,000
Discretionary incentive
Benefits package