Site Reliability Engineer III- Network

JPMorgan Chase & Co.

Hyderabad

On-site

INR 2,500,000 - 4,500,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase & Co. in Hyderabad is seeking a Site Reliability Engineer III to own day-to-day network reliability, lead incident management, and drive automation across production systems.

You will implement production-grade automation with Python, Shell, and Ansible; work on SD-WAN, SDA, and security components; and partner with engineering teams to improve observability, dashboards, and SLO-driven improvements.

Qualifications

  • Formal training or certification on site reliability engineering concepts.
  • 3+ years of applied experience in incident response and problem management.
  • Strong hands-on networking across enterprise routing/switching and security components.
  • Automation capability using Python, Shell, and Ansible in production operations.
  • SRE mindset with practical reliability concepts and risk analysis.

Responsibilities

  • Guides and assists others in building appropriate designs and gaining consensus on SRE best practices.
  • Lead day-to-day operational ownership for network services and complex troubleshooting.
  • Drive incident and problem management with structured investigations and RCAs.
  • Participate in major incident management with communications and mitigation support.
  • Design and implement production-grade automation using Python, Shell, and Ansible.
  • Engineer and support SD-WAN, SDA, and broader SND networking capabilities.
  • Improve observability with dashboards, high-signal alerting, and service health metrics.
  • Validate AI-assisted operational recommendations before applying changes.

Skills

incident response
problem management
networking
automation
sre mindset
independent work
observability
incident communications
ai capabilities
data sensitivity

Education

SRE training/certification
CCNP
CCNA
regulated environment experience

Tools

Python
Shell
Ansible
SD-WAN
SDA
Cisco ACI

Job description

There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Platforms team, you will solve complex and broad business problems with simple and straightforward solutions. Network SRE who owns troubleshooting and reliability improvements across network platforms. Leads problem management for recurring issues, drives automation-first operations, and partners with development teams to improve observability, alert quality, and resilience.

Job responsibilities
  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
  • Lead day-to-day operational ownership for network services, including complex troubleshooting and coordinated restoration.
  • Drive incident and problem management by running structured investigations, producing high-quality RCAs, and ensuring corrective/preventive actions are delivered.
  • Participate in major incident management, providing communications support, technical lead support, and mitigation execution.
  • Design and implement production-grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
  • Engineer and support software-defined networking capabilities, including SD-WAN, SDA, and broader SND.
  • Engineer and support routing and switching across enterprise networks.
  • Engineer and support security and L4–L7 network components, including firewalls, load balancers, and proxies.
  • Improve reliability through standardization, guardrails, repeatable runbooks, continuous validation, and observability (dashboards, high-signal alerting, service health metrics) in partnership with developers/platform teams.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience
  • Demonstrated experience in incident response and problem management, including end-to-end ownership of RCAs through closure.
  • Strong hands-on networking skills across enterprise routing/switching and security/L4–L7 components.
  • Strong automation capability using Python, Shell, and Ansible in production operations.
  • SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including familiarity with FMEA or equivalent methods).
  • Ability to work independently, prioritize effectively, and deliver with minimal oversight.
  • Experience supporting software-defined networking environments (e.g., SD-WAN, SDA, and related tooling).
  • Ability to build and operationalize monitoring/observability, including dashboards, alerting, and service health metrics.
  • Strong communication and coordination skills during high-severity incidents and cross-team restoration efforts.
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
Preferred qualifications, capabilities, and skills
  • Demonstrate experience with Cisco ACI / fabrics.
  • Hold relevant certifications such as CCNP (preferred), CCNA, or other vendor certifications.
  • Work effectively in a financial institution or other regulated environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer- Network
Lead Site Reliability Engineer- Network

JPMorgan Chase & Co. • Bengaluru

On-site
INR 3,000,000 - 7,000,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorganChase • Hyderabad

On-site
INR 2,600,000 - 4,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • Maharashtra

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorgan Chase & Co. • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineer II
Site Reliability Engineer II

United States Digital Space LLC • Maharashtra

On-site
INR 1,500,000 - 2,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 1,500,000 - 2,000,000
Sr Lead Software Engineer
Sr Lead Software Engineer

JPMorgan Chase & Co. • Bengaluru

On-site
INR 2,400,000 - 4,000,000
Director of Infrastructure Engineering Network Rapid Response
Director of Infrastructure Engineering Network Rapid Response

JPMC • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Network Site Reliability Engineer
Senior Network Site Reliability Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 600,000
Senior Network Site Reliability Engineer
Senior Network Site Reliability Engineer

NVIDIA • Hyderabad

On-site
INR 1,200,000 - 2,400,000