Site Reliability Engineer III- Network

JPMorganChase

Hyderabad

On-site

INR 2,500,000 - 4,200,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorganChase Hyderabad is seeking an experienced Site Reliability Engineer to design and operate enterprise network services, drive incident management, and advance reliability across complex, mission-critical systems.

You will lead day-to-day operations, implement automation with Python, Shell and Ansible, and partner with development teams to improve observability, guardrails, and recovery processes.

Qualifications

  • Formal training or certification in site reliability engineering concepts with 3+ years of hands-on experience.
  • Experience in incident response and problem management with end-to-end RCA ownership.
  • Strong networking skills across enterprise routing, switching and security.
  • Automation skills using Python, Shell, and Ansible in production.
  • SRE mindset with understanding of reliability, NF Rs and risk analysis.
  • Ability to work independently with minimal supervision.
  • Experience with software-defined networking environments (e.g., SD-WAN, SDA).
  • Ability to build monitoring/observability dashboards and alerts.

Responsibilities

  • Lead day-to-day ownership for network services, including complex troubleshooting and restoration.
  • Drive incident management and produce high-quality RCAs with corrective actions.
  • Participate in major incident management with communications support and mitigation.
  • Design and implement production-grade automation using Python, Shell, and Ansible.
  • Engineer SD-WAN, SDA, and software-defined networking capabilities.
  • Support routing and switching across enterprise networks and security components.
  • Improve reliability with standardization, runbooks, and observability in partnership with developers.
  • Apply enterprise AI capabilities to identify reliability risks and speed triage.

Skills

Incident response
Problem management
Networking
Automation
SRE mindset
Independent work
Cross-team coordination
Communication during incidents

Tools

Python
Shell
Ansible

Job description

Job Description:

There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the worlds most complex and mission‑critical systems.

Job responsibilities
  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
  • Lead day‑to‑day operational ownership for network services, including complex troubleshooting and coordinated restoration.
  • Drive incident and problem management by running structured investigations, producing high‑quality RCAs, and ensuring corrective/preventive actions are delivered.
  • Participate in major incident management, providing communications support, technical lead support, and mitigation execution.
  • Design and implement production‑grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
  • Engineer and support software‑defined networking capabilities, including SD‑WAN, SDA, and broader SND.
  • Engineer and support routing and switching across enterprise networks.
  • Engineer and support security and L4–L7 network components, including firewalls, load balancers, and proxies.
  • Improve reliability through standardization, guardrails, repeatable runbooks, continuous validation, and observability (dashboards, high-signal alerting, service health metrics) in partnership with developers/platform teams.
  • Applies enterprise‑authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse‑first improvements tied to SLO outcomes.
  • Uses enterprise‑authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post‑incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience
  • Demonstrated experience in incident response and problem management, including end-to-end ownership of RCAs through closure.
  • Strong hands‑on networking skills across enterprise routing/switching and security/L4–L7 components.
  • Strong automation capability using Python, Shell, and Ansible in production operations.
  • SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including familiarity with FMEA or equivalent methods).
  • Ability to work independently, prioritize effectively, and deliver with minimal oversight.
  • Experience supporting software‑defined networking environments (e.g., SD‑WAN, SDA, and related tooling).
  • Ability to build and operationalize monitoring/observability, including dashboards, alerting, and service health metrics.
  • Strong communication and coordination skills during high‑severity incidents and cross‑team restoration efforts.
  • Working knowledge of using enterprise‑authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI‑assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
Preferred qualifications, capabilities, and skills
  • Demonstrate experience with Cisco ACI / fabrics.
  • Hold relevant certifications such as CCNP (preferred), CCNA, or other vendor certifications.
  • Work effectively in a financial institution or other regulated environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer III- Network
Site Reliability Engineer III- Network

JPMorgan Chase & Co. • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Senior Network Site Reliability Engineer
Senior Network Site Reliability Engineer

NVIDIA • Maharashtra

On-site
INR 4,000,000 - 6,000,000
Senior Network Site Reliability Engineer
Senior Network Site Reliability Engineer

NVIDIA • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Network Site Reliability Engineer
Senior Network Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Lead Site Reliability Engineer- Network
Lead Site Reliability Engineer- Network

JPMorganChase • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Senior Network Site Reliability Engineer
Senior Network Site Reliability Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 600,000
Lead Site Reliability Engineer- Network
Lead Site Reliability Engineer- Network

JPMorgan Chase & Co. • Bengaluru

On-site
INR 3,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

ACI Worldwide • Maharashtra

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

ACI Worldwide, Inc. • Pune District

On-site
INR 900,000 - 1,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000