Infosys is seeking a Python SRE Senior Support Engineer in Chennai, India. The role involves contributing to production support and ensuring operational stability of applications. Responsibilities include managing the incident lifecycle, assisting in root cause analysis, and enhancing observability tools. Candidates should have a strong foundation in ITIL processes and hands-on experience with observability tools. Familiarity with Python, cloud fundamentals, and DevOps practices is preferred.
Qualifications
Strong foundation in incident, problem, and change management processes.
Hands-on experience with observability tools.
Ability to troubleshoot across infrastructure layers.
Familiarity with CI/CD pipelines and DevOps practices.
Ability to read and debug Java code.
Responsibilities
Contribute to production support and operational stability.
Manage the incident lifecycle and collaborate with teams.
Assist in root cause analysis and preventive measures.
Support observability setup and operational tasks.
Participate in on-call rotations and communicate effectively.
Skills
Incident, problem, and change management processes (ITIL v4 exposure preferred)
Observability tools (e.g., Splunk, Dynatrace)
Troubleshooting across infrastructure layers (Linux/Windows, networking, cloud)
CI/CD pipelines and DevOps practices
Read and debug Java code
Strong analytical skills
Tools
Splunk
Dynatrace
AppDynamics
Datadog
Prometheus
Grafana
Job description
Production support PYTHON SRE
Common Stack & Tools
Backend: Python
Python SRE Senior Support Engineer (59 Years)
Scope & Impact
Contribute to the production support and operational stability of assigned applications/domains.
Work closely with senior engineers and crossfunctional teams to ensure timely incident resolution and smooth daytoday operations.
Provide technical inferences around system behavior, performance, and recurring issues.
Support initiatives to maintain high availability, reliability, and resilience across customerfacing and critical applications.
Core Responsibilities
Ensure stable production operations through active monitoring, alert handling, and proactive issue detection.
Manage the incident lifecycle (P2–P5) including triage, analysis, coordination, and resolution under guidance for P1 issues.
Assist in root cause analysis, resolution of recurring issues, and implementation of preventive measures.
Maintain and update runbooks, SOPs, playbooks, and knowledge base to support consistent operations.
Support and enhance the observability setup—dashboards, alerts, logs, metrics.
Collaborate with DevOps/SRE teams on automation, self-heal scripts, and reduction of manual tasks.
Provide operational support across L1/L2 activities and contribute to L3 investigation under guidance.
Participate in oncall rotations and adhere to team schedules and workload priorities.
Communicate effectively with business and engineering teams on incident updates, impact, and resolutions.
Ensure compliance with change management, audit needs, and security standards.
Track and support improvements in key operational metrics such as SLA adherence, MTTA, and MTTR.
MustHave Skills
Strong foundation in incident, problem, and change management processes (ITIL v4 exposure preferred).
Handson experience with observability tools such as Splunk, Dynatrace, AppDynamics, Datadog, Prometheus, or Grafana.
Ability to troubleshoot across infrastructure layers—Linux/Windows OS, networking basics, and cloud fundamentals (AWS/Azure/GCP/OpenShift).
Familiarity with CI/CD pipelines, build/release processes, and DevOps practices.
Ability to read and debug Java code for issue diagnosis.
Strong analytical skills and ability to work collaboratively within support teams.
NicetoHave
Basic knowledge of Python and automation scripting.
AI & Responsible AI Expectations
Basic understanding of prompt engineering and ability to leverage tools such as Copilot for productivity.