Role & responsibilities
We are looking for an experienced L2/L3 Engineer responsible for troubleshooting production issues, performing root-cause analysis, resolving incidents, and ensuring the stability and availability of critical applications and infrastructure.
The engineer will work closely with L1 support, development, DevOps, infrastructure, and business teams to resolve complex technical issues and drive continuous improvement.
Key Responsibilities:
- Provide L2/L3 technical support for production applications and infrastructure.
- Monitor systems, investigate alerts, and respond to incidents within defined SLAs.
- Troubleshoot application, database, network, OS, and infrastructure issues.
- Perform log analysis and identify root causes of recurring incidents.
- Handle critical/P1/P2 incidents and participate in incident management and escalation.
- Perform RCA and create corrective/preventive action plans.
- Deploy application fixes, configuration changes, and patches following change-management processes.
- Work with development teams to troubleshoot application defects and performance issues.
- Monitor system health, availability, capacity, and performance.
- Automate repetitive operational and troubleshooting activities using scripting/tools.
- Maintain technical documentation, runbooks, knowledge articles, and operational procedures.
- Participate in on-call/shift support as required.
- Identify opportunities to improve system reliability, monitoring, automation, and operational efficiency.
Required Skills
- 57+ years of experience in production/application/infrastructure support.
- Strong troubleshooting and analytical skills.
- Good knowledge of Linux/Unix and/or Windows administration.
- Experience with SQL and relational databases.
- Strong understanding of application logs, APIs, HTTP/HTTPS, networking, and system performance.
- Experience with monitoring and logging tools such as Splunk, ELK, Grafana, Prometheus, Dynatrace, AppDynamics, or similar.
- Experience with ticketing and ITSM tools such as ServiceNow, Jira, or similar.
- Knowledge of incident, problem, and change management processes.
- Scripting experience in Shell, Python, PowerShell, or similar.
- Familiarity with Git, CI/CD, Docker/Kubernetes, and cloud platforms (AWS/Azure/GCP) is an advantage.
Preferred candidate profile
- SLA adherence and incident resolution time.
- Production availability and reliability.
- Reduction in recurring incidents.
- Quality and timeliness of RCA documentation.
- Automation and operational-efficiency improvements.
- Successful resolution of critical incidents.