L2 Support Engineer

Keka Inc.

Gurugram District

On-site

INR 1,500,000 - 4,500,000

Full time

48 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Keka Inc. is seeking a Site Reliability Engineer/DevOps professional in Gurugram to provide expert second-line support, collaborate with development for escalations, and drive proactive health monitoring.

You will work with product and delivery teams on release readiness and contribute to continuous improvement of infrastructure reliability and performance. The role emphasizes applying AI/ML techniques to detect anomalies, optimize logs and metrics, and ensure timely incident response and

Qualifications

  • 5+ years of direct experience in SRE/DevOps roles with high availability
  • ITIL fundamentals knowledge (Incident/Problem/Change)
  • Strong SQL/PostgreSQL and data-driven monitoring
  • Proficient in Windows and Linux environments, scripting (Python) preferred

Responsibilities

  • Provide second line client-facing technical support for escalated issues.
  • Collaborate with development for third line escalation and fixes.
  • Coordinate with product and delivery teams for readiness of releases and new enhancements.
  • Lead proactive health monitoring, reporting, and optimization initiatives.
  • Apply AI/ML techniques to detect anomalies and improve support operations.
  • Own incidents end-to-end, including 3rd line/change activity when needed.
  • Maintain knowledge base and share learnings with the team.
  • Ensure incidents comply with SLAs and OLA targets.

Skills

SRE/DevOps experience
AWS expertise
Incident management
Monitoring & Observability
Python scripting

Education

Bachelor’s degree in CS/Engineering

Tools

Splunk
CloudWatch
Grafana
Prometheus
Dotcom/Monolith monitoring

Job description

Provide second line client-facing technical support for issues escalated by first line support teams.

Apply strong technical skills and good business knowledge together with investigative techniques and problem-solving skills to identify and resolve issues efficiently and in a timely manner.

Work collaboratively with development team required for third line escalation.

Coordinate with product and delivery teams to ensure the Service Management team is ready for new releases and engaged in early design of new enhancements.

Work on initiatives and continuous improvement process around proactive application health monitoring, reporting, and technical support.

Apply AI/ML techniques to detect anomaly, predict alerting, and to enhance support operations.

Key Areas of The Teams Responsibilities Are
  • Proactive monitoring and management of business critical 24x7 real-time. Where required to rectify issues in a timely fashion to restore application functionality.
  • Ensure incidents are correctly processed, assessing business and technical impact and severity.
  • Taking ownership of application incidents and ensuring that they are resolved, this includes retaining ownership of incidents that require 3rd Line or IT Change activity to resolve.
  • Ensuring the communication to the business community remains active.
  • Application responsibilities will cover Application Infrastructure, Data Fixes, User Queries, User Education and Incident Investigation.
  • Monitoring of application events alerts, job schedules, capacity monitors and performance KPI's. Creation and ownership of change requests raised to address any of the above issues.
  • Proactively share knowledge with the team and update the knowledge base with support documentation (Confluence).
  • Work to provide services to agreed Service Level Targets and Operating Level Agreements.
  • Leverage AI Ops techniques to analyse logs, metrics, traces, and event data, enabling proactive trend identification and continuous optimization of system performance
Education and Hand on experience required.
  • 5+ years of direct experience in Site Reliability Engineering or DevOps roles, high availability, and incident response in AWS.
  • Strong knowledge on ITIL fundamentals (Incident management, problem management, change management).
  • Good understanding of Application Support processes
  • AWS: Understanding of AWS services flow (VPC, Cloud watch, RC2, EKS, should be able to debug issues)
  • Ideally familiar with monitoring tools such as Splunk, Cloudwatch, Dotcom and Monolith.
  • Expertise in SQL/PostgreSQL: Proficiency in advanced SQL techniques, query optimization, and experience with complex database systems.
  • Experience with advanced observability tools (e.g., Prometheus, Grafana, Splunk) for monitoring, logging, and tracing.
  • Experience in leading post-mortem analyses and implementing preventative measures to avoid recurrence of incidents.
  • Excellent problem-solving skills and the capacity to lead effectively under pressure during incident response and outage management.
  • Must understand operating systems most especially Windows and Linux. Good scripting experience (preferably including python) an advantage.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Support Engineer 2
Support Engineer 2

Acesoft Labs • Dadri

Hybrid
INR 600,000 - 1,200,000
Production Support (L2 Support)
Production Support (L2 Support)

TechBlocks • Dadri

Hybrid
INR 1,200,000 - 1,800,000
L2 Support System Administrator
L2 Support System Administrator

Acesoft Labs • Dadri

On-site
INR 1,200,000 - 1,800,000
Hybrid work model
Production Support Engineer
Production Support Engineer

Moofwd • Pune District

On-site
INR 600,000 - 1,000,000
Production support engineer
Production support engineer

Devon Software Services • Bengaluru

Hybrid
INR 600,000 - 1,000,000
Production Support / Application Support
Production Support / Application Support

Cloudxtreme • Pune District

On-site
INR 1,200,000 - 1,800,000
L3 Application Support
L3 Application Support

Altimetrik • Bengaluru

Hybrid
INR 850,000 - 1,350,000
Support Analyst L2
Support Analyst L2

PMC Commerce • Vadodara

On-site
INR 450,000 - 650,000
Technical Support Engineer - Data & Cloud Platforms
Technical Support Engineer - Data & Cloud Platforms

Luxoft • Gurugram District

On-site
INR 900,000 - 1,500,000
Application Support Engineer,production support, lead
Application Support Engineer,production support, lead

DMart • Coimbatore District, Bengaluru

On-site
INR 2,800,000 - 3,600,000