SRE

Sidglobal

Hyderabad

On-site

INR 600,000 - 900,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SID Global Solutions is seeking a proactive Site Reliability Engineer (SRE) / L1 Monitoring Engineer to join our core digital infrastructure operations team. You will monitor production systems, triage alerts, and ensure high availability for critical banking-related services.

Ideal candidates have 1–3 years in L1/monitoring roles, hands-on with Datadog, Dynatrace, Prometheus, and Grafana, and strong Linux/Kubernetes skills. This role requires a 24/7 shift schedule and on-site work in Hyderabad.

Qualifications

  • 1–3 years in an L1 Support, Infrastructure Monitoring, or junior SRE role.
  • Hands-on with Datadog, Dynatrace, Prometheus and Grafana for alerting and dashboards.
  • Practical understanding of Nginx and logs analysis.
  • Foundational Kubernetes knowledge and kubectl usage.
  • Familiarity with Google Cloud Platform core services.
  • Experience with API gateways, preferably Apigee.
  • Linux/Unix command-line proficiency and scripting basics.
  • Willingness to work a 24/7 rotating shift, including nights and weekends.

Responsibilities

  • Monitor production environments in real time and spot anomalies using observability stacks.
  • Acknowledge, triage and categorize alerts per SOPs and document incidents clearly.
  • Escalate unresolved P1/P2 issues to L2 or DevOps with logs and context.
  • Manage API traffic performance and monitor API gateway health with Apigee.
  • Support Nginx operations and troubleshoot routing or TLS issues.
  • Monitor Kubernetes cluster health and view pod logs and resource usage.

Skills

Observability tools
Datadog
Dynatrace
Prometheus
Grafana
Nginx
Kubernetes
kubectl
Google Cloud Platform
Apigee
Linux/Unix
24/7 shift readiness

Tools

Apigee
Kubernetes tooling
Datadog
Dynatrace
Prometheus
Grafana

Job description

Site Reliability Engineer (SRE) / L1 Monitoring Engineer
Job Summary

We are seeking a proactive and technically driven SRE / L1 Monitoring Engineer with 1 to 3 years of experience to join our core digital infrastructure operations team. In this role, you will serve as the first line of defense ensuring the high availability, security, and performance of critical financial services and digital banking applications. You will be responsible for real-time system monitoring, tracking alerts across modern observability stacks, performing initial triage on infrastructure bottlenecks, and managing API traffic performance. This is an excellent opportunity for an early-career engineer looking to scale their skills in a high-volume, secure cloud infrastructure environment.

Key Responsibilities
L1 Infrastructure Monitoring & Alerts
  • Real-time Surveillance: Actively monitor production environments, enterprise dashboards, and telemetry feeds using toolsets like Datadog, Dynatrace, and Grafana to spot anomalies before they impact end-users. [1, 2, 3]
  • Alert Triage: Acknowledge, validate, and categorize incoming infrastructure, database, and application alerts generated by Prometheus and application performance monitoring (APM) agents using predefined Standard Operating Procedures (SOPs). [1, 2, 3, 4, 5]
  • Incident Escalation: Document incident details clearly in the ticketing system and swiftly elevate unresolved P1/P2 issues to L2 engineers or specialized DevOps teams with complete log snippets and context.
Application Delivery & API Traffic Management
  • Nginx Operations: Monitor web server logs, verify reverse proxy configurations, and troubleshoot basic traffic routing or SSL/TLS certificate errors. [1, 2, 3]
  • API Gateways: Use Apigee to monitor API proxy performance, track error rates (5xx/4xx codes), track latency spikes, and check developer portal connectivity. [1, 2, 3, 4]
  • Kubernetes Support: Monitor cluster health, inspect pod statuses, view application logs using kubectl, and track resource usage (CPU/Memory limits). [1, 2]
Cloud Operations & Reliability
  • GCP Monitoring: Utilize Google Cloud logging, native monitoring tools, and integrated observability dashboards to check the health of virtual machines, storage, and networking layers. [1, 2, 3, 4]
  • Health Checks: Perform routine daily morning sanity checks and post-deployment validation steps for critical banking services.
  • Runbook Execution: Execute automated or manual scripts to restart failed services, clear disk space, or cycle pods safely in staging and production environments.
Required Qualifications & Technical Skills
  • Experience: 1 to 3 years of hands-on experience in an L1 Support, Infrastructure Monitoring, or Junior SRE role.
  • Observability Tools: Hands-on experience navigating and tracking alerts within Datadog, Dynatrace, Prometheus, and Grafana.
  • Web Servers: Practical understanding of Nginx (reverse proxy, load balancing, log analysis).
  • Containerization: Foundational knowledge of Kubernetes (K8s) (understanding pods, deployments, services, and basic troubleshooting commands like kubectl logs and kubectl get pods).
  • Cloud Platform: Familiarity with Google Cloud Platform (GCP) core services and cloud monitoring concepts.
  • API Management: Exposure to Apigee or equivalent API gateways for monitoring traffic flow and checking endpoint health.
  • Operating Systems: Strong command-line comfort in Linux/Unix environments for navigating directories and tailing logs.
  • Shift Flexibility: Readiness to work in a 24/7 rotating shift model (including night shifts and weekends) to maintain uninterrupted banking infrastructure support.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE
SRE

SID Global Solutions • Mumbai

On-site
INR 1,200,000 - 1,800,000
SRE
SRE

United States Digital Space LLC • Maharashtra

On-site
INR 900,000 - 1,500,000
SRE
SRE

Sidglobal • Mumbai

On-site
INR 600,000 - 1,200,000
SRE
SRE

Sidglobal • Airoli

On-site
INR 600,000 - 900,000
Associate DevOps Engineer & Enterprise Architect Site Reliability Engineering
Associate DevOps Engineer & Enterprise Architect Site Reliability Engineering

300005 Chief Executive's Office_00002555 • India

On-site
INR 800,000 - 1,200,000
SRE Lead
SRE Lead

Hdfc Securities • Mumbai

On-site
INR 3,500,000 - 5,500,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Sr. Site Reliability Engineer I
Sr. Site Reliability Engineer I

MetLife • Hyderabad

Hybrid
INR 1,500,000 - 2,300,000
Site Reliability Engineering (SRE) Lead
Site Reliability Engineering (SRE) Lead

SID Global Solutions • Hyderabad

On-site
INR 3,000,000 - 5,000,000