SRE

Sidglobal

Airoli

On-site

INR 600,000 - 900,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sidglobal invites an early-career Site Reliability Engineer to join our core digital infrastructure operations team in India. You will serve as the first line of defense for high availability, security, and performance of critical financial services applications.

This role focuses on real-time monitoring, triage, and incident escalation to maintain uninterrupted banking infrastructure. The ideal candidate has 1–3 years of hands-on experience with observability stacks (Datadog, Dynatrace,

Qualifications

  • 1 to 3 years of hands-on experience in an L1 support, infrastructure monitoring, or junior SRE role.
  • Hands-on experience with Datadog, Dynatrace, Prometheus, and Grafana.
  • Practical understanding of Nginx (reverse proxy, load balancing, log analysis).
  • Foundational knowledge of Kubernetes (pods, deployments, services; basic kubectl commands).
  • Familiarity with Google Cloud Platform (GCP) core services and monitoring concepts.
  • Exposure to Apigee or similar API gateways for monitoring traffic and endpoint health.
  • Strong Linux/Unix CLI skills for navigating logs and directories.
  • Willingness to work in a 24/7 rotating shift model including nights and weekends.

Responsibilities

  • Real-time surveillance of production environments, dashboards, and telemetry feeds using Datadog, Dynatrace, and Grafana to spot anomalies.
  • Acknowledge, triage, and categorize alerts from Prometheus and APM agents using SOPs.
  • Document incidents clearly and escalate unresolved P1/P2 issues to L2 or DevOps teams with logs.
  • Monitor API traffic and proxy performance with Apigee, track error rates and latency spikes.
  • Check cluster health, pod statuses, and logs via kubectl; monitor resource usage in Kubernetes.
  • Perform health checks and runbooks for critical banking services in staging and production.

Skills

1-3 years exp
Datadog
Dynatrace
Prometheus
Grafana
Nginx
Kubernetes (K8s)
Apigee
Linux/Unix
GCP
Shift rotation

Tools

Nginx
Kubernetes
Apigee

Job description

Job Title: Site Reliability Engineer (SRE) / L1Monitoring Engineer

Job Summary

We are seeking a proactive and technically driven SRE /L1 Monitoring Engineer with 1 to 3 years of experience to join our coredigital infrastructure operations team. In this role, you will serve as thefirst line of defense ensuring the high availability, security, and performanceof critical financial services and digital banking applications. You will beresponsible for real-time system monitoring, tracking alerts across modernobservability stacks, performing initial triage on infrastructure bottlenecks,and managing API traffic performance. This is an excellent opportunity for anearly-career engineer looking to scale their skills in a high-volume, securecloud infrastructure environment. [1 ]

Key Responsibilities
L1 Infrastructure Monitoring & Alerts
  • Real-timeSurveillance: Actively monitor production environments, enterprisedashboards, and telemetry feeds using toolsets like Datadog, Dynatrace,and Grafana to spot anomalies before they impact end-users. [1 , 2 , 3 ]
  • AlertTriage: Acknowledge, validate, and categorize incoming infrastructure,database, and application alerts generated by Prometheus andapplication performance monitoring (APM) agents using predefined StandardOperating Procedures (SOPs). [1 , 2 , 3 , 4 , 5 ]
  • IncidentEscalation: Document incident details clearly in the ticketing systemand swiftly escalation unresolved P1/P2 issues to L2 engineers or specialized DevOps teams with complete log snippets and context.
Application Delivery & API Traffic Management
  • NginxOperations: Monitor web server logs, verify reverse proxy configurations, and troubleshoot basic traffic routing or SSL/TLScertificate errors. [1 , 2 , 3 ]
  • APIGateways: Use Apigee to monitor API proxy performance, track error rates (5xx/4xx codes), track latency spikes, and check developerportal connectivity. [1 , 2 , 3 , 4 ]
  • KubernetesSupport: Monitor cluster health, inspect pod statuses, viewapplication logs using kubectl, and track resource usage (CPU/Memorylimits). [1 , 2 ]
Cloud Operations & Reliability
  • GCPMonitoring: Utilize Google Cloud logging, native monitoring tools, and integrated observability dashboards to check the health of virtualmachines, storage, and networking layers. [1 , 2 , 3 , 4 ]
  • HealthChecks: Perform routine daily morning sanity checks andpost-deployment validation steps for critical banking services.
  • RunbookExecution: Execute automated or manual scripts to restart failed services, clear disk space, or cycle pods safely in staging and production environments.
Required Qualifications & Technical Skills
  • Experience: 1 to 3 years of hands-on experience in an L1 Support, InfrastructureMonitoring, or Junior SRE role.
  • ObservabilityTools: Hands-on experience navigating and tracking alerts within Datadog,Dynatrace, Prometheus, and Grafana.
  • WebServers: Practical understanding of Nginx (reverse proxy, loadbalancing, log analysis).
  • Containerization: Foundational knowledge of Kubernetes (K8s) (understanding pods,deployments, services, and basic troubleshooting commands like kubectllogs and kubectl get pods).
  • CloudPlatform: Familiarity with Google Cloud Platform (GCP) core services and cloud monitoring concepts.
  • APIManagement: Exposure to Apigee or equivalent API gateways for monitoring traffic flow and checking endpoint health.
  • OperatingSystems: Strong command-line comfort in Linux/Unix environments for navigating directories and tailing logs.
  • ShiftFlexibility: Readiness to work in a 24/7 rotating shift model (including night shifts and weekends) to maintain uninterrupted bankinginfrastructure support
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE
SRE

Sidglobal • Mumbai

On-site
INR 600,000 - 1,200,000
SRE
SRE

United States Digital Space LLC • Maharashtra

On-site
INR 900,000 - 1,500,000
SRE
SRE

SID Global Solutions • Mumbai

On-site
INR 1,200,000 - 1,800,000
SRE
SRE

Sidglobal • Hyderabad

On-site
INR 600,000 - 900,000
Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Kameda Infologics Pvt. Ltd. • Thiruvananthapuram

On-site
INR 600,000 - 800,000
SRE -Site Reliability Engineer II
SRE -Site Reliability Engineer II

QUEST DIAGNOSTICS INC • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

FIS • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Competitive salary
Attractive range of benefits
Opportunity for skill growth
Site Reliability Engineer (SRE) – Core IT Infrastructure
Site Reliability Engineer (SRE) – Core IT Infrastructure

TECEZE • Chennai District

On-site
INR 1,000,000 - 2,000,000