Job Title: Site Reliability Engineer (SRE) / L1Monitoring Engineer
Job Summary
We are seeking a proactive and technically driven SRE /L1 Monitoring Engineer with 1 to 3 years of experience to join our coredigital infrastructure operations team. In this role, you will serve as thefirst line of defense ensuring the high availability, security, and performanceof critical financial services and digital banking applications. You will beresponsible for real-time system monitoring, tracking alerts across modernobservability stacks, performing initial triage on infrastructure bottlenecks,and managing API traffic performance. This is an excellent opportunity for anearly-career engineer looking to scale their skills in a high-volume, securecloud infrastructure environment. [1 ]
Key Responsibilities
L1 Infrastructure Monitoring & Alerts
- Real-timeSurveillance: Actively monitor production environments, enterprisedashboards, and telemetry feeds using toolsets like Datadog, Dynatrace,and Grafana to spot anomalies before they impact end-users. [1 , 2 , 3 ]
- AlertTriage: Acknowledge, validate, and categorize incoming infrastructure,database, and application alerts generated by Prometheus andapplication performance monitoring (APM) agents using predefined StandardOperating Procedures (SOPs). [1 , 2 , 3 , 4 , 5 ]
- IncidentEscalation: Document incident details clearly in the ticketing systemand swiftly escalation unresolved P1/P2 issues to L2 engineers or specialized DevOps teams with complete log snippets and context.
Application Delivery & API Traffic Management
- NginxOperations: Monitor web server logs, verify reverse proxy configurations, and troubleshoot basic traffic routing or SSL/TLScertificate errors. [1 , 2 , 3 ]
- APIGateways: Use Apigee to monitor API proxy performance, track error rates (5xx/4xx codes), track latency spikes, and check developerportal connectivity. [1 , 2 , 3 , 4 ]
- KubernetesSupport: Monitor cluster health, inspect pod statuses, viewapplication logs using kubectl, and track resource usage (CPU/Memorylimits). [1 , 2 ]
Cloud Operations & Reliability
- GCPMonitoring: Utilize Google Cloud logging, native monitoring tools, and integrated observability dashboards to check the health of virtualmachines, storage, and networking layers. [1 , 2 , 3 , 4 ]
- HealthChecks: Perform routine daily morning sanity checks andpost-deployment validation steps for critical banking services.
- RunbookExecution: Execute automated or manual scripts to restart failed services, clear disk space, or cycle pods safely in staging and production environments.
Required Qualifications & Technical Skills
- Experience: 1 to 3 years of hands-on experience in an L1 Support, InfrastructureMonitoring, or Junior SRE role.
- ObservabilityTools: Hands-on experience navigating and tracking alerts within Datadog,Dynatrace, Prometheus, and Grafana.
- WebServers: Practical understanding of Nginx (reverse proxy, loadbalancing, log analysis).
- Containerization: Foundational knowledge of Kubernetes (K8s) (understanding pods,deployments, services, and basic troubleshooting commands like kubectllogs and kubectl get pods).
- CloudPlatform: Familiarity with Google Cloud Platform (GCP) core services and cloud monitoring concepts.
- APIManagement: Exposure to Apigee or equivalent API gateways for monitoring traffic flow and checking endpoint health.
- OperatingSystems: Strong command-line comfort in Linux/Unix environments for navigating directories and tailing logs.
- ShiftFlexibility: Readiness to work in a 24/7 rotating shift model (including night shifts and weekends) to maintain uninterrupted bankinginfrastructure support