Educational Requirements
Bachelor of Engineering, Bachelor Of Science
Service Line
Cloud & Infrastructure Services
Responsibilities
- Deploy, configure, and maintain the core observability stack using Prometheus, Grafana, Alertmanager, and Loki.
- Dashboarding & Visualization: Collaborate with engineering teams to design and build comprehensive Grafana dashboards for application and infrastructure health monitoring.
- Alerting Strategy: Configure and fine-tune Alertmanager rules to ensure accurate, actionable alerts while minimizing alert fatigue.
- Log Management: Architect and manage centralized logging solutions using Loki to ensure efficient log aggregation and querying.
- System Optimization: Monitor the performance of the observability stack itself, optimizing resource usage and scaling infrastructure as needed.
- Continuous Improvement: Evaluate and integrate modern observability tools (such as VictoriaMetrics and VictoriaLogs) to enhance system correlation and analysis capabilities.
- Experience: 9+ years of hands‑on experience in DevOps, SRE, or Observability roles.
- Core Stack: Deep technical expertise in Prometheus, Grafana, Alertmanager, and Loki (PLG Stack).
- Automation: Proficiency in scripting (Python, Bash) and Infrastructure as Code (e.g., Terraform, Ansible).
- Infrastructure & OS: Strong working knowledge of Linux/Unix administration.
- Containerization: Experience monitoring containerized environments (Docker, Kubernetes).
Additional Responsibilities
Location of posting: Chennai, Bangalore, Hyderabad, Pune
Technical and Professional Requirements
Hands‑on knowledge or prior exposure to VictoriaMetrics (for scalable time‑series data) and VictoriaLogs. Experience with distributed tracing tools (e.g., Jaeger, Tempo, or VictoriaTraces). Understanding of service level indicators (SLIs) and service level objectives (SLOs).
Preferred Skills
- Technology: Server‑Virtualization – OpenStack.