Job Description
Experience - 4-6 years
Role Summary
Mid-level engineer to help build, operate, and continuously improve our monitoring and observability platform. You'll ensure reliability and performance by creating actionable alerting, clear dashboards, and resilient telemetry pipelines across infrastructure and applications. Stakeholder management and maintain customer satisfaction on Monitoring platform.
Key Responsibilities
- Design and maintain monitoring for services, infrastructure, and cloud resources (metrics, logs, traces).
- Build and tune alerting rules to reduce noise and improve actionable signal quality.
- Create dashboards and service health views for engineering and operations stakeholders.
- Support incident response by improving detection, triage, and post-incident learnings (RCA).
- Implement SLOs/SLIs and error budgets; track reliability/performance trends.
- Automate monitoring workflows (onboarding new services, alert routing, report generation).
- Partner with app teams to instrument code and standardize telemetry (OpenTelemetry where applicable).
- Maintain and optimize observability tooling, capacity, and cost (data retention, sampling, indexing).
- Document standards, runbooks, and operational procedures; mentor junior team members.
Required Qualifications
- 46 years of experience in monitoring/observability, SRE, platform engineering, or production operations.
- Strong fundamentals in Cloud Platforms, networking, and troubleshooting distributed systems.
- Hands‑on experience with at least one monitoring stack (Datadog, Dynatrace).
- Familiarity with tracing and instrumentation concepts (OpenTelemetry, Jaeger, Zipkin).
- Scripting/automation skills (Python, Go, or Bash) and comfort with APIs.
- Experience with on‑call/incident management and writing postmortems.
- Clear communication skills and ability to work cross‑functionally.
Preferred Qualifications
- Cloud experience (AWS/Azure/GCP) and managed monitoring services.
- Container/Kubernetes monitoring (Kubernetes, Helm, service meshes).
- Infrastructure as Code (Terraform/CloudFormation) and CI/CD (GitHub Actions, Jenkins, GitLab).
- Experience defining SLOs and implementing alerting based on user impact.
- Knowledge of IT service management tooling (ServiceNow) and alert routing.
Soft Skills & Ways of Working
- Bias for automation and measurable improvements (alert quality, MTTR, availability).
- Strong ownership mindset for production reliability.
- Pragmatic approach to standards to improve consistency without blocking delivery.
Job Classification
Industry: Recruitment / Staffing
Functional Area / Department: IT & Information Security
Role Category: IT Infrastructure Services
Role: IT Operations Management
Employment Type: Full time
Contact Details
Company: Cirruslabs
Location(s): Hyderabad
Ability to work in a team environment with members of varying skill levels. Highly motivated. Learns quickly.