We are looking for an experienced Consultant – Enterprise Observability Platform Engineer to design, develop, and govern an enterprise-wide observability platform. The ideal candidate will drive platform architecture, establish observability standards, enable developer self‑service, and support engineering resilience across cloud‑native, hybrid, and legacy environments.
Responsibilities
- Define and evolve the Enterprise Observability Platform architecture integrating logs, metrics, traces, events, and alerts
- Design reference architectures and observability patterns for cloud‑native and monolithic applications
- Evaluate and integrate observability tools including OpenTelemetry, Prometheus, Grafana, Elastic, Splunk, Datadog, and New Relic
- Define enterprise observability standards, SLOs, SLIs, instrumentation guidelines, and telemetry governance
- Build and maintain scalable, highly available, and cost‑optimized observability platform services
- Automate provisioning, onboarding, alert configuration, and tenant lifecycle management
- Enable self‑service capabilities through instrumentation kits, dashboards, alert templates, and troubleshooting guides
- Collaborate with Platform Engineering, Developer Experience, SRE, and Application teams to drive observability adoption
- Lead application onboarding, migration, and telemetry integration initiatives
- Define observability KPIs, telemetry baselines, and portfolio‑level monitoring strategies
Qualifications
- 5+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, or DevOps
- Strong understanding of Observability concepts – Logs, Metrics, Traces, Events, SLOs, SLIs, RED & USE models
- Hands‑on experience with Grafana, Prometheus, Elastic, Splunk, Datadog, OpenTelemetry, or New Relic
- Strong scripting and automation skills using Python, Go, Bash, or Terraform
- Experience with Kubernetes, Container Orchestration, AWS, and Azure
- Experience building or supporting Enterprise Observability Platforms
- Knowledge of Multi‑Tenant Observability Systems and Governance‑as‑Code
- Experience with CI/CD integration and developer enablement
- Strong troubleshooting, automation, and platform engineering mindset
- Excellent communication and stakeholder management skills