At Ensono, our Purpose is to be a relentless ally, disrupting the status quo and unleashing our clients to Do Great Things! We enable our clients to achieve key business outcomes that reshape how our world runs. As an expert technology adviser and managed service provider with cross-platform certifications, Ensono empowers our clients to keep up with continuous change and embrace innovation.
We can Do Great Things because we have great Associates. The Ensono Core Values unify our diverse talents and are woven into how we do business. These five traits are the key to achieving our purpose.
Job Responsibilities:
1. Enterprise Architecture & Strategy
- Architect and govern a unified observability framework covering metrics, logs, traces, and events using IBM Instana, Grafana, OpenTelemetry, Telegraf, and InfluxDB.
- Lead First-of-a-Kind (FOAK) implementations—evaluating new observability tech and converting them into secure, repeatable, production-ready patterns.
- Define enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management.
2. SRE & Service Reliability
- Define and govern Service Level Indicators (SLIs), Objectives (SLOs), and error budgets.
- Serve as the senior technical escalation point, leading major P1/P2 incident war rooms and conducting evidence-based Root Cause Analysis (RCA).
- Drastically reduce MTTD/MTTR and alert noise through event correlation, dynamic thresholds, and dependency mapping.
- Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps.
- Automate the deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors.
- Integrate observability platforms seamlessly with ITSM (ServiceNow), Netcool, and CI/CD pipelines.
- Design deep observability for Docker, Kubernetes, microservices, and multi-cloud environments (Azure/AWS/GCP).
- Correlate application APM telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies.
- Ensure secure-by-design telemetry pipelines (RBAC, TLS, secrets management, and image scanning).
5. Technical Leadership & Transition Management
- Lead complex Knowledge Transfer (KT) programs, vendor transitions, and operational readiness handovers for global 24x7 teams.
- Mentor cross-functional engineering teams and influence enterprise technology roadmaps.
Required Qualifications:
- Observability & APM: IBM Instana, Grafana (Enterprise & Alloy), Prometheus, OpenTelemetry, Telegraf, InfluxDB.
- Legacy/Traditional Monitoring: SolarWinds, Netcool, Elastic/Splunk.
- Infrastructure: Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.
- Automation & DevOps: Ansible, Terraform, Python, Bash, Webhooks, CI/CD (GitHub Actions/GitLab/Jenkins).
- ITSM/Operations: ServiceNow, ITIL 4, advanced Major Incident Management.
Qualifications & Experience
- 12+ years of total IT experience, with a minimum of 5 to 7 years functioning as a Lead Architect, SRE, or Principal Observability Engineer in a massive enterprise environment.
- Proven track record of migrating organizations from legacy monitoring to proactive, automated observability platforms.