The Site Reliability Engineer (Observability) to join its Platform Engineering organization and help develop, implement, and mature a robust enterprise observability platform. This is a highly technical, hands‑on role focused on improving visibility across complex infrastructure, applications, Kubernetes environments, cloud platforms, and Windows/Linux systems. The engineer will assess the current environment, identify observability gaps, recommend solutions, and implement the tooling, instrumentation, telemetry, dashboards, and automation required to improve system reliability and operational visibility. The ideal candidate combines strong SRE/DevOps engineering experience with deep observability expertise, particularly with Grafana and OpenTelemetry. Candidates must be comfortable getting into technical details, working directly with infrastructure and application teams, configuring monitoring agents, scripting solutions, and translating technical findings into actionable recommendations.
About the Role
The Site Reliability Engineer (Observability) will design, build, and maintain enterprise observability solutions, working closely with infrastructure and application teams to improve system reliability, performance, and operational visibility across diverse environments.
Responsibilities
- Design, build, implement, and maintain scalable enterprise observability solutions across applications, infrastructure, cloud, and network environments.
- Assess the existing technology environment to identify gaps in monitoring, telemetry, instrumentation, and operational visibility.
- Develop and implement solutions that improve visibility across Kubernetes clusters, microservices, AWS/cloud infrastructure, Windows/Linux systems, VMware/vSphere, and network environments.
- Implement and support observability technologies including Grafana, OpenTelemetry, Prometheus, Loki, InfluxDB, and related monitoring and telemetry platforms.
- Build and maintain telemetry pipelines to aggregate metrics, logs, and distributed traces and provide end‑to‑end system visibility.
- Install, configure, troubleshoot, and maintain observability and monitoring agents across Windows and Linux environments.
- Develop dashboards, visualizations, KPIs, and intelligent alerting strategies that improve system uptime, reliability, performance, and troubleshooting.
- Work hands‑on with Kubernetes and DevOps technologies, including troubleshooting complex infrastructure and application issues.
- Develop scripts and automation using Python, Bash, Go, or similar technologies to improve observability, monitoring, remediation, and operational efficiency.
- Support incident response, troubleshooting, root cause analysis, and automated remediation workflows.
- Partner with Platform Engineering, Application Development, Network Engineering, and Support teams to establish consistent observability standards across environments.
- Analyze system health and performance data and communicate risks, trends, and recommendations to technical and non‑technical stakeholders.
- Continuously improve instrumentation, telemetry collection, alerting strategies, monitoring standards, and observability best practices.
- Evaluate observability technologies and vendors, manage vendor relationships, and assess solutions against technical and business requirements.
- Translate vendor capabilities and technical evaluations into clear recommendations and implementation strategies for internal engineering teams.
- Develop documentation, operational standards, and reporting to promote transparency and shared accountability across engineering teams.
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience.
Required Skills
- 5-7 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Observability Engineering, or a closely related discipline.
- Strong hands‑on experience with DevOps engineering and Kubernetes, with the ability to troubleshoot, code, configure, and work deeply within technical environments.
- Strong knowledge of modern observability principles, including metrics, logs, traces, telemetry collection, and telemetry streaming.
- Hands‑on experience with Grafana and related observability technologies.
- Experience implementing or supporting OpenTelemetry, including telemetry collection and instrumentation.
- Experience installing, configuring, managing, and troubleshooting monitoring or telemetry agents on both Windows and Linux systems.
- Cloud infrastructure experience, with AWS strongly preferred.
- Strong scripting and automation capabilities, preferably using Python; Bash, PowerShell, or Go experience is also valuable.
- Experience monitoring and troubleshooting distributed systems, infrastructure, applications, or microservices.
- Ability to identify gaps in observability coverage and independently recommend and implement technical solutions.
- Experience working with technology vendors, evaluating solutions, managing vendor relationships, and communicating recommendations to internal stakeholders.
- Strong problem‑solving and root cause analysis capabilities.
- Ability to communicate complex technical findings clearly to engineering teams, leadership, and business stakeholders.
- Strong cross‑functional collaboration skills with the ability to work across Platform Engineering, Application Development, Network Engineering, Infrastructure, and Support teams.
Preferred Skills
- Experience with VMware and/or VMware Cloud.
- Experience with Prometheus, Loki, InfluxDB, or similar observability technologies.
- Experience with Git and CI/CD pipelines, including GitHub Actions.
- Experience supporting hybrid environments spanning cloud and on‑premises infrastructure.
- Relevant certifications such as Certified Kubernetes Administrator (CKA), AWS certifications, VMware VCP, or similar credentials.