Lead Platform Engineer (Observability Focus)
About The Role
We are looking for a Senior Platform Engineer with deep expertise in observability, cloud-native infrastructure, and large-scale distributed systems. This role is highly hands‑on and focuses on designing, building, and operating reliable, observable, and scalable platforms running on Kubernetes, with a strong preference for AWS. Knowledge of GCP is an edge.
Key Responsibilities
Reliability & Operations
- Design, implement, and maintain highly available and resilient systems in Kubernetes‑based environments.
- Define and enforce SLOs, SLIs, and error budgets.
- Lead incident response, RCA, and postmortems.
- Drive reliability improvements through automation.
Observability (Core Focus)
- Architect and operate observability platforms for metrics, logging, tracing, and alerting.
- Work with Prometheus, Alertmanager, Grafana, Splunk, Cribl, Datadog.
- Establish actionable alerting standards.
Cloud & Platform Engineering
- Build and manage infrastructure on AWS.
- Operate Kubernetes clusters (EKS preferred).
- Deploy services using Helm, ArgoCD and Argo rollout.
- Manage containerized workloads using Docker and containerd.
Automation & Tooling
- Strong Python skills with emphasis on reliability, automation, and observability tooling.
- Develop automation and tooling using Python.
- Create internal reliability and monitoring tools.
- Integrate CI/CD pipelines with observability and reliability checks.
Collaboration & Leadership
- Mentor junior engineers.
- Influence architecture decisions.
- Collaborate across engineering teams.
Required Qualifications
- 6+ years of relevant experience in SRE, DevOps, or Platform Engineering.
- Strong Python skills with experience building production‑grade automation and tooling.
- Strong programming experience in Python.
- Production experience with Kubernetes.
- Strong observability fundamentals.
- Experience with Helm, ArgoCD, Argo Rollout and Docker.
- Experience with AWS cloud.
- Strong Linux and networking fundamentals.
- Familiarity with the SDLC.
Preferred Qualifications
- Experience with OpenTelemetry and Observability tools.
- Experience with Kubernetes package manager (Helm) and deployment (ArgoCD / Argo Rollout).
- Multi‑cluster or multi‑region Kubernetes experience.
- Service mesh (Istio) and API Gateway (Kong) experience.
- Infrastructure‑as‑Code (Terraform preferred).
- Cloud cost optimization experience.
Technology Stack
- Programming & Automation: Python (strong proficiency, production‑grade tooling and automation).
- Containerization & Orchestration: Docker, AWS EKS.
- Packaging & Deployment: Helm, ArgoCD, Argo Rollout.
- Observability & Monitoring: Prometheus, Alertmanager, Grafana, OpenTelemetry, Datadog, Splunk, Cribl, Edge Collectors, AWS Cloud Watch.
- Platforms: AWS (primary and preferred), Google Cloud Platform (good to have).
- CI/CD & DevOps: Git‑based CI/CD pipelines, release automation, reliability checks.
- Infrastructure as Code: Terraform (preferred).
- Operating Systems & Networking: Linux, TCP/IP, DNS, load balancing.
Project Details / What You’ll Work On
- Build and operate a centralized observability platform for metrics, logs, traces, and alerting across Kubernetes workloads using the mentioned tooling for services running in AWS, on‑prem, and GCP (good to have).
- Contribute to the o11y center‑of‑excellence guiding teams to create their SLOs, SLIs, reduce MTTR and collaborate with developers to move toward an Observability‑Driven Development mindset.
- Support observability for services running on our Kubernetes ecosystem.
- Develop Python‑based automation and tooling for observability, SLO reporting, incident response, and operational workflows.
- Lead incident response for production issues, conduct blameless postmortems, and drive long‑term reliability improvements.
- Optimize platform scalability, performance, and cloud cost efficiency with a strong focus on GCP and AWS.
- Act as a technical leader, influencing architecture and mentoring teams on reliability and observability best practices.
(ref:hirist.tech)