Lead / Senior Engineer – Build, Release & Deploy (K8s Operations & Observability)
6 - 8
Full-Time
About the Role
We are looking for a Lead / Senior Engineer to own our Build, Release & Deploy pipeline along with Kubernetes operations and observability.
This role is central to how reliably and quickly we ship software — from CI/CD architecture to production-grade Kubernetes operations and end-to-end observability. You'll be the go-to owner for release engineering and platform reliability, working closely with engineering teams to reduce friction and increase confidence in every deployment.
Key Responsibilities
- Own the end-to-end build, release, and deployment pipeline — CI/CD architecture, release strategies (blue-green, canary, rolling), and rollback processes.
- Lead Kubernetes operations — cluster architecture, scaling, upgrades, networking, and workload
- Design and own the observability stack — metrics, logging, tracing, and alerting (e.g.,
- Define and drive SLIs/SLOs and error-budget practices in partnership with engineering teams.
- Build self-service deployment tooling and golden paths so product teams can ship independently and safely.
- Drive incident management — on-call practices, runbooks, postmortems, and continuous reliability improvements.
- Own infrastructure-as-code practices for cluster and pipeline configuration (Terraform, Helm, Argo
- Mentor engineers on CI/CD, Kubernetes, and observability best practices; act as the technical
- escalation point for release and platform issues.
Required Skills & Experience
- 7+ years of experience in DevOps / SRE / Platform Engineering, with at least 2+ years in a lead / senior capacity.
- Deep, hands-on expertise with Kubernetes in production — cluster management, Helm, operators, networking, and troubleshooting.
- Strong experience building and owning CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI, Argo CD, or similar).
- Solid experience with observability tooling — Prometheus, Grafana, ELK/EFK, OpenTelemetry,
- Proficiency with Infrastructure as Code (Terraform) and configuration management.
- Strong understanding of cloud platforms (AWS/Azure/GCP) and containerization (Docker).
- Scripting/programming proficiency (Python, Go, or Bash) for automation and tooling.
- Experience owning production reliability — on-call, incident response, and postmortem culture.
Good to Have
- Relevant certifications — CKA/CKAD/CKS, or cloud provider DevOps/Architect certifications.
- Experience with GitOps workflows and service mesh (Istio/Linkerd).
- Experience setting up a platform engineering or SRE practice from scratch.
What We Offer
- Ownership of the platform that powers every release across the engineering org.
- A culture that invests in tooling, automation, and developer experience.
- Competitive compensation, benefits, and growth opportunities.