Role Overview
As part of the Platform Operations team, you will:
- Own Kubernetes cluster operations in production environments
- Lead incident response and shift execution
- Ensure high availability, reliability, and operational excellence
Key Responsibilities
Kubernetes Platform Administration
- Manage and support production Kubernetes clusters (cloud/hybrid)
- Perform:
- Cluster installation, upgrades, patching, scaling
- Configure:
- Namespaces, RBAC, network policies, resource quotas
- Ensure:
- High availability, performance, and resilience
- Troubleshoot:
- Cluster, node, and application-level issues
Containers & Runtime Technologies
- Work with:
- Troubleshoot:
- Container lifecycle issues and runtime failures
- Enforce:
- Container security best practices
CNCF Ecosystem & Service Mesh
- Hands‑on with:
- Helm, operators, GitOps tools
- Monitoring, logging, and security tools
- Strong experience with Istio:
- Traffic management
- mTLS, observability, policy enforcement
- Implement service mesh architectures
Programming & Automation
- Use Go (Golang) for:
- Automation
- Tooling
- Kubernetes component
- Build automation for:
- Operational workflows
- Repetitive tasks
- Troubleshoot Go‑based services/controllers
Monitoring, Observability & Reliability
- Implement observability solutions:
- Proactively detect and resolve issues
- Conduct Root Cause Analysis (RCA)
- Drive reliability and performance improvements
Production Support & ITSM
- Provide L2/L3 support for Kubernetes platforms
- Operate within ITSM processes:
- Incident, Problem, Change Management
- Ensure:
- Maintain:
- SOPs, runbooks, documentation
Shift Lead Responsibilities
- Lead UK shift operations (3 PM – 1 AM IST)
- Coordinate team activities and workload
- Own:
- Major incident management
- Escalation and communication
- Ensure:
- Smooth shift handovers
- Operational discipline and compliance
- Act as primary escalation point during shift
Ideal Candidate Profile
- Strong hands‑on Kubernetes expert with operations mindset
- Proven ability to handle production incidents and lead shifts
- Calm, structured, and decisive under pressure
- Strong documentation and communication skills
- Comfortable working UK shift and global team environments
Why Join Us?
- Work on enterprise‑scale Kubernetes and cloud‑native platforms
- Exposure to advanced CNCF ecosystem & service mesh technologies
- Leadership opportunity through shift ownership
- Career growth into:
- Platform Engineering
- Site Reliability Engineering (SRE)
- Infrastructure Leadership
- Stable role based in Trichy supporting global systems