Join the ETP Growth Journey At ETP Group, we deliver the next generation of AI‑powered, cloud‑native SaaS platforms that are transforming retail and e‑Commerce operations across Asia Pacific. As we empower brands with the agility, intelligence, and innovation they need to grow in a dynamic market, we grow too. To do that we need the right people on board. We’re always on the lookout for passionate professionals who are smart, self‑motivated, and eager to make a real impact. If you love solving challenges, working with cutting‑edge technology, and being part of a collaborative and fast‑paced environment, ETP is the place for you. Here, you’ll find more than just a job. You’ll find the right opportunity to shape the future of unified commerce, whilst growing your career alongside a team that values innovation, ownership, and excellence. Ready to be part of our success story? Email your resume to careers@etpgroup.com — please include the position you’re applying for, a recent photograph, current and expected compensation, educational qualifications, work experience, and contact details.
Experience Required: 5+ years
Key Responsibilities
- Design, deploy, and manage Kubernetes clusters on AWS and GCP
- Set up and manage:
- EKS and/or GKE clusters
- Node groups and autoscaling policies
- Cluster networking and ingress controllers
- Implement namespace segregation and resource quotas
- Manage cluster upgrades, patching, and lifecycle management
- Provision infrastructure using Infrastructure‑as‑Code tools:
- Terraform (preferred)
- CloudFormation
- Automate cluster provisioning and environment setup
- Implement automated scaling using:
- Cluster Autoscaler
- Horizontal Pod Autoscaler
- Vertical Pod Autoscaler (preferred)
Hands‑on Experience Managing
- AWS:
- EC2, EKS, VPC, IAM, ELB (ALB/NLB), EBS, CloudWatch
- GCP:
- GKE, Compute Engine, VPC, IAM, Load Balancing
Additional Responsibilities
- VPCs, subnets, route tables
- Load balancers (ALB, NLB, GCP Load Balancer)
- Security groups and firewall rules
- IAM roles, service accounts, and access policies
Microservices Platform Operations
- Optimize resource utilization and cluster performance
- Troubleshoot pod, node, networking, and application issues
- Manage rolling deployments and zero‑downtime upgrades
Observability & Monitoring
- Implement and manage monitoring tools:
- Prometheus
- Grafana
- GCP Operations Suite
- Configure dashboards, alerts, and incident monitoring
- Perform root cause analysis and incident resolution
Security & Compliance
- Implement Kubernetes security best practices:
- RBAC
- Network policies
- Pod security standards
- Configure secrets management:
- AWS Secrets Manager
- GCP Secret Manager
- Ensure compliance with enterprise security standards
CI/CD & DevOps Integration
- Integrate Kubernetes with CI/CD pipelines:
- Jenkins
- GitLab CI / GitHub Actions
- Support container build and deployment workflows
- Work with Docker and Helm
Disaster Recovery & High Availability
- Design and implement HA Kubernetes architectures
- Implement backup and recovery strategies
- Participate in DR drills and recovery validation
Required Technical Skills
- 5+ years hands‑on Kubernetes experience
- Strong experience with EKS and/or GKE
- Experience managing production clusters
- Strong knowledge of:
- Pods, Deployments, StatefulSets
- Services, Ingress
- ConfigMaps, Secrets
Infrastructure as Code
- Terraform (preferred)
- CloudFormation
- Grafana
- ELK Stack or cloud‑native logging
Preferred Skills
- Experience with Service Mesh (Istio, Linkerd)
- Experience with Kafka, Redis, or Solr environments
- Multi‑environment setup (Dev, QA, Prod)
- Production incident handling (SRE practices)
Soft Skills
- Strong troubleshooting and analytical skills
- Ability to work in production‑critical environments
- Good communication and collaboration skills
- Flexible and adaptable mindset
Preferred Certifications
One or more of:
- AWS Certified Solutions Architect
- Google Professional Cloud DevOps Engineer
First 90 Days – Success Indicators
- Set up and manage Kubernetes environments
- Improve cluster reliability and scalability
- Implement monitoring and alerting frameworks
Ready to Transform Your Retail Enterprise?
Ready to Transform Your Retail Enterprise?