Skill required: Kubernetes cluster administration, scaling, upgrades, RBAC, namespaces, Helm charts, ingress controllers, workload troubleshooting and optimization
Qualifications: Highly skilled and hands-on Senior DevOps / Platform Engineer
Years of Experience: 5+ Years
Job Description
We are looking for a highly skilled and hands-on Senior DevOps / Platform Engineer
to design, build, automate, and manage scalable cloud-native platforms and DevOps ecosystems.
The ideal candidate should have deep expertise in Kubernetes, CI/CD, Infrastructure as Code,
observability, automation, tools, cloud cost optimization, cloud networking, and platform engineering,
along with strong troubleshooting and SRE capabilities.
The role requires close collaboration with Development, Security, SRE, and Infrastructure teams
to improve platform reliability, deployment efficiency, security, and developer productivity.
Key Responsibilities
- Design, implement, and manage scalable Kubernetes-based platforms
- Build and maintain cloud-native infrastructure and DevOps tooling
- Develop and manage CI/CD pipelines for application and infrastructure deployments
- Implement GitOps-based deployment strategies and platform automation
- Manage observability platforms including monitoring, logging, tracing, and alerting
- Drive Infrastructure as Code (IaC) and automation initiatives
- Ensure platform security, compliance, reliability, and high availability
- Troubleshoot production incidents and perform root cause analysis
- Collaborate with engineering and security teams for release management and platform improvements
- Optimize cloud infrastructure cost, scalability, and performance
- Drive DevSecOps and SRE best practices across teams
- Enable developer productivity through automation and AI-assisted engineering tools
Required Technical Skills
- Helm charts, ingress controllers, workload troubleshooting, and optimization
- Harbor
- Grafana
- ELK Stack / EFK Stack
- Loki
- Thanos
- Fluentd / Fluent Bit
- Cert-Manager
- NGINX Ingress Controller
CI/CD & DevSecOps
- GitHub Actions
- Jenkins
- CI/CD
- Blue-Green and Canary deployments
- Artifact and container image management
- Terraform and Pulumi
- Cloud provisioning and automation
- Load balancing and auto scaling
- VPC, DNS, VPN, firewall concepts
- High Availability and Disaster Recovery
- Multi-cloud and hybrid cloud exposure
Containerization & Orchestration
- Docker and container lifecycle management
- Cluster monitoring and optimization
- Metrics, logging, and tracing implementations
- Prometheus, Grafana, ELK, Loki, and Thanos
- Alerting, incident management, RCA, and performance tuning
Security & Compliance
- Kubernetes security best practices
- Vault and Kubernetes Secrets
- IAM and access control
- Container security scanning and DevSecOps
Automation & Scripting
- API integrations and workflow automation
Troubleshooting & Reliability Engineering
- Production incident handling
- SRE practices, capacity planning, and optimization
AI-Driven Engineering & Developer Productivity
- Cursor
- AI-assisted automation and troubleshooting
Soft Skills
- Strong communication and collaboration skills
- Strong ownership and problem-solving mindset
- Experience working in Agile and DevOps culture
- Ability to collaborate with Development, Security, and SRE teams
Preferred Qualifications
- Kubernetes certifications (CKA / CKAD / CKS)
- Cloud certifications (Azure / GCP)
- Exposure to platform engineering and Internal Developer Platforms (IDP)
Nice to Have
- Experience with AIOps platforms
- Exposure to enterprise security and compliance frameworksExperience with large-scale production environments and global deployments