We are seeking a Senior Cloud DevOps Engineer to design, implement, and support scalable, secure, and resilient cloud infrastructure. The ideal candidate will have strong hands‑on experience with AWS, Kubernetes, Terraform, GitOps, CI/CD, and cloud networking.
This role will focus on infrastructure automation, Kubernetes platform management, cloud reliability, and continuous improvement while working closely with application and engineering teams.
Key Responsibilities
- Design and manage AWS infrastructure using Terraform, Kubernetes, and infrastructure-as-code practices.
- Develop Terraform modules, manage remote state, and implement policy-as-code.
- Implement GitOps using tools such as Argo CD or Flux.
- Manage Kubernetes platforms, including EKS, cluster lifecycle, upgrades, scaling, and containerized workloads.
- Troubleshoot complex Kubernetes and cloud networking issues involving CNI, ingress controllers, service mesh, and load balancing.
- Build and maintain reliable CI/CD pipelines using Jenkins and related tools.
- Automate infrastructure provisioning and configuration using Ansible, Puppet, or Chef.
- Develop infrastructure-as-code and configuration-as-code across multiple environments.
- Monitor and optimize infrastructure for performance, availability, security, and cost.
- Work with monitoring and logging platforms such as Datadog, Prometheus, Grafana, or ELK.
- Manage AWS networking, including VPCs, security groups, IAM, DNS, VPN, and related services.
- Integrate third‑party tools, plugins, and internal automation scripts.
- Create technical documentation, runbooks, and architecture diagrams.
- Participate in on‑call rotations and improve incident response and reliability processes.
- Collaborate with application, engineering, security, and infrastructure teams.
- Mentor engineers and contribute to DevOps/SRE best practices.
Required Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Electrical Engineering, or equivalent experience.
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering.
- 5+ years of hands‑on experience with cloud platforms, primarily AWS.
- Advanced Linux administration and troubleshooting skills, including systemd, cgroups, namespaces, OS tuning, and production hardening.
- Strong Kubernetes experience beyond fundamentals, including CNI, networking, ingress, service mesh, and cluster management.
- Hands‑on experience with AWS EKS and Docker.
- Strong expertise in Terraform and infrastructure automation.
- Strong scripting skills with Python and Bash.
- Experience with Git, Jenkins, and artifact repositories.
- Strong understanding of DNS, DHCP, VPN, LDAP, VPCs, security groups, and IAM policies.
- Experience with monitoring/logging tools such as Datadog, Prometheus, Grafana, or ELK.
- Experience working in Agile environments using JIRA and Confluence.
Nice to Have
- Experience with Helm and Kubernetes Helm charts.
- Experience with Argo CD or Flux.
- Exposure to hybrid cloud and on-premises environments.
- Experience with HashiCorp Vault, AWS Secrets Manager, or other secrets‑management platforms.
- Experience mentoring engineers or leading DevOps/SRE initiatives.
- Experience with service mesh and advanced Kubernetes networking.