Overview
We are seeking an experienced Senior Cloud and DevOps Engineer to design, implement, and manage scalable cloud infrastructure and DevOps platforms across multi-cloud environments. The ideal candidate will have strong expertise in Kubernetes, Terraform, CI/CD automation, cloud-native technologies, observability, cost optimization, and platform reliability. The role also involves mentoring engineers, driving cloud transformation initiatives, and ensuring production stability.
Cloud Infrastructure and Platform Engineering
- Design, build, and manage cloud infrastructure on GCP, Azure, and AWS.
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Manage networking, IAM, security controls, storage, databases, and serverless services.
- Implement cloud governance, security, and compliance best practices.
- Optimize cloud spending through resource right-sizing, Spot instances, and infrastructure improvements.
Kubernetes And Container Platform
- Design and manage Kubernetes clusters in production environments.
- Implement and maintain service mesh technologies such as Istio.
- Develop custom Kubernetes controllers, webhooks, and platform automation solutions.
- Ensure high availability, scalability, and reliability of containerized workloads.
- Manage backup, disaster recovery, and cluster lifecycle operations.
DevOps And CI/CD
- Build and optimize CI/CD pipelines using GitLab, GitHub Actions, Helm, and ArgoCD.
- Implement GitOps practices for application deployment and infrastructure management.
- Automate provisioning, deployment, and operational workflows.
- Improve developer productivity through platform automation.
Observability And SRE
- Implement monitoring, logging, alerting, and tracing solutions using Datadog, Grafana, Prometheus, and Loki.
- Lead production incident management, root cause analysis, and post-incident reviews.
- Define and implement reliability standards, SLIs, SLOs, and operational excellence practices.
- Drive proactive monitoring and automation initiatives.
AI And Platform Innovation
- Evaluate and implement AI-driven operational solutions.
- Build intelligent observability and incident management workflows.
- Leverage AI agents to reduce operational overhead and improve platform efficiency.
Leadership And Collaboration
- Mentor and guide DevOps engineers.
- Collaborate with development, architecture, security, and product teams.
- Lead technical discussions and cloud modernization initiatives.
- Drive adoption of engineering best practices across teams.
Requirements
- Bachelor's Degree in Engineering, Computer Science, Information Technology, or a related field.
- Cloud Platforms: Google Cloud Platform (GCP), Microsoft Azure, Amazon Web Services (AWS).
- Infrastructure as Code: Terraform.
- Containerization and Orchestration: Kubernetes, Docker, and Istio Service Mesh.
- CI/CD and GitOps: GitLab CI/CD, GitHub Actions, ArgoCD, Helm.
- Observability: Datadog, Grafana, Prometheus, Loki.
- Programming and Scripting: Bash/Shell Scripting.
Preferred Qualifications
- HashiCorp Terraform Associate Certification.
- Certified Kubernetes Application Developer (CKAD).
- Azure AZ-900 Certification.
- Experience with Cloudflare, Spot.io, Velero, PostgreSQL, and AI-driven platform engineering.
- Experience leading technical teams and mentoring engineers.
This job was posted by Priyanka R N from Falabella.