Senior DevOps Engineer with Python
Location: Bengaluru
Job Code: 14607
Experience: 7–9 years
We are seeking a highly skilled Senior DevOps Engineer with expertise in Core DevOps, AWS, and Python/GoLang to join our team. The ideal candidate will design and implement scalable, reliable AWS infrastructure and collaborate closely with developers and SREs to ensure systems are resilient, scalable, and efficient.
Responsibilities
- Design and implement scalable, reliable AWS infrastructure.
- Develop and maintain automation tools and CI/CD pipelines using Jenkins or GitHub Actions.
- Build, operate, and maintain Kubernetes clusters for container orchestration.
- Leverage Infrastructure as Code tools like Terraform or CloudFormation for consistent environment provisioning.
- Automate system tasks using Python or Golang scripting.
- Collaborate with developers and SREs to ensure systems are resilient, scalable, and efficient.
- Monitor and troubleshoot system performance using observability tools such as Datadog, Prometheus, and Grafana.
- Architect, build, and operate highly available, scalable, and secure cloud infrastructure primarily on AWS.
- Design VPCs, IAM, compute, storage, and load balancing solutions following AWS best practices.
- Define and implement infrastructure scalability, high availability, and disaster recovery strategies.
- Support multi‑AZ and multi‑region architectures for production workloads.
- Design and maintain Terraform‑based infrastructure using modules, workspaces, and remote state backends.
- Integrate Terraform workflows into CI/CD pipelines with automated validation and provisioning.
- Drive infrastructure standardization, governance, and reusability.
- Design, deploy, and operate Kubernetes clusters (EKS and on‑premise).
- Own Kubernetes control‑plane and node lifecycle management, including upgrades.
- Manage EKS add‑ons such as VPC CNI, CoreDNS, kube‑proxy, metrics‑server, CSI drivers, and ALB Ingress Controller.
- Design and operate highly available Kubernetes clusters across multiple availability zones and multi‑cluster architectures.
- Package, deploy, and manage Kubernetes applications using Helm charts.
- Design and maintain reusable, versioned Helm charts with environment‑specific values and override strategies.
- Implement Pod Security Standards and enforce policies using OPA Gatekeeper or Kyverno.
- Design and maintain RBAC governance with least‑privilege access models.
- Secure the container supply chain with image scanning, SBOM generation, Cosign signing/verification, and registry governance.
- Implement GitOps workflows using ArgoCD or FluxCD for declarative deployments and environment promotion.
- Enable automated rollbacks, drift detection, and environment consistency.
- Develop automation, tooling, and platform services using Python or Golang.
- Write Shell scripts for infrastructure automation, cluster operations, and tooling integration.
- Perform advanced Linux administration and troubleshooting across compute, networking, and storage layers.
- Lead and participate in production incident management, including triage, mitigation, and root‑cause analysis.
- Diagnose complex failures across cloud infrastructure, Kubernetes, networking, and CI/CD systems.
Qualifications
- Bachelor’s degree in Computer Science, Information Technology, or a related field (B.Tech).
- 6–9 years of experience as a DevOps Engineer with a strong focus on AWS.
- Hands‑on experience with containerization tools such as Docker.
- Expertise in managing Kubernetes clusters in production.
- Experience creating Helm charts for applications.
- Proficiency in creating and managing CI/CD workflows with Jenkins or GitHub Actions.
- Strong background in Infrastructure as Code (Terraform or CloudFormation).
- Automation and scripting skills using Python or Golang.
- Strong analytical and problem‑solving abilities.
- Excellent communication and collaboration skills.
- Certifications in Kubernetes or Terraform are a plus.
Preferred Qualifications
- Configuration management using Ansible.
- Basic understanding of AI & ML concepts.
- Experience with observability and service mesh tools such as Prometheus, Grafana, Loki/ELK, Datadog, CloudWatch, Istio, or Envoy.