We are seeking a Senior Azure Platform & DevOps Engineer to design, implement, automate, and support enterprise cloud infrastructure and platform services within Microsoft Azure. This role is responsible for managing Azure-based infrastructure, Kubernetes platforms, CI/CD pipelines, GitOps workflows, Infrastructure as Code (IaC), observability solutions, and cloud security controls. The ideal candidate combines deep Azure expertise with strong DevOps, automation, and platform engineering experience to drive scalable, secure, and highly available cloud solutions.
Responsibilities
Key Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and platform services.
- Ensure cloud environments are secure, scalable, resilient, and aligned with enterprise standards.
- Manage the full lifecycle of Kubernetes clusters, including provisioning, configuration, upgrades, patching, scaling, and decommissioning.
- Support production workloads running on AKS and troubleshoot complex platform issues.
- Enhance platform reliability, availability, and operational efficiency.
GitOps & Container Management
- Implement and maintain GitOps deployment methodologies using ArgoCD.
- Develop and manage Helm charts and container deployment standards.
- Support Docker image management and release processes.
DevOps & CI/CD Automation
- Build, maintain, and optimize CI/CD pipelines using GitHub Actions and Azure DevOps.
- Standardize build, test, and deployment processes across environments.
- Improve deployment quality, consistency, and release velocity.
Infrastructure as Code (IaC)
- Provision and manage cloud infrastructure using Terraform and Terragrunt.
- Develop reusable and version-controlled infrastructure modules.
- Implement automation to ensure consistent deployments across environments.
Databricks Platform Support
- Support Azure Databricks platform administration and infrastructure operations.
- Manage platform integrations, access controls, networking connectivity, and environment configurations.
- Assist application and data teams with infrastructure-related Databricks requirements.
Automation & Scripting
- Develop automation tools and operational scripts using Python and Bash.
- Automate infrastructure deployment, maintenance, monitoring, and operational tasks.
- Reduce manual effort through scalable automation solutions.
Cloud Security & Networking
- Design and implement Azure networking solutions including VNets, NSGs, firewalls, private endpoints, and routing.
- Configure and maintain role-based access controls (RBAC) and security policies.
- Ensure compliance with organizational security and governance requirements.
Monitoring, Observability & Optimization
- Implement and maintain cloud monitoring solutions using Azure Monitor and Application Insights.
- Develop performance monitoring and alerting strategies.
- Optimize cloud resource utilization and cost efficiency.
- Build and support enterprise observability platforms for metrics, logging, and alerting.
Skills
Must have
Required Skills
- Virtual Machines
- Container & Platform Technologies
- Docker
- ArgoCD
- GitOps methodologies
- DevOps & Automation
- GitHub Actions
- Source control and release management
- Infrastructure as Code
- Terraform
- Monitoring & Observability
- Application Insights
- Loki
- Promtail
- Fluentd
- Programming & Scripting
- Python
- Bash/Shell scripting
- Security & Networking
- Network Security Groups (NSGs)
- Virtual Networks (VNets)
- Firewall configuration
- Cloud security best practices
________________________________________
Required Experience
- 8+ years of experience designing and supporting Azure infrastructure platforms.
- 8+ years of Kubernetes platform administration and engineering experience.
- 8+ years working with cloud networking, security, and enterprise infrastructure.
- 8+ years supporting containerized application platforms.
- 6+ years implementing DevOps practices and CI/CD automation.
- 6+ years developing automation solutions with Python, Bash, or similar scripting languages.
- 5+ years implementing Infrastructure as Code using Terraform and Terragrunt.
- 5+ years supporting observability and monitoring platforms.
- 5+ years administering or supporting Azure Databricks environments.
- Experience supporting production systems in large-scale enterprise environments.
________________________________________
Preferred Qualifications
- Experience supporting multi-region and highly available cloud architectures.
- Experience implementing site reliability engineering (SRE) practices.
- Familiarity with FinOps and cloud cost optimization methodologies.
- Experience with enterprise governance, compliance, and security frameworks.
- Experience supporting data and analytics platforms.
________________________________________
Education Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
- Equivalent combination of education and professional experience may be considered.
________________________________________
Success Measures
The successful candidate will:
- Maintain highly available and secure Azure and Kubernetes platforms.
- Improve deployment reliability through automated CI/CD and GitOps processes.
- Reduce operational overhead through automation and Infrastructure as Code.
- Optimize cloud performance, observability, and cost efficiency.
- Provide rapid resolution of production incidents and platform issues.
- Drive continuous improvement across cloud, DevOps, security, and platform engineering practices.