A leading technology firm is seeking a Senior Site Reliability Engineer (Azure) to build, maintain, and scale cloud-native infrastructure. This role will work with development and operations teams to ensure reliable and efficient systems. Key responsibilities include designing Azure environments, managing Kubernetes clusters, and enhancing CI/CD pipelines using Terraform and Databricks. The ideal candidate will have 4+ years in Site Reliability Engineering or cloud infrastructure, strong Azure skills, and a collaborative mindset.
Qualifications
Minimum 4 years of experience in Site Reliability Engineering, DevOps, or cloud infrastructure roles.
Strong hands-on experience with Azure cloud services.
Proficiency with Java and Infrastructure-as-Code tools including Terraform and Terragrunt.
Strong experience with Kubernetes (preferably AKS) and container orchestration.
Experience working with Databricks in production environments.
Proficiency with CI/CD tooling, especially GitHub Workflows/Actions and ArgoCD.
Strong understanding of observability tooling, including Grafana (Prometheus, Loki, Tempo preferred).
Ability to collaborate in cross-functional environments and communicate effectively.
Responsibilities
Design, implement, and maintain Azure cloud infrastructure using best practices.
Manage and optimize Kubernetes clusters and containerized workloads.
Build and maintain Infrastructure-as-Code solutions.
Develop, maintain, and enhance CI/CD pipelines.
Support Databricks environments and associated integrations.
Implement and improve observability using various tools.
Automate operational tasks to improve efficiency.
Participate in on-call rotations, incident response, and root-cause analysis.
Collaborate with developers to improve application performance.
Identify opportunities for cost optimization and infrastructure security enhancements.
Skills
Site Reliability Engineering
DevOps
Azure cloud services
Kubernetes
Terraform
CI/CD
Observability tools
Education
Master’s degree in Computer Science or related field
Tools
Terraform
Terragrunt
GitHub Workflows/Actions
ArgoCD
Grafana
Prometheus
Loki
Tempo
Databricks
Job description
A leading technology firm is seeking a Senior Site Reliability Engineer (Azure) to build, maintain, and scale cloud-native infrastructure. This role will work with development and operations teams to ensure reliable and efficient systems. Key responsibilities include designing Azure environments, managing Kubernetes clusters, and enhancing CI/CD pipelines using Terraform and Databricks. The ideal candidate will have 4+ years in Site Reliability Engineering or cloud infrastructure, strong Azure skills, and a collaborative mindset.