An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Cloud Soft Solutions is seeking a Senior SRE in Hyderabad to build reliable Azure and GCP platforms with Terraform and Python. You will run AKS/GKE clusters, implement SLO-driven reliability, and collaborate on multi-region disaster recovery in a healthcare context.
Ideal candidates have 5+ years in SRE/DevOps with Terraform expertise, Kubernetes experience, and strong Python/CI-CD skills. The role offers a hybrid work arrangement in Hyderabad and a competitive INR salary package.
Senior SRE for UnitedHealth Group / Optum's healthcare-scale platforms in Hyderabad — build reliable Azure / GCP infrastructure with Terraform and Python, run AKS / GKE clusters, and drive SLO-led reliability engineering. 5+ years, hybrid Hyderabad, INR 24-44 LPA band.
Build and maintain reliable cloud platforms using Terraform for infrastructure-as-code on Azure and GCP, supporting healthcare-scale systems where uptime directly affects patients and providers.Manage Kubernetes clusters (AKS and / or GKE): cluster lifecycle and upgrades, autoscaling, Workload Identity for least-privilege IAM, ingress, and platform add-ons via Helm.Implement observability, alerting and automated remediation — define and track SLIs / SLOs, instrument services with Prometheus / Grafana and cloud-native monitoring, and reduce toil through self-healing automation.Integrate Python scripting for custom reliability tooling, CI/CD optimisation, and infrastructure management; conduct blameless post-mortems and codify learnings into runbooks.Ensure compliance, security and cost governance in a regulated healthcare environment — secrets management, encryption, audit trails (HIPAA-aware), and FinOps reporting.Collaborate cross-functionally to deliver resilient microservices and disaster-recovery strategies, including multi-region failover and tested restore procedures.
5+ years of SRE / DevOps / Platform Engineering experience with a strong reliability mindset.Terraform expertise on Azure and / or GCP — reusable modules, remote state, plan reviews.Production Kubernetes experience (AKS or GKE preferred).Strong Python for automation tooling and CI/CD; comfortable with Bash.Hands-on observability (Prometheus, Grafana, cloud-native monitoring) and incident-response tooling.SLI / SLO definition, error budgets, and post-incident review discipline.Experience in regulated or large-scale environments (healthcare, BFSI) is a strong plus; HIPAA awareness valued.Bachelor's in CS / IT or equivalent. Azure / GCP and CKA certifications a plus.
Optum (the technology and health-services arm of UnitedHealth Group, a Fortune-5 company) runs one of the largest healthcare-technology engineering centres in Hyderabad. Reliability and compliance are first-class priorities here, which makes it an excellent environment to deepen genuine SRE skills — SLOs, error budgets, multi-cloud Terraform and disaster recovery — inside a stable, well-funded organisation.
INR 24-44 LPA fixed band for senior profiles plus annual bonus (market-rate estimate; confirmed at offer). Comprehensive health coverage for self + family + parents. Hybrid model in Hyderabad. Sponsored Azure / GCP / Kubernetes certifications and learning budget. Strong leave, parental-leave and wellness benefits typical of a Fortune-5 employer.