Get more replies from employers
Send a job-specific resume in minutes.
Aura Recruitment Solutions is seeking a Senior DevOps / Site Reliability Engineer (SRE) for a remote role based in India. You will design, build, and maintain robust cloud infrastructure, CI/CD pipelines, and high-availability platforms.
Key focus areas include AWS/Azure, Kubernetes, Terraform, Docker, and observability tools. You will collaborate with software, data, and security teams to improve platform reliability, automation, and scalability, while driving DevSecOps and cost optimization
This is a remote position.
Location: India
Work Mode: Remote
Employment Type: Full-time
Experience:5+ Years
We are seeking an experienced Senior DevOps / Site Reliability Engineer (SRE) to design, build, and maintain robust cloud infrastructure, automated CI/CD pipelines, and high-availability operational platforms. The ideal candidate brings strong hands-on expertise in AWS/Azure, Kubernetes, Terraform, Docker, CI/CD tools (GitHub Actions/Azure DevOps/Jenkins), and observability frameworks (Prometheus/Grafana/ELK).
In this role, you will work closely with software engineering, data, and security teams to drive platform reliability, infrastructure automation, developer efficiency, and system scalability.
Infrastructure as Code (IaC): Provision, manage, and scale cloud infrastructure using Terraform, CloudFormation, or Ansible.
Container Orchestration: Manage production-grade Kubernetes (EKS/AKS) clusters, helm charts, container runtimes, and ingress controllers.
CI/CD Automation: Design, optimize, and secure automated build, test, and deployment pipelines using GitHub Actions, Azure DevOps, or Jenkins.
Observability & Monitoring: Implement enterprise monitoring, log aggregation, and tracing systems using Prometheus, Grafana, Datadog, or ELK Stack.
Reliability & Incident Management: Own system uptime, define SLOs/SLIs, maintain incident response protocols, and lead post-mortem root cause analyses (RCA).
Security & Compliance: Enforce DevSecOps best practices, secret management (HashiCorp Vault/Key Vault), IAM policies, and cloud vulnerability scanning.
Cost & Performance Optimization: Audit cloud resource utilization, implement autoscaling strategies, and optimize cloud spending.
Experience: 6+ years of hands-on experience in DevOps, Site Reliability Engineering, or Cloud Infrastructure.
Cloud Architecture: Deep expertise in AWS or Azure public cloud services and architecture.
Containerization & Orchestration: Strong experience with Docker and production Kubernetes administration.
Infrastructure as Code: Proficiency in Terraform for declarative infrastructure provisioning.
CI/CD: Hands-on experience building multi-environment CI/CD pipelines with GitHub Actions, Azure DevOps, or Jenkins.
Scripting: Strong scripting capabilities in Python, Bash, or Go for operational automation.
Monitoring & Alerting: Solid background in setting up metrics, dashboards, and alerts via Prometheus and Grafana.
Experience with Service Mesh (Istio/Linkerd), GitOps workflows (ArgoCD/Flux), and Serverless infrastructure.
Familiarity with DevSecOps integrations (SonarQube, Trivy, Snyk).
About Our Client
Our client is an innovative financial technology platform headquartered in London, UK, delivering global payment rails, real-time transaction clearing, and open-banking API infrastructure. Operating under strict regulatory standards across European, APAC, and Americas markets, the organization prioritizes high availability, zero-downtime deployments, and robust security protocols.
Their engineering culture centers around a strong DevSecOps and Site Reliability Engineering (SRE) discipline, where automated chaos engineering, proactive observability, and self-healing infrastructure are core priorities. They foster a collaborative, blameless post-mortem environment where engineers have high ownership over system uptime, performance benchmarks, and modern multi-cloud deployment strategies.