Site Reliability Engineer

MishiPay

Bengaluru

On-site

INR 2,500,000 - 3,800,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MishiPay is seeking an Azure Cloud Infrastructure and SRE Engineer to own and operate our cloud platform across Dev, staging, production, and disaster recovery environments. You will drive reliability, scalability, and cost optimization while partnering with the InfoSec team to uphold security standards.

You will implement and maintain SLOs, observability, and incident response processes, and lead cloud migrations from Azure to GCP as part of modernization efforts.

Qualifications

  • 5+ years of hands-on experience managing production cloud infrastructure.
  • Strong Azure experience across AKS, VMs/VMSS, Networking, Storage, Monitoring and Security services.
  • Deep understanding of Site Reliability Engineering (SRE), SLOs, error budgets and incident management.
  • Experience with CI/CD pipelines using GitHub Actions or Azure DevOps.
  • Infrastructure as Code using Terraform.
  • Scripting skills in Python and Bash.
  • Experience with caching (Redis) and search/observability indexing (Elasticsearch).

Responsibilities

  • Own and operate Azure infrastructure across Dev, Staging, Production, and DR environments with high uptime targets.
  • Lead incident response, triage, RCA, and post-incident reviews for production incidents.
  • Define and maintain SLOs and alerting policies; build observability across logs/metrics/traces.
  • Manage AKS internals and migrate VM/VMSS workloads to Kubernetes.
  • Ensure database health across PostgreSQL, MySQL, MongoDB; manage backups and DR drills.
  • Build and maintain CI/CD pipelines and environment isolation.
  • Oversee Azure networking components and cloud cost optimization.
  • Support security, governance, and vulnerability remediation with InfoSec.

Skills

Azure
Kubernetes
SRE
CI/CD
Terraform
Python
Bash
Datadog
Observability
Incident management

Tools

AKS
VMs
VMSS
Azure Monitor
Datadog
Prometheus
Grafana
ELK
GitHub Actions
Azure DevOps
Terraform
Python
Bash
Cloudflare
Redis
Elasticsearch

Job description

We are looking for a proactive and detail-oriented Azure Cloud Infrastructure and SRE Engineer to drive the scalability, reliability, and evolution of our cloud platform. This role demands strong ownership across infrastructure operations, CI/CD pipelines, cloud migration, observability, performance tuning and incident response. You will also work closely with our InfoSec team to proactively identify and eliminate infrastructure vulnerabilities, ensuring compliance and security best practices.

You must have around 5 years of experience in both DevOps and SRE. You will have experience working on high-scale platforms that serve millions of users or process large volumes of real-time transactions with strict uptime and latency requirements. You will also have strong Azure and Kubernetes experience.

Please also note the additional requirements listed below, as we cannot consider anyone who doesn't have what we require. This is a role for someone who can hit the ground running and take ownership immediately.

You’ll work closely with the Director of Engineering and other squad members, alongside the Product, Payment, Security and Delivery teams, achieving the roadmap which has been set against our top business priorities. You’ll work on getting rid of tech debt, deploying best-in-class systems and architecture and ensuring that we can scale to 1000s of stores while maintaining system performance at over 99.9% at the push of a button.

If you’re a startup enthusiast with the required experience, who is passionate about solving complex problems and wants to learn something new every day, we’d absolutely love to speak to you!

Responsibilities
  • Own and operate Azure infrastructure across Dev, Staging, Production, and DR environments with 99.999% uptime SLAs.
  • Lead incident response; own triage, resolution and RCA for all production incidents.
  • Define and maintain SLOs, error budgets and alerting policies; build observability coverage across logs, metrics and traces (Datadog, Azure Monitor, Sentry).
  • Manage AKS internals - pods, deployments, ingress, autoscaling (HPA) - and drive migration of VM/VMSS workloads to Kubernetes.
  • Own database operational health and performance tuning across PostgreSQL, MySQL and MongoDB; manage backups and DR drills.
  • Build and maintain CI/CD pipelines with versioned deployments and environment isolation.
  • Manage Azure networking components including Application Gateway, Traffic Manager and Cloudflare (DNS, CDN, WAF).
  • Ensure infrastructure scales reliably with business growth while continuously optimising cloud costs through right-sizing, cleanup and spend monitoring.
  • Ensure security, compliance and governance; collaborate with InfoSec on vulnerability remediation.
  • Implement and manage caching (Redis) and search/observability indexing (Elasticsearch).
  • Lead cloud-to-cloud migration from Azure to GCP.
  • Support engineering, QA and support teams with access to cloud infrastructure and databases.
Requirements
  • 5+ years of hands-on experience managing production cloud infrastructure.
  • Strong Azure experience across AKS, VMs, VMSS, Networking, Storage, Monitoring and Security services.
  • Deep understanding of Site Reliability Engineering (SRE), SLOs, Error Budgets and Incident Management.
  • Hands-on experience with Datadog, Azure Monitor, Prometheus, Grafana or ELK.
  • Strong Kubernetes (AKS), Docker and Helm experience.
  • Experience with PostgreSQL, MySQL and MongoDB performance tuning and operations.
  • Experience building CI/CD pipelines using GitHub Actions or Azure DevOps.
  • Infrastructure as Code using Terraform.
  • Strong scripting skills in Python and Bash.
  • Experience with Cloudflare, Redis and Elasticsearch.
  • Understanding of networking, security, VPNs, firewalls and cloud governance.
Bonus Experience
  • Experience with GCP (GKE, Cloud SQL, VPC).
  • Azure or Kubernetes certifications.
  • Startup or scale-up experience.

This job was posted by Sanjana Supriya from Mishipay.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NCR Voyix • Chennai District

On-site
INR 3,000,000 - 5,400,000
SRE Engineer
SRE Engineer

Prodapt Solutions Private Limited • Chennai District

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Questhiring • Gurugram District

Hybrid
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ACG World • Mumbai

On-site
INR 2,200,000 - 3,200,000
Site Reliability Engineer
Site Reliability Engineer

S P A Enterprise Info Services • Chennai District

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer— Azure, AKS, Terraform | Global Platform team — Hyderabad
Senior Site Reliability Engineer— Azure, AKS, Terraform | Global Platform team — Hyderabad

CareerXperts Consulting • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer (Azure) - S
Senior Site Reliability Engineer (Azure) - S

Tata Consultancy Services • Kolkata District, Chennai District, Bengaluru

On-site
INR 2,400,000 - 4,200,000