Experience: 7+ years (flexible based on depth in AWS + Kubernetes + production ops)
Role Summary: We are looking for an experienced Site Reliability Engineer (SRE) to own reliability, scalability, automation, and operational excellence for cloud-native platforms running on AWS and Kubernetes (EKS). You will build and operate CI/CD pipelines, standardize infrastructure provisioning, implement monitoring/alerting, drive incident response, and partner with architects and engineering teams to deliver secure, cost- efficient, highly available systems.
Key Responsibilities:
Reliability & Operations (SRE Core)
- Own availability, latency, performance, and capacity for production workloads.
- Define and track SLIs/SLOs, error budgets, and reliability KPIs.
- Improve MTTR through automation, runbooks, and self-healing patterns.
- Operate and evolve AWS EKS clusters (multi-namespace, multi-env).
- Manage deployments using Helm/Kustomize, and enable safe rollouts (blue/green, canary).
CI/CD & Release Engineering
- Build and maintain CI/CD pipelines using GitHub Actions / Jenkins / GitLab CI (based on org standard).
- Container build & security scanning workflows (SAST/DAST/image scanning) and SBOM where required.
- Promote “everything in Git” culture: infra/app configs + deployment manifests. Infrastructure as
- Provision cloud infra using Terraform / CloudFormation / CDK (as applicable).
- Manage multi-account AWS access patterns (dedicated Dev accounts, IAM roles, federation/SSO).
- Implement cost controls: tagging standards, budgeting alerts, right-sizing, savings plans guidance.
Observability (Monitoring, Logging, Tracing)
- Own observability stack: Prometheus, Grafana, CloudWatch, Alertmanager.
- Implement metrics and dashboards for K8s cluster health + application SLIs.
- Enable distributed tracing (OpenTelemetry / X-Ray / Jaeger) where needed.
Security & Governance (DevSecOps)
- Implement edge and app protection using AWS WAF / CloudFront, and coordinate rollout safely.
- Enforce least privilege IAM, secrets handling, KMS encryption, secure networking.
- Support vulnerability remediation (CVEs), patching processes, and audit readiness.
- Partner with Engineering + Architecture to review designs and keep solutions simple.
- Create/maintain runbooks, SOPs, onboarding guides, and deployment standards.
- Raise infra requests and coordinate with platform/GT teams to ensure timely provisioning.
Must-Have Skills
- Strong hands-on experience with AWS (EKS, EC2, IAM, VPC, ALB/NLB, CloudWatch, S3, RDS).
- Deep experience managing Kubernetes in production, preferably EKS.
- Strong CI/CD ownership: pipelines, environments, release controls, rollback strategies.
- Observability experience: Prometheus + Grafana, alerting, dashboards, incident response.
- Infrastructure as Code: Terraform (preferred) or CloudFormation/CDK.
- Solid Linux + networking fundamentals (DNS, TLS, load balancing, routing).
Good-to-Have Skills
- GitOps tools: Argo CD / Flux.
- Experience with WAF tuning, bot rules, rate limiting, and safe production rollout.
- PostgreSQL/MySQL operations basics (performance, monitoring, connection pooling).
- Experience with high-volume ingestion/data pipelines is a plus.
AWS,GitHub Actions , Jenkins , GitFlows, Terraforms, EKS