DevOps & Site Reliability Engineer

Keka Technologies Private Limited

Coimbatore District

On-site

INR 1,200,000 - 2,000,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Aivar Innovations is seeking a skilled DevOps & Site Reliability Engineer to join our Platform Engineering team in India. You will own reliability, scalability, and operational excellence of cloud-native infrastructure, designing and operating CI/CD pipelines, Kubernetes workloads on AWS, and observability stacks.

This hands-on role emphasizes automation, reducing toil, and championing SRE practices across the organization. You will drive cost-efficiency and secure, scalable deployments.

Qualifications

  • 3–5 years of hands-on experience in a DevOps, Platform Engineering, or Site Reliability Engineering role.
  • Proven track record managing production Kubernetes clusters at scale — EKS strongly preferred.
  • Deep working knowledge of AWS services, including compute, networking, storage, and managed databases.
  • Strong proficiency in IaC tools: Terraform, AWS CDK, or CloudFormation.
  • Solid scripting ability in Python, Bash, or Go for automation and tooling development.

Responsibilities

  • Design, provision, and maintain production-grade AWS infrastructure using IaC tooling.
  • Manage multi-cluster Kubernetes environments (EKS) with upgrades, scaling, network policies, and security hardening.
  • Administer AWS services such as RDS, ElastiCache, MSK, S3, CloudFront, Route 53, and VPC networking.
  • Build, maintain, and improve CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tooling.
  • Champion GitOps workflows using ArgoCD or Flux for Kubernetes deployments.
  • Develop and maintain Helm charts and Kustomize overlays for consistent deployments.
  • Define and track SLOs/SLIs and maintain observability stacks with Prometheus, Grafana, and Datadog.
  • Lead incident response, post-incident reviews, and long-term remediation tracking.

Skills

Kubernetes administration
AWS cloud services
IaC (Terraform/CloudFormation/AWS CDK)
Scripting (Python/Bash/Go)
CI/CD automation
GitOps (ArgoCD/Flux)

Tools

Terraform
CloudFormation
AWS CDK
GitHub Actions
Jenkins
ArgoCD
Flux
Helm
Kustomize
Prometheus
Grafana
Datadog
EKS

Job description

Aivar is an AI-first technology partner where cutting-edge technology meets industry expertise to supercharge your projects. Our AI-augmented teams accelerate development, reduce time-to-market, and deliver exceptional code quality. We bring together the best minds in tech to craft scalable, repeatable solutions that drive real momentum for your business.

Role Overview

Aivar Innovations is looking for a skilled DevOps & Site Reliability Engineer to join our Platform Engineering team. In this role you will own the reliability, scalability, and operational excellence of our cloud-native infrastructure. You will partner closely with product engineering teams to design and operate CI/CD pipelines, Kubernetes-based workloads on AWS, and observability stacks — ensuring that Aivar's accelerator platforms run with the highest possible uptime and performance.

This is a hands-on role with significant ownership. You will be expected to drive automation-first thinking, reduce toil, and champion SRE principles across the organisation.

Requirements
KEY RESPONSIBILITIES
Infrastructure & Platform Operations
  • Design, provision, and maintain production-grade AWS infrastructure using Infrastructure-as-Code (Terraform / CloudFormation).
  • Manage and optimise multi-cluster Kubernetes environments (EKS) — including cluster upgrades, node scaling, network policies, and security hardening.
  • Administer and tune AWS-managed services: RDS, ElastiCache, MSK (Kafka), S3, CloudFront, Route 53, and VPC networking.
  • Implement and own cost-optimisation strategies — right-sizing, spot usage, savings plans, and tagging governance.
CI/CD & Automation
  • Build, maintain, and improve CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tooling.
  • Drive automation of operational tasks — provisioning, patching, scaling events, and incident runbooks — to minimise manual intervention.
  • Champion GitOps workflows using ArgoCD or Flux for Kubernetes application delivery.
  • Develop and maintain Helm charts and Kustomize overlays for consistent, repeatable deployments across environments.
Site Reliability & Observability
  • Define and track SLOs, SLIs, and error budgets for production services.
  • Build and operate observability platforms using Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent tooling.
  • Lead incident response efforts — on-call participation, post-incident reviews, and long-term remediation tracking.
  • Perform chaos engineering experiments and game days to proactively surface reliability risks.
Security & Compliance
  • Enforce security best practices across the platform — IAM least-privilege, secrets management (AWS Secrets Manager / Vault), container image scanning.
  • Collaborate with the security team on vulnerability management and patching cadence.
Preferred Technical Qualifications
  • 3 – 5 years of hands-on experience in a DevOps, Platform Engineering, or Site Reliability Engineering role.
  • Proven track record managing production Kubernetes clusters at scale — EKS strongly preferred.
  • Deep working knowledge of AWS services, including compute (EC2, ECS, Lambda), networking (VPC, ALB, NLB, Transit Gateway), storage, and managed databases.
  • Strong proficiency in at least one IaC tool: Terraform (preferred), AWS CDK, or CloudFormation.
  • Solid scripting ability in Python, Bash, or Go for automation and tooling development.
PREFERRED QUALIFICATIONS
  • AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect certification.
  • Certified Kubernetes Administrator (CKA) or Certified Kubernetes Security Specialist (CKS).
  • Experience with service mesh technologies such as Istio or Linkerd.
  • Prior experience operating AI/ML or data-intensive workloads on Kubernetes (GPU node pools, Karpenter, etc.).
  • Exposure to multi-cloud or hybrid-cloud environments.
Why You’ll Love Working at Aivar
  • Learn from Experts: Work directly with former AWS leaders and AI pioneers.
  • Direct Ownership: Lead high-impact "greenfield" projects from concept to global launch.
  • Modern Tech: Master the latest Generative AI frameworks and cloud-native architectures.
  • Real-World Impact: Build mission-critical systems used by major global enterprises.
  • Rapid Growth: Scale your career quickly in a high-speep
Diversity and Inclusion

Aivar Innovations is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to gender, gender identity, sexual orientation, religion, disability, age, marital status, caste, or any other protected characteristic, and we are committed to building a diverse, inclusive, and respectful workplace for everyone.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps & Site Reliability Engineer
DevOps & Site Reliability Engineer

Aivar Innovations • Coimbatore District

On-site
INR 1,500,000 - 2,300,000
Mentorship
Ownership of greenfield projects
Modern cloud-native tech exposure
+2
Associate AI Architect
Associate AI Architect

Keka Technologies Private Limited • India

On-site
INR 4,000,000 - 6,500,000
Senior MLOps / AI Platform Engineer
Senior MLOps / AI Platform Engineer

Keka Technologies Private Limited • Coimbatore District

On-site
INR 3,000,000 - 6,000,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Aivar Innovations • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Learn from Experts
Direct Ownership of projects
Modern Generative AI stacks
+2
Senior Enterprise Account Manager
Senior Enterprise Account Manager

Keka Technologies Private Limited • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Learn from Experts
Direct Ownership
Modern Tech
+2
Associate Principal Site Reliability Engineer
Associate Principal Site Reliability Engineer

Ten Eleven Ventures • Bengaluru

On-site
INR 2,500,000 - 3,500,000
AI INFRASTRUCTURE ENGINEER
AI INFRASTRUCTURE ENGINEER

AI Innovation and Inclusion Initiative (A4I) • Bengaluru

On-site
INR 2,800,000 - 4,400,000
Senior DevOps Engineer
Senior DevOps Engineer

Avathon • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Senior Enterprise Account Manager – New Business / Growth
Senior Enterprise Account Manager – New Business / Growth

Keka Technologies Private Limited • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Associate Principal Site Reliability Engineer
Associate Principal Site Reliability Engineer

Saviynt • Bengaluru

On-site
INR 4,000,000 - 6,000,000