Platform SRE

YASH Technologies

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

YASH Technologies in Bengaluru seeks a seasoned Platform Reliability Engineer to own end‑to‑end production Kubernetes on AWS, drive performance tuning, and ensure reliability for AI/MCP workloads. You will mentor vendors and stakeholders and shape platform direction.

You will optimise memory usage, container runtimes, and serverless components, implement scalable CI/CD, and lead incident response in a regulated environment.

Qualifications

  • 5–10 years in software, DevOps, SRE or production support.
  • Working knowledge of GCP and/or Azure; multi-cloud.
  • CI/CD with GitHub Actions, Jenkins or GitLab.
  • Scripting in Python and/or Bash; YAML.

Responsibilities

  • Own production Kubernetes/EKS and AWS infrastructure reliability.
  • Tune memory and runtime performance; diagnose issues.
  • Deploy and operate LLM-backed AI services in cloud.
  • Monitor and troubleshoot platform at scale.

Skills

AWS expertise
Kubernetes (EKS)
Memory tuning
Serverless tuning
AI platform ops
Database ops
Terraform/CF
CI/CD pipelines
SRE/DevOps
Monitoring

Education

Bachelor's degree or equivalent

Tools

Helm
GitHub Actions
Jenkins
GitLab
Terraform

Job description

  • Owns reliability and performance of the platform end to end - leads incident response, drives cluster and runtime tuning standards, and sets direction for the AI/MCP platform.
  • Expected to mentor and influence vendor and stakeholder decisions.
  • Looking for a strong AWS & Kubernetes (EKS) expert to manage production-scale cloud infrastructure, platform reliability, performance tuning, and automation.
  • Experience with AI/LLM platform operations, monitoring, troubleshooting, and cloud-native environments is highly preferred.
  • This is a deep platform and performance-engineering role - owning production Kubernetes, AWS infrastructure, and the reliability of LLM-backed and agentic services in a regulated environment.
  • It goes well beyond standard SRE: we are looking for an engineer who tunes clusters, diagnoses memory and runtime behaviour, and operates AI platform workloads at production scale
Mandatory Skills:
  • Deep AWS expertise - End-to-end production deployment on AWS is a must - EC2, ECS, EKS, Lambda, S3, IAM, RDS, API Gateway, VPC, EFS, SNS, SQS, EventBridge, CodeBuild.
  • Kubernetes (deep & mandatory) - Production EKS ownership - cluster configuration (CPU/ RAM sizing, node groups, autoscaling), workload fine-tuning (requests/limits, HPA/VPA, eviction policies), and hands-on Helm chart management.
  • Memory expertise - Container vs. runtime memory models, OOMKill diagnosis, cgroup behaviour, JVM/Python heap tuning, and database buffer-pool configuration.
  • Serverless fine-tuning - Lambda memory/concurrency/cold-start optimization, ECS/Fargate task sizing, and serverless cost-performance tradeoffs.
  • Database operations - Running and tuning RDS or equivalent in production; StatefulSet-based DB deployments in K8s a strong plus.
  • AI platform & MCP environment - Hands-on deployment or operation of LLM-backed services, MCP servers, or agentic pipelines on cloud infrastructure.
Core Requirements:
  • 5-10 years in software development, DevOps, SRE or production support roles.
  • Working knowledge of GCP and/or Azure - multi-cloud integrations, platform differences, hybrid workloads.
  • CI/CD with GitHub Actions, Jenkins or GitLab.
  • Scripting proficiency in Python and/or Bash; YAML fluency.
  • Terraform and/or CloudFormation for infrastructure provisioning.
  • ITIL framework across Incident, Change, Problem and CAPA management.
  • Monitoring with Splunk and/or Grafana - including infra-level resource and memory dashboards.
  • ServiceNow and JIRA; strong ITSM discipline.
  • Bachelor's degree or equivalent practical experience
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technology Manager
Technology Manager

Wolters Kluwer • Pune District

Hybrid
INR 2,800,000 - 5,000,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Senior Director- Infrastructure, Operations & App Support
Senior Director- Infrastructure, Operations & App Support

Randstad • Hyderabad

Hybrid
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
DevOps & SRE Lead
DevOps & SRE Lead

Syngentagroup • Pune District

On-site
INR 4,000,000 - 7,000,000
Platform Engineer
Platform Engineer

United States Digital Space LLC • Maharashtra

On-site
INR 1,500,000 - 2,800,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
SRE Engineer
SRE Engineer

Prodapt Solutions Private Limited • Chennai District

On-site
INR 1,800,000 - 3,000,000
Cloud Operations Lead
Cloud Operations Lead

NewVision Software • Pune District

On-site
INR 3,000,000 - 5,500,000