Associate Principal Site Reliability Engineer

Saviynt

Bengaluru

On-site

INR 4,000,000 - 6,000,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Saviynt is hiring a Staff SRE Engineer to own reliability across our cloud-native, AI-driven platform. You’ll work at the intersection of distributed systems, Kubernetes operations, and AI-powered automation to build systems that scale and self-heal.

You'll design self-healing infrastructure, implement LLM-powered operational tooling using APIs like OpenAI, and lead incident response with postmortems to drive continuous reliability improvements for a world-class identity security platform.

Qualifications

  • 8+ years in SRE / DevOps / Platform Engineering.
  • Strong hands-on experience with AWS at scale and production Kubernetes.
  • Proficient in Python or Go for internal tooling.
  • Experience with monitoring, alerting, and incident management.

Responsibilities

  • Own uptime, reliability, and performance of services on AWS + Kubernetes (EKS).
  • Design and implement self-healing infrastructure using automation and AI agents.
  • Build LLM-powered operational tooling using APIs like the OpenAI API for alert triage, incident summaries, root cause analysis, and runbook automation.
  • Manage and scale Kubernetes workloads: deployments, autoscaling, resource optimization, cluster reliability and cost efficiency.
  • Build and evolve observability systems: metrics (Prometheus), dashboards (Grafana), logs (ELK/OpenSearch), tracing (OpenTelemetry).
  • Define and enforce SLOs/SLAs and automate infrastructure with Terraform and CI/CD.
  • Lead incident response, postmortems, and continuous reliability improvements; introduce chaos engineering.

Skills

Distributed systems
Python
Go
Kubernetes
AWS
LLM APIs

Education

Bachelor's degree in CS/EE

Tools

Terraform
Prometheus
Grafana
OpenTelemetry
OpenSearch
OpenAI API

Job description

Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world’s leading brands, Fortune 500 companies and government institutions. For more information, please visit www.saviynt.com .


We’re a fast-moving AI Security Company building AI-native infrastructure and applications powered by LLMs and autonomous agents. Our stack is deeply integratedwith AWS, Kubernetes, and OpenAI-based systems, and we’re rethinking reliability in aworld where software can reason, adapt, and self-heal.


We’re hiring a Staff SRE Engineer to own reliability across our cloud-native and AI-driven platform. You’ll work at the intersection of distributed systems, Kubernetesoperations, and LLM-powered automation, building systems that don’t just scale—butthink and fix themselves.


Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world’s leading brands, Fortune 500 companies and government institutions. For more information, please visit www.saviynt.com .


We’re a fast-moving AI Security Company building AI-native infrastructure and applications powered by LLMs and autonomous agents. Our stack is deeply integratedwith AWS, Kubernetes, and OpenAI-based systems, and we’re rethinking reliability in aworld where software can reason, adapt, and self-heal.


We’re hiring a Staff SRE Engineer to own reliability across our cloud-native and AI-driven platform. You’ll work at the intersection of distributed systems, Kubernetesoperations, and LLM-powered automation, building systems that don’t just scale—butthink and fix themselves.


WHAT YOU WILL BE DOING


  • Own uptime, reliability, and performance of services running on AWS + Kubernetes (EKS)

  • Design and implement self-healing infrastructure using automation and AI agents

  • Build LLM-powered operational tooling using APIs such as the OpenAI API for:

    • Intelligent alert triage

    • Incident summarization

    • Root cause analysis

    • Runbook automation



  • Manage and scale Kubernetes workloads:

    • Deployments, autoscaling, resource optimization

    • Cluster reliability and cost efficiency



  • Build and evolve observability systems:

    • Metrics (Prometheus), dashboards (Grafana)

    • Logs (ELK / OpenSearch)

    • Tracing (OpenTelemetry)



  • Define and enforce SLOs, SLAs, and error budgets tied to business metrics

  • Automate infrastructure using Terraform and CI/CD pipelines

  • Lead incident response, postmortems, and continuous reliability improvements

  • Introduce chaos engineering practices to proactively test system resilience


WHAT YOU BRING


  • 8+ years in SRE / DevOps / Platform Engineering

  • Strong hands-on experience with:

    • AWS infrastructure at scale

    • Kubernetes (production-grade clusters)



  • Proven ability to debug complex distributed systems under pressure

  • Strong coding skills (Python or Go)—you build internal platforms and tools

  • Experience implementing monitoring, alerting, and incident management systems


Bonus (AI / LLM Focus)


  • Experience working with LLM APIs such as the OpenAI API

  • Familiarity with agent frameworks like:

    • LangChain

    • AutoGen



  • Built or experimented with:

    • AI agents for DevOps / SRE workflows

    • Retrieval-Augmented Generation (RAG) systems

    • Vector databases (Pinecone, Weaviate, etc.)



  • Exposure to AIOps or intelligent automation systems


If required for this role, you will:


  • Complete security & privacy literacy and awareness training during onboarding and annually thereafter

  • Review (initially and annually thereafter), understand, and adhere to Information Security/Privacy Policies and Procedures such as (but not limited to):

    • Data Classification, Retention & Handling Policy

    • Incident Response Policy/Procedures

    • Business Continuity/Disaster Recovery Policy/Procedures

    • Mobile Device Policy

    • Account Management Policy

    • Access Control Policy

    • Personnel Security Policy

    • Privacy Policy




Saviynt is an amazing place to work. We are a high-growth, Platform as a Service company focused on Identity Authority to power and protect the world at work. You will experience tremendous growth and learning opportunities through challenging yet rewarding work which directly impacts our customers, all within a welcoming and positive work environment. If you're resilient and enjoy working in a dynamic environment you belong with us!


Saviynt is an equal opportunity employer and we welcome everyone to our team. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.


We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE/Senior SRE
SRE/Senior SRE

Saviynt • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Associate Principal Site Reliability Engineer
Associate Principal Site Reliability Engineer

Ten Eleven Ventures • Bengaluru

On-site
INR 2,500,000 - 3,500,000
Platform Support Engineer -
Platform Support Engineer -

Saviynt • Bengaluru

On-site
INR 1,800,000 - 3,600,000
Manager, Customer Support
Manager, Customer Support

Saviynt • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Senior Software Engineer
Senior Software Engineer

Saviynt • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Associate Principal Engineer - Technology and Cloud Alliances
Associate Principal Engineer - Technology and Cloud Alliances

Saviynt • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Associate Principal Engineer/Senior Software Engineer
Associate Principal Engineer/Senior Software Engineer

Saviynt • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Principal Engineer - Java/AI
Principal Engineer - Java/AI

Saviynt • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Sr. Manager - Cloud Ops
Sr. Manager - Cloud Ops

Saviynt Inc. • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
High-growth work environment
Learning opportunities
Dynamic work culture
Associate Principle Engineer
Associate Principle Engineer

Saviynt Inc. • Bengaluru

On-site
INR 2,500,000 - 5,500,000