Senior Site Reliability Engineer

Clarus Advisers

Hyderabad

On-site

INR 1,800,000 - 2,800,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Clarus Advisers is seeking a DevOps Specialist to design, build, and operate large-scale cloud-native production systems across Azure, AWS, or GCP. You will own SLI/SLO/SLAs, lead on-call incident response, and drive automation with Terraform and ArgoCD while improving observability with Prometheus, Grafana, and OpenTelemetry.

You will work on production AI/ML workloads, MLOps/LLMOps, and cloud FinOps, ensuring security controls and scalable deployments across multiple environments.

Qualifications

  • 6–9 years of experience in DevOps, SRE, Platform Engineering, or Software Engineering.

Responsibilities

  • Design, build, operate, and continuously improve large-scale cloud-native production systems.
  • Define and own SLIs, SLOs, SLAs, error budgets, and reliability objectives.
  • Lead production incident response, on-call operations, root-cause analysis, and reliability improvements.
  • Build and manage CI/CD pipelines, Kubernetes platforms, infrastructure automation, and multi-environment deployments.
  • Implement Infrastructure as Code using Terraform and deployment automation using ArgoCD.
  • Develop production-grade observability across metrics, logging, and distributed tracing.
  • Work with monitoring tools such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, Azure Monitor, and AWS CloudWatch.
  • Operate AI/ML and GenAI workloads in production, addressing reliability, performance, model drift, output variance, and train/serve skew.
  • Support MLOps/LLMOps platforms and AI control‑plane capabilities such as model gateways and guardrails.
  • Implement Kubernetes/Docker-based solutions and cloud-native networking across multiple environments.
  • Conduct load and performance testing using LoadRunner, k6, or JMeter.
  • Implement chaos engineering, capacity planning, autoscaling, and resilience testing.
  • Drive cloud and AI FinOps, including GPU, inference, and token-cost attribution and optimization.
  • Implement security controls including RBAC, least privilege, secrets management, deployment approvals, and segregation of duties.
  • Collaborate with software engineering, security, data/AI, and product teams to improve platform reliability and operational excellence.

Skills

SRE principles
On-call operations
CI/CD
Cloud cost optimization
Security controls
Communication

Education

Bachelor's degree in Computer Science/Engineering

Tools

Kubernetes
Docker
Terraform
ArgoCD
Prometheus
Grafana
OpenTelemetry
Datadog
Dynatrace
Splunk
Azure Monitor
AWS CloudWatch

Job description

Company Overview

Our client is a leading technology organization focused on building scalable, cloud-native platforms and intelligent software solutions.

Position Overview

We are looking for a DevOps Specialist Engineer with strong experience in Site Reliability Engineering, cloud platform engineering, software development, and production operations. The ideal candidate will have hands‑on experience operating large-scale distributed systems across Azure, AWS, or GCP, along with exposure to AI/ML, GenAI, LLMOps/MLOps, observability, performance engineering, and cloud cost optimization. The role requires a strong engineering mindset and the ability to build reliable, secure scalable, and highly automated production platforms.

Responsibilities
  • Design, build, operate, and continuously improve large-scale cloud-native production systems.
  • Define and own SLIs, SLOs, SLAs, error budgets, and reliability objectives.
  • Lead production incident response, on-call operations, root-cause analysis, and reliability improvements.
  • Build and manage CI/CD pipelines, Kubernetes platforms, infrastructure automation, and multi-environment deployments.
  • Implement Infrastructure as Code using Terraform and deployment automation using tools such as ArgoCD.
  • Develop production-grade observability across metrics, logging, and distributed tracing.
  • Work with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, Azure Monitor, and AWS CloudWatch.
  • Operate AI/ML and GenAI workloads in production, addressing reliability, performance, model drift, output variance, and train/serve skew.
  • Support MLOps/LLMOps platforms and AI control‑plane capabilities such as model gateways and guardrails.
  • Implement Kubernetes/Docker-based solutions and cloud-native networking across multiple environments.
  • Conduct load and performance testing using tools such as LoadRunner, k6, or JMeter.
  • Implement chaos engineering, capacity planning, autoscaling, and resilience testing.
  • Drive cloud and AI FinOps, including GPU, inference, and token-cost attribution and optimization.
  • Implement security controls including RBAC, least privilege, secrets management, deployment approvals, and segregation of duties.
  • Collaborate with software engineering, security, data/AI, and product teams to improve platform reliability and operational excellence.
Skills & Experience
  • 6-9 years of experience in DevOps, SRE, Site Reliability Engineering, Platform Engineering, or Software Engineering.
  • Bachelor's degree in Computer Science, Software Engineering, Data Science, Machine Learning, or a related discipline.
  • Strong programming experience in one or more of Python, C#/.NET, Go, Java, or Bash.
  • Strong hands‑on experience with Azure, AWS, or GCP; Azure/AWS preferred.
  • Experience with Kubernetes, Docker, Terraform, and cloud-native architectures.
  • Strong understanding of CI/CD, GitHub, Azure DevOps (ADO), and ArgoCD.
  • Experience with production observability and monitoring using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, CloudWatch, or Azure Monitor.
  • Strong understanding of SRE principles, SLIs, SLOs, SLAs, error budgets, incident management, and production on-call operations.
  • 3+ years of experience operating or supporting large-scale production systems.
  • Experience with AI/ML or GenAI workloads in production and familiarity with Azure OpenAI, AWS Bedrock, or Vertex AI.
  • Knowledge of MLOps/LLMOps, MLflow, LangFuse, LangSmith, or equivalent AI/agent orchestration platforms.
  • Experience with load/performance testing, capacity planning, autoscaling, and chaos engineering.
  • Understanding of cloud cost optimization/FinOps, including AI/GPU/inference and token-cost management.
  • Knowledge of DevSecOps, RBAC, secrets management, security controls, and environment integrity.
  • Strong understanding of OOP/OOD, data structures, algorithms, code instrumentation, and software engineering practices.
  • Ability to understand and work with business context, sequence, activity, state, entity-relationship, and data-flow diagrams.
  • Exposure to XP, Lean, SRE, and AI-augmented/spec-driven software development is an advantage.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Specialist Engineer SRE, Cloud & Applied AI
DevOps Specialist Engineer SRE, Cloud & Applied AI

Clarus Advisers • Hyderabad

On-site
INR 1,200,000 - 2,400,000
AI SRE/ AI Site Reliability Engineer
AI SRE/ AI Site Reliability Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Senior DevOps Engineer
Senior DevOps Engineer

Responsive • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000