DevOps Specialist Engineer SRE, Cloud & Applied AI

Clarus Advisers

Hyderabad

On-site

INR 1,200,000 - 2,400,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Clarus Advisers seeks a DevOps Specialist Engineer to design, build, and run large-scale cloud-native production systems. You will own reliability objectives, incident response, and deployment automation across Azure, AWS, or GCP.

You will contribute to AI/ML workloads in production, work with observability tools, and drive cost optimization and security controls. A strong engineering mindset is essential.

Qualifications

  • Bachelor's degree in CS, software engineering, data science or a related field.
  • 5+ years of DevOps/SRE experience with production-scale systems.
  • Hands-on experience with cloud platforms (Azure, AWS, GCP).
  • Strong knowledge of CI/CD, IaC, monitoring/observability, and incident management.

Responsibilities

  • Design, build, operate, and improve large-scale cloud-native production systems.
  • Define and own SLI/SLOs/SLAs, error budgets, and reliability objectives.
  • Lead production incident response, on-call operations, RCA, and improvements.
  • Build and manage CI/CD pipelines, Kubernetes platforms, and infrastructure automation.

Skills

SRE
Cloud platforms (Azure AWS GCP)
Kubernetes
Terraform
CI/CD pipelines
Observability
Python
OpenTelemetry

Education

Bachelor's degree in Computer Science/Software Engineering

Tools

Kubernetes
Docker
Terraform
ArgoCD
Prometheus
Grafana
Datadog
Dynatrace

Job description

Company Overview

Our client is a leading technology organization focused on building scalable, cloud-native platforms and intelligent software solutions. The organization combines software engineering, cloud technologies, Site Reliability Engineering, and Applied AI to deliver highly resilient and production-ready products at scale.

Position Overview

We are looking for a DevOps Specialist Engineer with strong experience in Site Reliability Engineering, cloud platform engineering, software development, and production operations. The ideal candidate will have hands‑on experience operating large‑scale distributed systems across Azure, AWS, or GCP, along with exposure to AI/ML, GenAI, LLMOps/MLOps, observability, performance engineering, and cloud cost optimization. The role requires a strong engineering mindset and the ability to build reliable, secure, scalable, and highly automated production platforms.

Responsibilities
  • Design, build, operate, and continuously improve large-scale cloud-native production systems.
  • Define and own SLIs, SLOs, SLAs, error budgets, and reliability objectives.
  • Lead production incident response, on‑call operations, root‑cause analysis, and reliability improvements.
  • Build and manage CI/CD pipelines, Kubernetes platforms, infrastructure automation, and multi‑environment deployments.
  • Implement Infrastructure as Code using Terraform and deployment automation using tools such as ArgoCD.
  • Develop production‑grade observability across metrics, logging, and distributed tracing.
  • Work with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, Azure Monitor, and AWS CloudWatch.
  • Operate AI/ML and GenAI workloads in production, addressing reliability, performance, model drift, output variance, and train/serve skew.
  • Support MLOps/LLMOps platforms and AI control‑plane capabilities such as model gateways and guardrails.
  • Implement Kubernetes/Docker‑based solutions and cloud‑native networking across multiple environments.
  • Conduct load and performance testing using tools such as LoadRunner, k6, or JMeter.
  • Implement chaos engineering, capacity planning, autoscaling, and resilience testing.
  • Drive cloud and AI FinOps, including GPU, inference, and token‑cost attribution and optimization.
  • Implement security controls including RBAC, least privilege, secrets management, deployment approvals, and segregation of duties.
  • Collaborate with software engineering, security, data/AI, and product teams to improve platform reliability and operational excellence.
Skills & Experience
  • 69 years of experience in DevOps, SRE, Site Reliability Engineering, Platform Engineering, or Software Engineering.
  • Bachelor's degree in Computer Science, Software Engineering, Data Science, Machine Learning, or a related discipline.
  • Strong programming experience in one or more of Python, C#/.NET, Go, Java, or Bash.
  • Strong hands‑on experience with Azure, AWS, or GCP; Azure/AWS preferred.
  • Experience with Kubernetes, Docker, Terraform, and cloud‑native architectures.
  • Strong understanding of CI/CD, GitHub, Azure DevOps (ADO), and ArgoCD.
  • Experience with production observability and monitoring using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, CloudWatch, or Azure Monitor.
  • Strong understanding of SRE principles, SLIs, SLOs, SLAs, error budgets, incident management, and production on‑call operations.
  • 3+ years of experience operating or supporting large‑scale production systems.
  • Experience with AI/ML or GenAI workloads in production and familiarity with Azure OpenAI, AWS Bedrock, or Vertex AI.
  • Knowledge of MLOps/LLMOps, MLflow, LangFuse, LangSmith, or equivalent AI/agent orchestration platforms.
  • Experience with load/performance testing, capacity planning, autoscaling, and chaos engineering.
  • Understanding of cloud cost optimization/FinOps, including AI/GPU/inference and token‑cost management.
  • Knowledge of DevSecOps, RBAC, secrets management, security controls, and environment integrity.
  • Strong understanding of OOP/OOD, data structures, algorithms, code instrumentation, and software engineering practices.
  • Ability to understand and work with business context, sequence, activity, state, entity‑relationship, and data‑flow diagrams.
  • Exposure to XP, Lean, SRE, and AI‑augmented/spec‑driven software development is an advantage.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior DevOps Engineer
Senior DevOps Engineer

Responsive • Bengaluru

On-site
INR 3,500,000 - 7,000,000
DevOps Engineer
DevOps Engineer

NeuralGarage • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Senior DevOps Engineer
Senior DevOps Engineer

Anaptyss • Dadri

On-site
INR 2,400,000 - 4,200,000
Lead SDE - DevOps
Lead SDE - DevOps

Flourish Ventures • Chennai District

On-site
INR 2,000,000 - 3,000,000
Inclusive and people-first culture
Health & wellness programs
Comprehensive medical insurance
+2
Devops Engineer
Devops Engineer

Larsen & Toubro • Chennai

On-site
INR 1,000,000 - 1,500,000
DevOps Engineer
DevOps Engineer

KnowledgeWorks Global Ltd. • Mumbai

On-site
INR 1,800,000 - 3,000,000
Forward Deployment Engineer (DevOps, AI Deployment)
Forward Deployment Engineer (DevOps, AI Deployment)

PwC Acceleration Centers • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior DevOps Engineer
Senior DevOps Engineer

Friar Service India Private Limited • Pune District

On-site
INR 4,000,000 - 6,000,000
DevOps Architect
DevOps Architect

Zycus • Mumbai

On-site
INR 4,000,000 - 7,000,000