Senior SRE

CloudRaft

India

On-site

INR 2,500,000 - 4,500,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary
Premium health insurance & wellness
AI stack & GPU infrastructure
Collaborative learning environment
Opportunity to lead and deliver

Job summary

CloudRaft is seeking passionate Site Reliability Engineers to join our growing team. You will take end-to-end ownership of designing, building, operating, and scaling mission-critical infrastructure for our partners.

You will be responsible for ensuring reliability, performance, security, and operational excellence while driving automation, improving system efficiency, and implementing innovative solutions.

Qualifications

  • Bachelor’s degree in Computer Science, IT or a related field.
  • 5+ years of experience in SRE, Platform Engineering, or DevOps.
  • Strong expertise in Kubernetes and cloud-native technologies.
  • Proficiency in Python or Go or Node.js.
  • Familiarity with CI/CD tools and modern deployment practices.
  • Proficiency in open source observability stacks and IaC (Terraform/Pulumi).
  • CKA/CKAD Certified (Brownie points!).
  • Excellent problem-solving abilities and communication skills.
  • Inclination toward open-source contributions is advantageous.

Responsibilities

  • Manage and maintain Kubernetes clusters across cloud platforms including OpenShift, AWS, Azure, GCP.
  • Implement and manage CI/CD pipelines using Jenkins, GitHub Actions, Argo CD, or GitLab CI/CD.
  • Design and maintain observability stacks with Prometheus, Grafana, Loki, OpenTelemetry, etc.
  • Optimize system performance and resolve production issues; participate in 24x7 on-call coverage.
  • Apply SRE principles with SLIs/SLOs to uphold reliability.
  • Automate infrastructure and tasks using Go/Python and Terraform/Pulumi.
  • Explore AI-driven automation for SDLC and AIOps.
  • Stay updated on AI and GPU infrastructure trends.
  • Share knowledge via technical writing and presentations.

Skills

Kubernetes
CI/CD tools
Cloud-native
Go
Python
Node.js
OpenTelemetry
Terraform
OpenShift
CKA/CKAD

Education

Bachelor’s degree in Computer Science / IT or related

Tools

Jenkins
GitHub Actions
Argo CD
GitLab CI/CD
Prometheus
Grafana
Loki
OpenTelemetry
Terraform
Pulumi

Job description

About CloudRaft

CloudRaft is a premier cloud-native consulting and engineering company that helps ambitious startups and digital-first organizations build, scale, and operate mission‑critical platforms. We partner with innovators at the forefront of artificial intelligence, developer productivity, observability, digital commerce, and enterprise software—enabling them to accelerate growth with resilient, scalable, and production‑ready cloud infrastructure.

Our experience spans organizations developing AI safety and governance platforms, AI Cloud, AI agent ecosystems, developer tooling, observability solutions, digital health products, customer engagement platforms, and technology-driven franchise networks. By combining deep expertise in Platform Engineering, Kubernetes, DevOps, Observability, and Cloud Native technologies, CloudRaft helps high‑growth companies move faster, operate more reliably, and focus on building category‑defining products.

Job Description

We are looking for passionate Site Reliability Engineers (SREs) to join our growing team. In this role, you will take end‑to‑end ownership of designing, building, operating, and scaling mission‑critical infrastructure for our partners. You will be responsible for ensuring reliability, performance, security, and operational excellence while driving automation, improving system efficiency, and implementing innovative solutions. Working at the intersection of software engineering and operations, you will help create resilient platforms that enable fast‑growing organizations to scale with confidence.

Responsibilities
  • Manage and maintain Kubernetes clusters across cloud platforms, including OpenShift, Amazon EKS, Azure AKS, and Google GKE.
  • Implement and manage CI/CD pipelines using tools such as Jenkins, GitHub Actions, Argo CD, or GitLab CI/CD.
  • Design and maintain observability stacks with tools including Prometheus, Grafana, Loki, OpenTelemetry, and related technologies. Be part of the team who support open source projects like Prometheus, Thanos, Mimir, CloudNativePG, Istio and more.
  • Optimize system performance and resolve production issues. Be part of the on call roster to provide 24x7 coverage for the critical production systems.
  • Implement SRE principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs), to uphold system reliability.
  • Automate infrastructure and operational tasks using programming languages such as Go or Python, and Infrastructure as Code (IaC) tools like Terraform.
  • Apply agentic AIto automate the SDLC lifecycle, AIOps and automation.
  • Learn about emerging technologies, including AI, GPU Infrastructure
  • Contribute to knowledge sharing through technical writing and presentations.
Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, or a related field.
  • 5+ years of experience in SRE, Platform Engineering, or DevOps Engineer.
  • Strong expertise in Kubernetes, cloud‑native technologies, on‑premise and major cloud platforms (AWS, Azure, GCP).
  • Proficiency in programming languages such as Python or Go or Node.js.
  • Familiarity with CI/CD tools and modern deployment practices.
  • Proficiency in one or more open source observability stacks and Infrastructure as Code (Terraform/Pulumi).
  • CKA/CKAD Certified (Brownie points!)
  • Excellent problem‑solving abilities and communication skills.
  • Inclination toward open‑source contributions is advantageous.
Benefits
  • Competitive salary
  • Premium health insurance and various health & wellness benefits from a leading insurance provider through Plum
  • Opportunity to work on the latest AI stack and GPU infrastructure
  • Collaborative and supportive work environment full of learning
  • Chance to take a front seat where you lead and deliver
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer(SRE)
Site Reliability Engineer(SRE)

Cloudraft Technologies Private Limited. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Premium health insurance
Opportunity to work on cutting-edge technologies
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Senior SRE Engineer
Senior SRE Engineer

EPAM Systems • Gurugram District

On-site
INR 3,000,000 - 5,000,000
Lead SRE
Lead SRE

Jobgether • India

On-site
INR 3,000,000 - 6,500,000
Mentorship programs
Structured leadership development
Generous paid time off
+1
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Chennai District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Okta • Bengaluru

On-site
INR 1,500,000 - 2,500,000