Cloud Engineer

Cloudifyops

Chennai District

On-site

INR 1,200,000 - 1,800,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

CloudifyOps is seeking an experienced SRE/DevOps professional to own on‑call incidents, drive end‑to‑end resolution, and help shape an AI‑powered monitoring tool. You will manage dashboards, alerts, and logs across environments while contributing to infrastructure on AWS using Terraform and Kubernetes.

Join a hands‑on, on‑call role that values curiosity, proactive problem solving, and clear communication with both engineers and clients.

Responsibilities

  • Handle the on‑call rotation and own incidents end‑to‑end triage, mitigation, escalation where needed, and clean resolution.
  • Write clear, structured RCAs after every significant incident what happened, when, why, and what changes going forward.
  • Maintain and improve the monitoring stack across environments dashboards, alerting rules, log pipelines, and distributed traces. Treat noisy alerts as a problem to fix, not something to mute.
  • Provision and manage cloud infrastructure on AWS using Terraform. This is a hands‑on role not just reviewing what others set up.
  • Work with Kubernetes across multiple environments, debugging pod and node issues.
  • Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management.
  • Track application performance using APM tooling and JVM metrics: spot anomalies, investigate degradation, and flag systemic issues before they become incidents.
  • Contribute to the AI monitoring tool initiative: prototype, test, iterate. This is early‑stage work and needs someone willing to figure things out, not just execute a finished design.

Skills

Incident mgmt
RCA writing
Monitoring
SRE

Tools

AWS
Kubernetes
Terraform
Linux
Jenkins
Rancher
Docker
Grafana
Prometheus
ELK/EFK

Job description

Job Description
Culture at CloudifyOps :

Working at CloudifyOps is a rewarding experience! Great people, a work environment that thrives on creativity, and the opportunity to take on roles beyond a defined job description are just some of the reasons you should work with us.

About the Role :

Were looking for someone who genuinely wants to understand why systems fail, not just respond to alerts. This role sits at the crossroads of cloud infrastructure and production reliability. You'll own monitoring, handle on-call, and be the person who digs in when things go wrong. At the same time, we're building an AI‑powered pipeline monitoring tool and need someone curious enough to contribute to shaping it, not just watching over it.

What you'll do:
  • Handle the on‑call rotation and own incidents end‑to‑end triage, mitigation, escalation where needed, and clean resolution. You don’t pass the baton and disappear.
  • Write clear, structured RCAs after every significant incident what happened, when, why, and what changes going forward. These go to clients, so they need to work for both an engineer and a non‑technical reader.
  • Maintain and improve the monitoring stack across environments dashboards, alerting rules, log pipelines, and distributed traces. Treat noisy alerts as a problem to fix, not something to mute.
  • Provision and manage cloud infrastructure on AWS using Terraform. This is a hands‑on role not just reviewing what others set up.
  • Work with Kubernetes across multiple environments, debugging pod and node issues.
  • Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management.
  • Track application performance using APM tooling and JVM metrics: spot anomalies, investigate degradation, and flag systemic issues before they become incidents.
  • Contribute to the AI monitoring tool initiative: prototype, test, iterate. This is early‑stage work and needs someone willing to figure things out, not just execute a finished design.
Tech Stack:
  • Cloud & Infrastructure : AWS, Kubernetes (K8s), Terraform, Linux
  • Observability & Metrics : Prometheus, Grafana, APM (Datadog / New Relic / Kfuse), JVM Metrics & GC Analysis, ELK / EFK Stack, Distributed Tracing
  • CI/CD & Platform : Jenkins, ArgoCD, Rancher, Git, Docker
  • Good to Have(Not Mandatory) : Python / Bash scripting, OpenTelemetry, Zenduty / OpsGenie, ML / AI basics
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Engineer
Cloud Engineer

CloudifyOps Pvt Ltd • Chennai District

On-site
INR 900,000 - 1,300,000
DevOps Engineer
DevOps Engineer

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Cloud DevOps Engineer
Cloud DevOps Engineer

Lightrains Technolabs • Thiruvananthapuram

On-site
INR 800,000 - 1,600,000
5 Days Working
Flexible Timings
Additional equipment budget
+1
Cloud Engineer DevOps AWS
Cloud Engineer DevOps AWS

Unified Consultancy Services • Delhi

Hybrid
INR 800,000 - 1,200,000
Cloud Engineer
Cloud Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 800,000 - 1,200,000
DevOps Engineer I - (AWS, Azure, GCP)
DevOps Engineer I - (AWS, Azure, GCP)

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,000,000 - 1,500,000
DevOps Engineer
DevOps Engineer

NAVVYASA CONSULTING PRIVATE LIMITED • Gurugram District

On-site
INR 800,000 - 1,200,000
DevOps Engineer II- (AWS)
DevOps Engineer II- (AWS)

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Sr. DevOps Engineer II
Sr. DevOps Engineer II

CloudifyOps Pvt Ltd • Chennai District

On-site
INR 1,000,000 - 1,500,000