Cloud Engineer

CloudifyOps Pvt Ltd

Chennai District

On-site

INR 900,000 - 1,300,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

CloudifyOps Pvt Ltd is seeking an observability/DevOps professional to own on-call incidents, drive end-to-end resolution, and improve monitoring across AWS, Kubernetes, and CI/CD stacks. You will contribute to an AI-powered monitoring tool initiative and communicate clearly with both engineers and clients.

The role emphasizes proactive incident handling, deep logs/metrics tracing, and building robust alerting and dashboards in a fast-paced environment.

Qualifications

  • 2.5 to 5 years of work experience in the role.
  • Strong fundamentals in distributed systems under load.
  • Ability to read logs, metrics and traces to diagnose issues.
  • Ownership of monitoring gaps and writing precise RCAs.
  • Willingness to handle on-call incidents and drive resolution.
  • Curiosity about AI tooling and experimentation.

Responsibilities

  • Handle the on-call rotation and own incidents end-to-end, including triage, mitigation, escalation, and resolution.
  • Write clear RCAs after significant incidents for engineers and non-technical readers.
  • Maintain and improve the monitoring stack across environments: dashboards, alerts, logs, and traces.
  • Provision and manage cloud infrastructure on AWS using Terraform; hands-on involvement.
  • Work with Kubernetes across environments, debugging pod and node issues.
  • Monitor CI/CD pipeline health via Jenkins and support Rancher-based workloads.
  • Track application performance with APM tools and JVM metrics to spot anomalies and degradation.
  • Contribute to the AI monitoring tool initiative: prototype, test, and iterate.

Tools

Prometheus
Grafana
APM tools
ELK/EFK
JVM Metrics
Kubernetes
Terraform

Job description

CloudifyOps is a company with DevOps and Cloud in our DNA. CloudifyOps enables businesses to become more agile and innovative through a comprehensive portfolio of services that addresses hybrid IT transformation, Cloud transformation, and end-to-end DevOps Workflows.

We are a proud Advance Partner of Amazon Web Services and have deep expertise in Microsoft Azure and Google Cloud Platform solutions.

We are passionate about what we do. The novelty and the excitement of helping our customers accomplish their goals drives us to become excellent at what we do.

Job Description

Culture at CloudifyOps :

Working at CloudifyOps is a rewarding experience! Great people, a work environment that thrives on creativity, and the opportunity to take on roles beyond a defined job description are just some of the reasons you should work with us.

About the Role :

We’re looking for someone who genuinely wants to understand why systems fail, not just respond to alerts. This role sits at the crossroads of cloud infrastructure and production reliability. You’ll own monitoring, handle on-call, and be the person who digs in when things go wrong. At the same time, we’re building an AI-powered pipeline monitoring tool and need someone curious enough to contribute to shaping it, not just watching over it.

What you’ll do:

Handle the on-call rotation and own incidents end-to-end triage, mitigation, escalation where needed, and clean resolution. You don’t pass the baton and disappear.

Write clear, structured RCAs after every significant incident what happened, when, why, and what changes going forward. These go to clients, so they need to work for both an engineer and a non-technical reader.

Maintain and improve the monitoring stack across environments dashboards, alerting rules, log pipelines, and distributed traces. Treat noisy alerts as a problem to fix, not something to mute.

Provision and manage cloud infrastructure on AWS using Terraform. This is a hands-on role not just reviewing what others set up.

Work with Kubernetes across multiple environments, debugging pod and node issues.

Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management.

Track application performance using APM tooling and JVM metrics: spot anomalies, investigate degradation, and flag systemic issues before they become incidents.

Contribute to the AI monitoring tool initiative: prototype, test, iterate. This is early-stage work and needs someone willing to figure things out, not just execute a finished design.

Observability & Metrics : Prometheus · Grafana · APM (Datadog / New Relic / Kfuse) · JVM Metrics & GC Analysis · ELK / EFK Stack · Distributed Tracing

Good to Have(Not Mandatory) : Python / Bash scripting · OpenTelemetry · Zenduty / OpsGenie · ML / AI basics

On-call here is real.Incidents happen outside business hours and when they do, it’s your responsibility to pick them up and drive them forward. That’s not unusual for this type of role but we want to be direct about it upfront.

Client expectations are high. You’ll produce RCAs, incident timelines, and status communications that clients read closely. Your writing needs to be clear, structured, and free of vagueness. “We investigated and fixed the issue” isn’t good enough. What was the issue, why did it happen, what was the business impact, and what prevents recurrence.

We expect precision regarding your own work. After a change, an incident, or a deployment, you should be able to clearly explain what you did and why without being prompted. Ownership doesn’t end when the alert clears.

Who we’re looking for:

The ideal candidate should have 2.5 years to 5 years of work experience.

Strong fundamentals. You understand how distributed systems actually behave under load, not just that a dashboard went red. You can read logs, metrics, and traces together.

Ownership without prompting. If you find a gap in monitoring coverage, you close it. If an RCA feels incomplete, you go back and make it precise. You don’t wait to be asked.

Writes clearly under pressure. During an incident, your updates should help not add noise. After one, your documentation should be good enough that anyone picking it up six months later understands what happened.

Curious about what comes next. The AI tooling initiative needs someone interested in figuring it out, not just waiting for a ticket. Some comfort with experimentation and ambiguity goes a long way here.

Equal opportunity employer

CloudifyOps is proud to be an equal opportunity employer with a global culture that embraces diversity. We are committed to providing an environment free of unfair discrimination and harassment. We do not discriminate based on age, race, color, sex, religion, national origin, disability, pregnancy, marital status, sexual orientation, gender reassignment, veteran status, or other protected category.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Engineer
Cloud Engineer

Cloudifyops • Chennai District

On-site
INR 1,200,000 - 1,800,000
DevOps Engineer
DevOps Engineer

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,500,000 - 2,100,000
DevOps Engineer II- (AWS)
DevOps Engineer II- (AWS)

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,800,000 - 2,800,000
DevOps Engineer I - (AWS, Azure, GCP)
DevOps Engineer I - (AWS, Azure, GCP)

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,000,000 - 1,500,000
DevOps Engineer
DevOps Engineer

Cloudifyops • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Senior Cloud Engineer
Senior Cloud Engineer

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,200,000 - 1,600,000
Sr. DevOps Engineer II
Sr. DevOps Engineer II

CloudifyOps Pvt Ltd • Chennai District

On-site
INR 1,000,000 - 1,500,000
Cloud DevOps Architect
Cloud DevOps Architect

CloudifyOps Pvt Ltd • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Cloud Engineer
Cloud Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 800,000 - 1,200,000
ZR_248_Manager CloudOps
ZR_248_Manager CloudOps

Priority Technology Holdings, Inc. • Chandigarh

On-site
INR 3,500,000 - 7,000,000
Five days working
Complimentary meal per day
Internet reimbursement
+2