IT Lead

Jobtailor

Pune District

On-site

INR 1,200,000 - 1,800,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Jobtailor in Pune, Maharashtra, India seeks an experienced Observability/SRE professional to design and operate telemetry systems for business-critical applications. You will implement standards for metrics, logs, and traces and drive automation across monitoring platforms.

The role emphasizes runbook creation, post-incident reviews, and collaboration with incident teams. Local office attendance is required part of each week, with IST hours and occasional on-call duties.

Qualifications

  • 4–8 years in Observability/SRE/Platform/Monitoring roles supporting SaaS or enterprise applications.
  • Hands-on experience with monitoring and logging tools.
  • Strong grasp of distributed tracing, RED/USE/golden signals, SLI/SLO/SLA, and error budgets.
  • Proficiency in Python.

Responsibilities

  • Design and operate the telemetry backbone for internal platforms and business-critical applications.
  • Define and implement standards for metrics, logs, traces, and profiling using tools such as OpenTelemetry collectors and exporters.
  • Establish golden signals, SLIs/SLOs, and health checks for priority services.
  • Automate baselining and anomaly detection.
  • Create executive and on-call dashboards, service health views, and dependency maps.
  • Develop alerting policy as code and reduce false positives through suppression and deduplication.
  • Implement auto-remediation runbooks.
  • Partner with Incident and Problem Management to accelerate triage, reduce MTTR, and drive durable RCAs and prevention actions.
  • Integrate observability with CI/CD, feature flags, incident tooling, CMDB/service catalog, Slack, and Zoom
  • Coach product and platform teams on instrumentation patterns, trace context, and SLO thinking
  • Contribute reusable modules and templates
  • Lead telemetry hygiene, monitoring platform cost/usage optimization, and performance tuning
  • Ensure monitoring data is handled according to policy and implement role-based access and guardrails for sensitive logs and metrics
  • Participate in an on-call rotation with follow-the-sun support

Skills

Python
SQL
Monitoring
Logging
Anomaly Detection
SLI/SLO
Telemetry Insights
Performance Tuning

Tools

OpenTelemetry
CI/CD
Slack
Zoom
CMDB
Service Catalog
Datadog

Job description

Responsibilities
  • Design and operate the telemetry backbone for internal platforms and business-critical applications
  • Define and implement standards for metrics, logs, traces, and profiling using tools such as OpenTelemetry collectors and exporters
  • Establish golden signals, SLIs/SLOs, and health checks for priority services
  • Automate baselining and anomaly detection
  • Create executive and on-call dashboards, service health views, and dependency maps
  • Develop alerting policy as code and reduce false positives through suppression and deduplication
  • Implement auto-remediation runbooks
  • Partner with Incident and Problem Management to accelerate triage, reduce MTTR, and drive durable RCAs and prevention actions
  • Integrate observability with CI/CD, feature flags, incident tooling, CMDB/service catalog, Slack, and Zoom
  • Coach product and platform teams on instrumentation patterns, trace context, and SLO thinking
  • Contribute reusable modules and templates
  • Lead telemetry hygiene, monitoring platform cost/usage optimization, and performance tuning
  • Ensure monitoring data is handled according to policy and implement role-based access and guardrails for sensitive logs and metrics
  • Participate in an on-call rotation with follow-the-sun support
Requirements
  • 4–8 years in Observability/SRE/Platform/Monitoring roles supporting SaaS or enterprise applications
  • Hands-on experience with monitoring and logging tools
  • Strong grasp of distributed tracing, RED/USE/golden signals, SLI/SLO/SLA, and error budgets
  • Proficiency in common languages such as Python
  • Experience with AWS
  • Comfortable operating within Incident/Problem/Change frameworks
  • Ability to create runbooks, RCAs, and post-incident reviews
  • SQL or log query languages
  • Ability to translate telemetry into insights and narratives
  • Clear communication and collaboration skills; calm during outages
  • Service maps/dependency modeling, synthetic/RUM design, APM transaction tuning, and log schema governance are nice to have
  • Experience integrating observability with CMDB/service catalog and feature flag systems is nice to have
  • Certifications such as AWS or Datadog are nice to have
  • Must be physically located and plan to work from Maharashtra; the requisition location is Pune, India
  • Must attend the local office for part of the week
  • Core hours aligned to IST, with occasional off-hours participation for major incidents or change windows
Core Competencies

Demonstrates expertise in designing and operating telemetry systems for business-critical applications, with a strong focus on observability, incident management, and automation. Proficient in implementing monitoring standards and translating telemetry data into actionable insights.

Highest-signal resume keywords
  • Observability Engineering
  • Distributed Tracing
  • AWS Proficiency
  • Runbook Creation
  • Incident Management
Hard Skills
  • Python
  • SQL
  • Monitoring Tools
  • Logging Tools
  • Anomaly Detection
  • Metrics Standards
  • Service Level Indicators
  • Service Level Objectives
  • Telemetry Insights
  • Performance Tuning
Soft Skills
  • Clear Communication
  • Collaboration
  • Calm Under Pressure
Certifications & Qualifications
  • AWS Certification
  • Datadog Certification
Industry Keywords
  • SaaS Applications
  • Monitoring Frameworks
  • Incident Response
  • Change Management
  • Telemetry Hygiene
Tools & Technologies
  • OpenTelemetry
  • CI/CD Integration
  • Slack
  • Zoom
  • CMDB
  • Service Catalog
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Pune District

On-site
INR 2,250,000 - 2,750,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
AI Observability & Monitoring Engineer
AI Observability & Monitoring Engineer

Elabs Infotech • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Observability Engineer
Observability Engineer

Coforge • Hyderabad, Pune District, Greater Noida

Hybrid
INR 2,400,000 - 4,200,000
Lead Engineer Observability Integration
Lead Engineer Observability Integration

Hapag-Lloyd • Chennai District

On-site
INR 3,500,000 - 6,000,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
SRE Expert
SRE Expert

HCLTech • Bengaluru

On-site
INR 1,500,000 - 2,400,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Kiya.ai - Observability Integration Lead
Kiya.ai - Observability Integration Lead

Infrasoft Technologies Ltd • Mumbai

On-site
INR 1,200,000 - 1,600,000
SRE - Site Reliability Engineering
SRE - Site Reliability Engineering

Build & Hire • Pune District

On-site
INR 1,500,000 - 2,300,000