Operations Specialist

OEC

Chennai District

On-site

INR 800,000 - 1,300,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OEC is seeking an Operations Specialist (L2) to join the Monitoring Team within the Enterprise Operations Group. You will own L2 alert triage, incident response, dashboard development, and service onboarding, contributing to tooling standards and best practices while keeping global platforms running reliably.

You will work with Datadog as the primary observability stack, perform day-to-day server monitoring and infrastructure support, and collaborate with DevOps, Cloud, and Engineering to ensure

Qualifications

  • A bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline is required.
  • ,

Responsibilities

  • Design, configure, and maintain Datadog monitors, composite alerts, and notification channels for production and non-production environments.
  • Own the L2 alert triage process — investigate, diagnose, and resolve escalated alerts from L1 with root cause identification.
  • Tune alert thresholds and suppression rules to minimise noise and reduce false positives.
  • Manage SLO definitions, tracking, and reporting across assigned services.
  • Participate in on-call rotation for critical monitoring alerts and P1/P2 incident response.
  • Build and maintain Datadog dashboards covering infrastructure health, application performance, log analytics, and business KPIs.
  • Configure and manage Datadog APM, distributed tracing, and error tracking for key services.
  • Implement and manage log ingestion pipelines, parsing rules, and log-based monitors.

Skills

Datadog monitoring
Alert triage
Dashboard development
SLO management
Incident response
Scripting (Python/Bash/PowerShell)
Cloud infrastructure (AWS)
Kubernetes/Docker
On-call coordination

Education

Bachelor's degree in Computer Science/IT/Engineering

Tools

Datadog
Kubernetes
Docker
Prometheus
Grafana
Jira/ServiceNow

Job description

OEC is in the middle of a multi-year journey to modernise its infrastructure and applications, transforming how we build, deploy, and operate software at global scale. The Operations Specialist (L2) is a core member of the Monitoring Team within the Enterprise Operations Team, responsible for the configuration, maintenance, and continuous improvement of OEC's observability stack, primarily powered by Datadog.

In this role you will own L2 alert triage, incident response, dashboard development, and service onboarding - contributing to the team's tooling standards and best practices. You will also handle day-to-day server monitoring, infrastructure support, and service management tasks that keep OEC's global platforms running reliably.

Key Responsibilities & Duties (essential to the job)
  • Design, configure, and maintain Datadog monitors, composite alerts, and notification channels for production and non-production environments.
  • Own the L2 alert triage process — investigate, diagnose, and resolve escalated alerts from L1, ensuring timely root cause identification.
  • Tune alert thresholds and suppression rules to minimise noise and reduce false positive rates.
  • Manage SLO (Service Level Objective) definitions, tracking, and reporting across assigned services.
  • Participate in the on-call rotation for critical monitoring alerts and P1/P2 incident response.
  • Build and maintain Datadog dashboards covering infrastructure health, application performance, log analytics, and business KPIs.
  • Configure and manage Datadog APM (Application Performance Monitoring), distributed tracing, and error tracking for key services.
  • Implement and manage log ingestion pipelines, parsing rules, and log-based monitors.
Service onboarding & integrations
  • Onboard new services and infrastructure components into the Datadog monitoring framework in collaboration with DevOps, Cloud, and Engineering teams.
  • Configure and maintain Datadog integrations with cloud platforms (AWS, Azure, Rackspace), Kubernetes, containerised workloads, and third-party tools.
  • Support the migration of legacy monitoring tooling (SCOM, Dynatrace, Grafana, Prometheus, Pingdom, Apigee, Redgate, IDERA) into Datadog.
Incident response & runbooks
  • Act as L2 escalation point during incidents — perform deep-dive investigations using metrics, logs, traces, and dashboards.
  • Create, maintain, and improve runbooks for common alert scenarios, incident response procedures, and post-incident remediation steps.
  • Contribute to post-incident reviews, documenting root cause findings and identifying monitoring gaps to prevent recurrence.
  • Alerts system owners, stakeholders, and management of degraded system status and Priority 1 and Priority 2 incidents; issues tickets for incident and problem escalation.
  • Leverage Datadog AI capabilities including Watchdog, anomaly detection, outlier detection, and Bits AI to improve proactive issue detection.
  • Implement forecast-based monitors for capacity management (CPU, memory, disk).
  • Use AI-based alert correlation and deduplication features to reduce alert fatigue.
  • Contribute to automation of repetitive monitoring tasks using Datadog’s API and scripting tools.
Standards & governance
  • Adhere to and actively contribute to monitoring standards, tagging strategies, and naming conventions.
  • Participate in regular alert quality reviews, dashboard audits, and SLO compliance checks.
  • Maintain accurate documentation of monitoring configurations, integrations, and architectural decisions.
  • Creates and maintains knowledge articles to be used internally and/or externally for training, best practices, solutions, or processes relating to applications, environments, and related technologies.
Infrastructure operations & support
  • Analyses and troubleshoots Microsoft and UNIX/Linux server configurations and processes.
  • Performs moderately complex database administration.
  • Monitors data centre networks, infrastructure bandwidth, application health, servers, and other infrastructure; coordinates and communicates with vendors.
  • Diagnoses and researches (using knowledge base) application incidents, monitoring alerts, and service requests; provides assistance and guidance to associate operations specialists.
  • Adheres to all incident and service request processes and procedures in accordance with established Service Level Agreements (SLAs).
  • Diagnoses, researches, and resolves Level-2 technical hardware and software incidents, monitoring alerts, and service requests.
  • Works on ad-hoc projects to support Infrastructure Engineering or other departments, as requested.
  • Demonstrates a flexible and adaptable approach to work and adjusts to shifts in priorities as the needs of the business change.
  • Collaborates with DevOps, Cloud, and application teams to align monitoring coverage with service requirements.
Education

A bachelor's degree from an accredited college or university in Computer Science, Information Technology, Engineering, or a related discipline is required. In the absence of a degree, equivalent work experience directly related to the key responsibilities of the role will be considered as a substitute for the degree.

Experience, Skills and Key Competencies

At least 3-6 years of experience in an infrastructure, DevOps, SRE, or monitoring engineering role is required, with hands-on experience in Datadog and working knowledge of cloud infrastructure and containerised environments.

Must also be able to demonstrate the following skills and abilities:

Technical skills
  • Hands-on experience with Datadog - monitors, dashboards, APM, log management, SLOs, and synthetic monitoring.
  • Solid understanding of cloud infrastructure (AWS required; Azure or GCP a plus) and containerised environments (Docker, Kubernetes).
  • Familiarity with observability concepts: metrics, traces, and logs.
  • Experience with at least one scripting language (Python, Bash, or PowerShell) for automation and tooling.
  • Working knowledge of monitoring and alerting tools such as Prometheus, Grafana, Dynatrace, or similar platforms.
  • Understanding of ITSM concepts and ticketing workflows (Jira, JSM, or ServiceNow).
  • Good working knowledge of Microsoft and UNIX/Linux server configurations, Active Directory, and internet protocols (HTTP, FTP, DNS, IP, SSH).
  • Strong analytical and problem-solving skills with the ability to diagnose complex infrastructure issues under pressure.
  • Clear written and verbal communication skills; able to produce runbooks and incident summaries for both technical and non-technical audiences.
  • Collaborative team player who can work effectively across Engineering, DevOps, and Operations disciplines.
  • Self-motivated with a continuous improvement mindset and a proactive approach to identifying and resolving issues before they elevate.
  • Flexible and adaptable approach to work; able to adjust to shifts in priorities as the needs of the business change.
Preferred Qualifications
  • Datadog Fundamentals or Datadog APM certification (or actively working towards one).
  • At least 1 year of experience with CI/CD pipelines (Jenkins, GitHub Actions, or Azure DevOps).
  • Knowledge of network monitoring and distributed systems architecture.
  • Exposure to ITIL or ITSM frameworks (Incident, Problem, and Change Management).
  • Familiarity with Atlassian tools (Jira, JSM, Confluence) for documentation and ticket management.
Special Position Requirements
  • Must be able to read, write, understand, and speak fluent English.
  • Must be able to work flexible shifts including holidays and weekends, to provide support across locations and time zones.
  • Must be able to participate in an on-call rotation for critical monitoring alerts and P1/P2 incident response.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Operations Specialist
Operations Specialist

OEConnection LLC • Chennai District

On-site
INR 1,800,000 - 2,500,000
Datadog Monitoring specialist
Datadog Monitoring specialist

Genpact • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Datadog Engineer
Datadog Engineer

Keylabsintelli • Hyderabad

On-site
INR 800,000 - 1,200,000
Software Engineer
Software Engineer

algoleap • Hyderabad

On-site
INR 1,200,000 - 2,400,000
DevOps Engineer
DevOps Engineer

Anblicks Inc. • Hyderabad

On-site
INR 1,200,000 - 2,000,000
Production and Support Engineer
Production and Support Engineer

Devon Software Services • Bengaluru

On-site
INR 800,000 - 1,200,000
Datadog Administration and Operations (Servicenow)
Datadog Administration and Operations (Servicenow)

HP • Bengaluru

On-site
INR 2,500,000 - 3,500,000
Software Engineer
Software Engineer

Algoleap Technologies • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Datadog
Datadog

Randstad Digital • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Network Admin/Engineer
Network Admin/Engineer

Ascendion • Bengaluru

On-site
INR 1,000,000 - 1,500,000