Observability Architect - Datadog

GSPANN

Hyderabad, Pune District, Gurugram District

On-site

INR 1,800,000 - 3,600,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GSPANN is seeking a Datadog Platform & Observability Architect with 10+ years of experience to design, implement, operate, and improve an enterprise observability platform. The role enables engineering and operations teams to collect telemetry, detect issues early, and manage observability governance and costs at scale.

Responsibilities include configuring monitors, dashboards, and integrations, and driving proactive issue detection in collaboration with engineering teams.

Qualifications

  • Strong hands-on experience with Datadog and observability tools.
  • Experience building dashboards, monitors, alerts, and integrations.
  • Experience with Infrastructure monitoring, APM, logs, and synthetic monitoring.

Responsibilities

  • Design, configure, and maintain Datadog monitors, dashboards, alerts, and integrations.
  • Implement observability across Infra, Application and Network layers.
  • Develop end-to-end dashboards and proactive alerts for services.
  • Integrate Datadog with ITSM and enterprise tools; enable monitoring-as-code.

Skills

Datadog
APM & Tracing
Log Management
Terraform
Kubernetes
Azure/AWS
SRE/Observability
Python/PowerShell/Bash

Tools

Terraform
Kubernetes
Azure
AWS
Datadog

Job description

Hello,

Greeting of the day

Role - Datadog Platform & Observability Architect
Experience - 10+ Years
Notice Period - Immediate to 30 Days
Role Summary

We are seeking a Datadog Platform & Observability architect to design, implement, operate, and continuously improve our enterprise observability platform. This role enables engineering and operations teams to collect reliable telemetry, detect issues early, troubleshoot efficiently collaborating with Engineering teams, and manage observability governance and costs at scale.

Key Responsibilities
  • Design, configure, and maintain Datadog monitors, dashboards, alerts, and integrations.
  • Should have complete understanding of observability to build around Infra, Application and Network
  • Should have exposure implementing Observability solutions for ERP and eComm applicaitons
  • Should be able to effectively utilize AI modules of Datadog for auto detection and prevention
  • Implement and support Infrastructure Monitoring, APM, Log Management, Synthetic Monitoring, RUM, and API monitoring.
  • Configure application and infrastructure telemetry, including metrics, logs, traces, and events.
  • Build and optimize Datadog log pipelines, parsing rules, indexes, retention filters, and enrichment.
  • Configure and troubleshoot Datadog Agents across Windows, Linux, cloud, containers, and Kubernetes environments.
  • Develop application-specific dashboards providing end-to-end visibility into application health and performance.
  • Create proactive alerts for critical applications, APIs, infrastructure components, and business services.
  • Review existing monitors and continuously improve alert thresholds, noise reduction, duplicate alerts, and actionable notifications.
  • Support application teams with troubleshooting using logs, traces, metrics, and dependency information.
  • Perform monitoring gap assessments and recommend improvements.
  • Integrate Datadog with ITSM and collaboration platforms such as Ivanti, ServiceNow, Teams, Slack, or other enterprise tools.
  • Support automation and event-driven incident creation/enrichment.
  • Work with DevOps/SRE teams to implement monitoring-as-code using Terraform or similar tools.
  • Participate in major incident troubleshooting and support RCA/problem management activities.
  • Create and maintain SOPs, monitoring standards, technical documentation, and knowledge articles.
  • Conduct knowledge-transfer sessions and help support teams effectively use Datadog for incident triage.
  • Track monitoring effectiveness and demonstrate measurable improvements such as reduced alert noise, faster detection, lower MTTR, and fewer repeat incidents.
Required Technical Skills
  • Strong hands-on experience with Datadog.
  • Datadog Infrastructure Monitoring and Agent configuration.
  • APM and Distributed Tracing.
  • Log Management and Log Pipelines.
  • Datadog dashboards, monitors, alerts, and notification configuration.
  • Synthetic/API monitoring.
  • Experience troubleshooting Windows and Linux systems.
  • Good understanding of cloud platforms, preferably Azure and/or AWS.
  • Knowledge of containers and Kubernetes monitoring.
  • Understanding of REST APIs, JSON, and integration concepts.
  • Scripting knowledge in Python, PowerShell, Bash, or similar languages.
  • Experience with Terraform or Infrastructure-as-Code is preferred.
  • Strong troubleshooting and root-cause analysis skills.
  • Datadog RUM and Digital Experience Monitoring experience.
  • Network monitoring experience.
  • Azure Monitor / Log Analytics knowledge.
  • Exposure to application performance optimization and capacity planning.
Good to Have
  • Workato or other integration/automation platform experience.
  • SRE/Observability implementation experience.
  • CI/CD monitoring experience.
  • ITIL knowledge covering Incident, Problem, and Change Management.
Key Expectations

The Architect should not only maintain existing monitoring but continuously identify opportunities to move operations from reactive support to proactive issue detection and prevention.

The person should be capable of independently identifying monitoring gaps, proposing solutions, implementing improvements, and working across multiple technical teams

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Datadog Engineer
Datadog Engineer

Keylabsintelli • Hyderabad

On-site
INR 800,000 - 1,200,000
Datadog Monitoring specialist
Datadog Monitoring specialist

Genpact • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Software Engineer
Software Engineer

algoleap • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Full Stack Observability Engineer
Full Stack Observability Engineer

Ikrux • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Datadog
Datadog

Randstad Digital • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Engineer(SRE)
Staff Engineer(SRE)

Nagarro • Gurugram District

Hybrid
INR 1,400,000 - 2,200,000
Operations Specialist
Operations Specialist

OEC • Chennai District

On-site
INR 800,000 - 1,300,000
Datadog Architect
Datadog Architect

GSPANN Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,100,000
Looking For Dynatrace / Datadog Engineer
Looking For Dynatrace / Datadog Engineer

Tech Mahindra • Hyderabad, Pune District

On-site
INR 1,500,000 - 2,100,000
Datadog Administration and Operations (Servicenow)
Datadog Administration and Operations (Servicenow)

HP • Bengaluru

On-site
INR 2,500,000 - 3,500,000