Experience: 10+ years
Extensive experience with DataDog (RUM/Synthetic/Log/APM) like installing, configuring, alert and dashboard
Responsibilities
1. System and Network Monitoring
- Continuously monitor servers, networks, applications, and databases to ensure availability and performance.
- Utilize monitoring tools to track system health, resource usage, and network traffic patterns.
- Set up alerting mechanisms to notify the relevant teams of performance issues or outages.
- Monitor and proactively manage operational alerts, incidents, and infrastructure issues, promptly addressing problems to maintain optimal uptime and service reliability.
2. Incident Detection and Response
- Identify and analyze incidents reported by monitoring tools or users.
- Respond to alerts promptly and coordinate with IT support to elevate and resolve issues.
- Document incidents, including root causes and remedies, to help inform future prevention strategies.
- Actively manage and coordinate incident responses using text message and ServiceNow.
3. Performance Analysis and Reporting
- Generate reports on system performance, availability and incidents to provide insights to management.
- Analyze trends to predict potential issues and advise on capacity planning and resource allocation.
- Monitor adherence to service level agreements (SLAs) and report on compliance.
- Establish the Service Level Objectives (SLO) and Service Level Indicators (SLI) to enhance application performance.
4. Tool Management and Optimization
- Maintain and optimize monitoring tools and systems to ensure effective performance.
- Implement new monitoring solutions as needed to enhance coverage and functionality.
- Stay updated on the latest monitoring technologies to improve monitoring capabilities.
- Automation and Optimization
- Maintain and ensure consistent availability and performance of Datadog observability solutions, including Application Performance Monitoring (APM), Infrastructure Monitoring, Log Management, Synthetic Monitoring, and Real User Monitoring (RUM).
- Utilize Terraform to manage infrastructure efficiently and securely, ensuring stability and compliance.
- Eliminate manual effort by leveraging PowerShell and Python.
- Observability coverage specifically to Kubernetes and Docker, ensuring comprehensive monitoring and performance optimization of containerized environments.
- Create, update, and manage Datadog monitors and dashboards based on incoming requests, ensuring they accurately reflect operational requirements and KPIs.
- Customize and fine‑tune monitoring solutions to meet specific team or application needs, enhancing visibility and system health awareness.
5. Collaboration and Communication
- Work collaboratively with IT teams to understand monitoring requirements and expectations.
- Communicate findings and performance metrics effectively to stakeholders, ensuring clarity and understanding.
- Provide knowledge sharing sessions to educate teams on monitoring practices and tool usage.
Qualifications
Minimum Qualifications
- Bachelor's/Master's in CS/IT or equivalent practical experience.
Preferred Qualifications
- Stay updated on the latest monitoring technologies to improve monitoring capabilities.
- Automation and Optimization
- Maintain and ensure consistent availability and performance of Datadog observability solutions, including Application Performance Monitoring (APM), Infrastructure Monitoring, Log Management, Synthetic Monitoring, and Real User Monitoring (RUM).