Sr. Application Observability Engineer
Location: Pune Employment Type: full-time Designation: Sr. Application Observability Engineer
Job Description — Senior Application Observability Engineer
Position
Senior Application Observability Engineer
Job Summary
The Senior Application Observability Engineer is the technical lead responsible for building and advancing enterprise observability capabilities across applications, APIs, integrations, and cloud platforms. This role owns the observability strategy, with primary responsibility for Dynatrace, ensuring end-to-end visibility into application performance, user experience, infrastructure health, and business transactions.
Working closely with Application Support, Engineering, SRE, Infrastructure, Security, and Product teams, the engineer designs proactive monitoring, intelligent alerting, operational dashboards, and performance analytics to reduce incident impact, improve system reliability, and enhance operational maturity.
Key Responsibilities
Observability Strategy & Engineering
- Own and evolve the enterprise application observability strategy using Dynatrace.
- Design and implement end-to-end monitoring for applications, APIs, databases, integrations, and cloud workloads.
- Configure OneAgent, synthetic monitoring, real user monitoring (RUM), distributed tracing, and business transaction monitoring.
- Define monitoring standards, tagging strategy, naming conventions, and governance across business applications.
Monitoring & Alerting
- Build actionable dashboards and executive operational views for engineering and business stakeholders.
- Develop intelligent alerting policies to minimize false positives and improve incident detection.
- Continuously optimize monitoring coverage, thresholds, and anomaly detection.
Performance & Reliability
- Analyse application performance bottlenecks using traces, metrics, logs, and dependency mapping.
- Lead root cause analysis for critical incidents and recurring problems.
- Identify reliability improvement opportunities and recommend preventive solutions.
Collaboration & Operational Excellence
- Partner with Engineering Operations (EOG), Development, Infrastructure, SRE, Security, and Product teams during incident response.
- Support application onboarding into Dynatrace and establish observability best practices across Lines of Business (LOBs).
- Drive adoption of observability capabilities through documentation, knowledge sharing, and technical mentoring.
Required Qualifications
- Bachelor’s degree in computer science, Information Technology, Engineering, or related discipline.
- 8–10 years of experience in Application Performance Monitoring (APM), Observability, SRE, or Production Engineering.
- 3+ years of hands-on Dynatrace experience in enterprise environments.
- Strong understanding of application monitoring, distributed tracing, metrics, logs, and real user monitoring.
- Experience with cloud-based applications and microservices architectures.
- Hands‑on experience designing dashboards, alerts, and operational monitoring solutions.
- Experience performing incident analysis, root cause analysis, and performance optimization.
- Excellent analytical, troubleshooting, and stakeholder communication skills.
Preferred Skills
- Dynatrace OneAgent administration
- Synthetic Monitoring & RUM
- Azure or AWS cloud monitoring
- Grafana, Prometheus, Splunk, ELK, or App Insights
- CI/CD integration for observability
- ITIL Incident & Problem Management
Good to Have
- Dynatrace Professional or Associate Certification
- Site Reliability Engineering (SRE) practices
- Automation using PowerShell, Python, or Bash
- API monitoring and REST services
- ServiceNow integration for incident automation
Systems Plus offers GCC capabilities to help organizations enhance their agility, access specialized talent, and optimize costs while driving innovation. With over 20 delivered GCCs and more than three decades of experience in providing disruptive solutions, Systems Plus helps your organization become a global value organization. At Systems Plus, we provide access to a global talent pool, standardize processes, and enhance collaboration by providing a central hub for shared services and support.