Role Description
We are seeking a hands-on Application Support Engineer with a strong focus on Observability Engineering and Platform Reliability to join the Technology Operations team.
This role is responsible for designing, implementing, and optimizing observability solutions across cloud-based applications and infrastructure. The ideal candidate will bring experience with Splunk, Google Analytics (or Google Analytics 4), and/or Bindplane, along with a strong foundation in AWS, automation, and Infrastructure as Code (IaC).
This role combines application support, observability engineering, and automation, with a focus on improving system visibility, telemetry quality, ing accuracy, and operational reliability.
Key Responsibilities
- Observability & Monitoring (PRIMARY FOCUS)
Design, implement, and maintain observability solutions using tools such as Splunk, Google Analytics, Bindplane, and Cloud-native monitoring tools
- Develop and optimize logging, metrics, and tracing strategies across applications and infrastructure
- Build and maintain dashboards, s, and anomaly detection mechanisms to proactively identify system issues
- Integrate telemetry data across multiple sources to improve end-to-end system visibility
- Improve signal-to-noise ratio in ing and reduce fatigue through tuning and correlation
Application & Production Support
- Troubleshoot issues across application, infrastructure, and integration layers in production and non-production environments
- Support application health monitoring and drive improvements in system reliability and performance
- Participate in incident response and root cause analysis using observability data
Automation & Platform Engineering
- Build and enhance automation using Ansible and scripting (Python/Bash) for observability deployment and management
- Implement observability components using Infrastructure as Code (Terraform, CloudFormation, etc.)
- Standardize observability patterns and reusable components across supported platforms
Cloud & Integration
- Support applications hosted in AWS environments, including integration with telemetry and monitoring platforms
- Collaborate with engineering teams to instrument applications for better observability (logs, metrics, traces)
- Support third-party enterprise platforms (SAP, Oracle, Axway, Qlik) with observability integrations
Operational Excellence
- Maintain systems aligned with N‑1 patching standards
- Document observability patterns, dashboards, ing strategies, and operational procedures
- Contribute to continuous improvement initiatives focused on reducing MTTR and improving system insight