We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what’s next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem.
We are seeking an experienced Observability Engineer to help build and maintain highly reliable, secure, and scalable enterprise platforms. This role is ideal for professionals with 5–8 years of experience in Site Reliability Engineering (SRE), observability, monitoring, and platform operations. The successful candidate will play a key role in enhancing platform reliability, monitoring application and infrastructure health, automating operational processes, and ensuring superior customer experience through proactive incident management and continuous improvement initiatives.
- Role: Observability Engineer
- Experience: 5 to 8 Years
- Job Type: Full-Time Employment
What You'll Do:
- Design, implement, and maintain enterprise observability solutions across infrastructure, platforms, and applications.
- Develop and manage dashboards, alerts, monitoring policies, and service health metrics using Grafana and related tools.
- Monitor system performance, availability, and reliability to ensure seamless business operations.
- Investigate production incidents, perform root cause analysis, and drive timely issue resolution.
- Create and maintain operational runbooks, support documentation, and standard operating procedures.
- Collaborate with engineering, cloud, infrastructure, security, and application teams to improve platform resilience.
- Implement automation solutions to reduce manual tasks and improve operational efficiency.
- Drive reliability engineering initiatives focused on availability, scalability, and performance optimization.
- Apply IAM policies, access governance controls, and security best practices across supported environments.
- Support change management, release activities, and production readiness assessments.
- Analyze service trends and proactively identify opportunities for performance and reliability improvements.
- Participate in incident response, problem management, and service improvement initiatives.
- Ensure compliance with enterprise security, governance, and operational standards.
Expertise You'll Bring:
- 5–8 years of experience in Observability, Site Reliability Engineering, Production Support, Cloud Operations, or Platform Engineering.
- Strong hands-on expertise in IAM, SRE practices, and Grafana.
- Experience with monitoring, logging, alerting, and observability platforms.
- Strong understanding of application performance monitoring and infrastructure monitoring concepts.
- Experience investigating production issues, conducting root cause analysis, and implementing preventive measures.
- Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Experience with automation and scripting using Shell Scripting, Python, PowerShell, or similar technologies.
- Understanding of incident management, problem management, and IT service management processes.
- Experience with DevOps practices, CI/CD pipelines, and operational excellence methodologies.
- Knowledge of security, access management, compliance, and governance requirements.
- Strong analytical, troubleshooting, and problem-solving skills.
- Experience working in Agile and cross-functional delivery environments.
- Excellent communication, stakeholder management, and documentation skills.
- Ability to work independently while managing multiple priorities in a fast-paced environment.
- Strong ownership mindset with a focus on customer experience, reliability, and operational excellence.
- Competitive salary and benefits package
- Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
- Opportunity to work with cutting-edge technologies
- Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
- Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents
Values-Driven, People-Centric & Inclusive Work Environment:
Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.
- We support hybrid work and flexible hours to fit diverse lifestyles.
- Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
- If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment
“Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.”