Observability Engineer

Persistent Systems Limited

Bengaluru

Hybrid

INR 1,200,000 - 2,400,000

Full time

25 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid work
Flexible hours
Long Service awards
Group term life insurance
Mediclaim hospitalization

Job summary

Persistent Systems Limited is seeking an Observability Engineer to design, implement, and maintain scalable observability across enterprise platforms. You will build dashboards, set alerts, and drive incident management with a focus on reliability and security.

The role requires 5–8 years of SRE/observability experience, cloud exposure, and strong scripting skills. The position offers hybrid work arrangements, collaboration across engineering and security teams, and opportunities to advance

Qualifications

  • 5–8 years of experience in Observability, SRE, Production Support, Cloud Operations, or Platform Engineering.
  • Strong hands-on expertise in IAM, SRE practices, and Grafana.
  • Experience with monitoring, logging, alerting, and observability platforms.
  • Strong understanding of application and infrastructure monitoring concepts.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Automation and scripting using Shell, Python, or PowerShell.

Responsibilities

  • Design, implement, and maintain enterprise observability solutions across infrastructure, platforms, and applications.
  • Develop and manage dashboards, alerts, monitoring policies, and service health metrics using Grafana.
  • Monitor system performance, availability, and reliability to ensure business continuity.
  • Investigate production incidents, perform root cause analysis, and drive timely resolution.
  • Create and maintain runbooks, deployment guides, and standard operating procedures.
  • Collaborate with cross-functional teams to improve platform resilience and security.
  • Automate repetitive tasks to improve efficiency and reduce manual work.
  • Support change management, release activities, and production readiness.

Skills

IAM
SRE practices
Grafana
Monitoring
Logging
Alerting
Observability platforms
Cloud platforms
Automation scripting
Shell scripting
Python
PowerShell
CI/CD
DevOps
Incident management

Tools

Grafana

Job description

We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what’s next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem.

We are seeking an experienced Observability Engineer to help build and maintain highly reliable, secure, and scalable enterprise platforms. This role is ideal for professionals with 5–8 years of experience in Site Reliability Engineering (SRE), observability, monitoring, and platform operations. The successful candidate will play a key role in enhancing platform reliability, monitoring application and infrastructure health, automating operational processes, and ensuring superior customer experience through proactive incident management and continuous improvement initiatives.

  • Role: Observability Engineer
  • Experience: 5 to 8 Years
  • Job Type: Full-Time Employment
What You'll Do:
  • Design, implement, and maintain enterprise observability solutions across infrastructure, platforms, and applications.
  • Develop and manage dashboards, alerts, monitoring policies, and service health metrics using Grafana and related tools.
  • Monitor system performance, availability, and reliability to ensure seamless business operations.
  • Investigate production incidents, perform root cause analysis, and drive timely issue resolution.
  • Create and maintain operational runbooks, support documentation, and standard operating procedures.
  • Collaborate with engineering, cloud, infrastructure, security, and application teams to improve platform resilience.
  • Implement automation solutions to reduce manual tasks and improve operational efficiency.
  • Drive reliability engineering initiatives focused on availability, scalability, and performance optimization.
  • Apply IAM policies, access governance controls, and security best practices across supported environments.
  • Support change management, release activities, and production readiness assessments.
  • Analyze service trends and proactively identify opportunities for performance and reliability improvements.
  • Participate in incident response, problem management, and service improvement initiatives.
  • Ensure compliance with enterprise security, governance, and operational standards.
Expertise You'll Bring:
  • 5–8 years of experience in Observability, Site Reliability Engineering, Production Support, Cloud Operations, or Platform Engineering.
  • Strong hands-on expertise in IAM, SRE practices, and Grafana.
  • Experience with monitoring, logging, alerting, and observability platforms.
  • Strong understanding of application performance monitoring and infrastructure monitoring concepts.
  • Experience investigating production issues, conducting root cause analysis, and implementing preventive measures.
  • Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Experience with automation and scripting using Shell Scripting, Python, PowerShell, or similar technologies.
  • Understanding of incident management, problem management, and IT service management processes.
  • Experience with DevOps practices, CI/CD pipelines, and operational excellence methodologies.
  • Knowledge of security, access management, compliance, and governance requirements.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Experience working in Agile and cross-functional delivery environments.
  • Excellent communication, stakeholder management, and documentation skills.
  • Ability to work independently while managing multiple priorities in a fast-paced environment.
  • Strong ownership mindset with a focus on customer experience, reliability, and operational excellence.
  • Competitive salary and benefits package
  • Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
  • Opportunity to work with cutting-edge technologies
  • Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
  • Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents
Values-Driven, People-Centric & Inclusive Work Environment:

Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.

  • We support hybrid work and flexible hours to fit diverse lifestyles.
  • Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
  • If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment

“Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.”

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Observability Engineer
Observability Engineer

Persistent Systems Limited • Pune District

Hybrid
INR 1,800,000 - 3,000,000
Hybrid work
Company-sponsored education
Employee engagement events
+1
SRE Observability Engineer
SRE Observability Engineer

Persistent Systems • Pune District

On-site
INR 1,500,000 - 2,000,000
Competitive salary and benefits package
Culture focused on talent development
Quarterly growth opportunities
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Persistent Systems Limited • Pune District

On-site
INR 2,500,000 - 4,200,000
Competitive salary
Benefits package
Talent development
+4
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Persistent Systems Limited • Pune District

On-site
INR 1,400,000 - 2,200,000
Hybrid work
Long Service awards
Company-sponsored education
Lead Engineer - Observability Engineering
Lead Engineer - Observability Engineering

Levi Strauss • Bengaluru

On-site
INR 4,000,000 - 6,500,000
Health check-up and OPD coverage
Best-in-class leave plan
Mental well-being support
+1
Site Reliability Engineer (SRE) / Observability Engineer
Site Reliability Engineer (SRE) / Observability Engineer

N Human Resources & Management Systems • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Hybrid work
Certification reimbursement
Structured learning
Site Reliability Engineer
Site Reliability Engineer

Persistent Systems • Hyderabad

On-site
INR 600,000 - 1,800,000
Competitive salary
Quarterly promotion cycles
Company-sponsored education and certifications
+3
Site Reliability Engineer (GCP)
Site Reliability Engineer (GCP)

Persistent Systems Limited • Hyderabad

Hybrid
INR 4,200,000 - 6,600,000
Hybrid work
Flexible hours
Education sponsorship
+3
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Principal Embedded Observability & Telemetry Architect
Principal Embedded Observability & Telemetry Architect

Mployee.me • Pune District

On-site
INR 4,000,000 - 6,000,000
Competitive salary
Hybrid work option
Education sponsorship
+2