Observability Engineer

Weekday (YC W21)

Chennai District

On-site

INR 3,500,000 - 7,000,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TalentBridge is seeking a Senior Observability Engineer to design, implement, and scale enterprise-grade observability platforms. You will drive visibility across distributed systems and lead migration from legacy monitoring to modern cloud-native solutions.

Join a collaborative engineering team, shaping best practices for metrics, logs, traces, and alerting. Strong OpenTelemetry, Grafana, Prometheus, OpenShift/Kubernetes, and Helm expertise are required.

Qualifications

  • Must have hands-on experience with observability tooling and architectures.
  • Strong experience designing and deploying end-to-end observability stacks.
  • Proficiency with Kubernetes/OpenShift and cloud-native monitoring.

Responsibilities

  • Design, deploy, and scale end-to-end observability solutions across enterprise environments.
  • Lead migration from legacy monitoring to modern observability stacks.
  • Build dashboards, alerts, and standards to improve incident response.

Skills

OpenTelemetry
Grafana Enterprise
Prometheus/PromQL
Grafana dashboards
OpenShift/Kubernetes
Helm charts
Observability platforms
SRE/DevOps

Education

Bachelor's or Master's degree

Tools

ITRS Geneos
Mimir
Loki
Tempo
Grafana

Job description

This role is for one of our clients

Company Name: TalentBridge

Seniority level: Mid-Senior level

Min Experience: 6+ years

Location: Chennai, Bangalore, Hyderabad, Mumbai, Pune

JobType: full-time

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

We are looking for a highly skilled Senior Observability Engineer to design, implement, and scale enterprise‑grade observability platforms that provide deep visibility into distributed systems. This role is ideal for professionals passionate about improving system reliability, performance, and operational excellence through modern observability practices.

As a key member of the engineering team, you will lead the design and deployment of comprehensive monitoring, logging, and tracing solutions while driving the migration from traditional monitoring platforms to a modern observability ecosystem. You will collaborate closely with platform, infrastructure, and application teams to establish best practices and improve the overall health and resilience of mission‑critical systems.

Key Responsibilities
  • Design, develop, and manage scalable end-to-end observability solutions covering metrics, logs, traces, and alerting across enterprise environments
  • Lead the migration from legacy monitoring platforms to modern observability frameworks and cloud‑native monitoring solutions
  • Deploy, administer, and optimize observability platforms running on Kubernetes or OpenShift environments
  • Build reusable dashboards, alerts, and monitoring standards to improve operational visibility and incident response
  • Develop and maintain Helm charts for deployment and lifecycle management of observability components
  • Implement automation for deployment, configuration management, and operational workflows using Python or Bash scripting
  • Collaborate with engineering and application teams to define observability standards and integrate monitoring into development workflows
  • Analyze system performance, identify bottlenecks, and recommend improvements that enhance platform reliability and scalability
  • Provide technical leadership, architectural guidance, and strategic recommendations for observability initiatives
  • Support production operations by troubleshooting complex monitoring and infrastructure issues
  • Contribute to continuous improvement initiatives and drive adoption of observability best practices across engineering teams
Must‑Have Skills
  • Strong hands‑on experience with OpenTelemetry for instrumentation and telemetry collection
  • Expertise in the Grafana Enterprise Stack, including Mimir, Loki, and Tempo
  • Experience administering and scaling ITRS Geneos in enterprise environments
  • Strong knowledge of Prometheus and PromQL
  • Hands‑on experience with Grafana, including dashboard creation, alerting, and data source management
  • Experience administering OpenShift or Kubernetes clusters
  • Expertise in developing and managing Helm Charts for Kubernetes deployments
  • Experience designing, deploying, and scaling enterprise observability platforms
Good‑to‑Have Skills
  • Experience with Google Cloud Observability or other cloud‑native monitoring solutions
  • Automation and scripting experience using Python or Bash
  • Familiarity with CI/CD pipelines and enterprise deployment processes
  • Experience implementing observability in cloud‑native or microservices‑based architectures
Preferred Qualifications
  • Bachelor's or Master's degree in Computer Science, Information Technology, or a related field
  • 8-15 years of experience in Platform Engineering, Site Reliability Engineering (SRE), DevOps, Infrastructure Engineering, or Observability Engineering
  • Proven experience implementing observability solutions at enterprise scale
  • Strong understanding of distributed systems, container platforms, and cloud‑native technologies
Soft Skills
  • Strong analytical and problem‑solving abilities
  • Strategic thinking with the ability to influence technical direction
  • Excellent communication and stakeholder management skills
  • Ability to collaborate effectively across cross‑functional teams
  • Strong leadership, mentoring, and relationship‑building capabilities
  • Service‑oriented mindset with a focus on operational excellence
  • Ability to manage multiple initiatives in a fast‑paced environment
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer
Observability Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Hyderabad

On-site
INR 4,200,000 - 7,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Mumbai

On-site
INR 4,000,000 - 6,000,000
AWS Observability Engineer
AWS Observability Engineer

Tata Consultancy Services • New Delhi, Dadri

On-site
INR 1,500,000 - 2,500,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Observability Engineer
Observability Engineer

Crest Data • Ahmedabad District

Hybrid
INR 2,000,000 - 3,600,000
Software Engineer -Observability
Software Engineer -Observability

1203 Barclays Global Serv. Cent • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Site Reliability Engineer (SRE) / Observability Engineer
Site Reliability Engineer (SRE) / Observability Engineer

N Human Resources & Management Systems • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work
Certification reimbursement
Structured learning
Observability Engineer (SRE)
Observability Engineer (SRE)

Cloudstepin • Hyderabad

On-site
INR 1,200,000 - 1,800,000