Digital Technology Senior Specialist - Observability & AI Ops

bakerhughes

Pune District, Mumbai

On-site

INR 2,200,000 - 4,200,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Baker Hughes is seeking a Senior Specialist for Observability & AI Ops to shape and implement our Digital Technology strategy. You will lead the deployment and operation of an Elastic Stack-based observability platform across cloud, on-premises, and hybrid environments.

Responsibilities include onboarding thousands of devices, building dashboards and alerts, automating ingestion pipelines, and collaborating with Cloud, Infra, and SRE teams to improve reliability and visibility.

Qualifications

  • Experience in enterprise-scale observability or SRE environments.
  • Proven ability to design and operate elastic-based observability platforms.
  • Strong scripting and automation capabilities to improve deployment and runbooks.

Responsibilities

  • Implement and manage the lifecycle of an enterprise observability platform based on Elastic Stack (Elasticsearch, Kibana, Logstash, Beats, Elastic APM, Fleet).
  • Onboard 20,000+ devices into the observability solution.
  • Onboard infrastructure, cloud services, applications, and platforms into observability and monitoring.
  • Create and maintain dashboards, visualizations, alerts, and operational reports.
  • Configure data collection (logs, metrics, traces, uptime, APM) across environments.
  • Assist with data ingestion, parsing, enrichment and retention policies.
  • Support incident investigation, troubleshooting, and root-cause analysis using observability data.
  • Collaborate with cloud, infra, app teams and SRE for reliability improvements.
  • Drive automation via scripting and IaC; contribute to upgrade and capacity planning.
  • Maintain runbooks and documentation for runbooks and consumption guides.

Skills

SRE/DevOps practices
Cloud operations
Automation scripting
Observability platforms
Kubernetes & Docker
Incident management

Tools

Elastic Stack
Kibana
Logstash
Beats
Elastic Agent
Fleet
Elastic APM
Grafana
Prometheus
Kubernetes
Docker

Job description

DT Senior Specialist - Observability & AI Ops

Would you like to help shape and implement our Digital Technology teams' strategic direction?

Are you passionate about helping improve observability and digital operations?

Join our Digital Technology team! We operate at the heart of Baker Hughes digital transformation journey. Our team delivers enterprise observability and AIOps capabilities that help technology teams detect issues earlier, troubleshoot faster, and improve service performance across cloud, infrastructure, and application environments.

Partner with the best

As an AIOps & Observability Engineer, you will support the implementation, enhancement, onboarding, and day-to-day operations of observability platforms, with a focus on Elastic Stack capabilities and practical SRE-driven operational outcomes.

As a Senior AI Ops Engineer, you will be responsible for:
  • Implement and manage the life cycle of enterprise observability platform based on elastic tech stacks [not limited to] Kibana, Logstash, Beats, Elastic Agent, Fleet, Elastic APM components, etc.
  • Onboard full suite of 20,000 plus devices under observability umbrella
  • Onboarding infrastructure, cloud services, applications, and platforms into observability and monitoring solutions.
  • Building and maintaining dashboards, visualizations, alerts, and operational reports for technology teams.
  • Configuring log, metric, trace, uptime, and APM data collection across supported environments.
  • Assisting with data ingestion, parsing, enrichment, and retention activities.
  • Supporting incident investigation, troubleshooting, and root cause analysis using observability data.
  • Collaborating with cloud, infrastructure, application, and SRE teams to improve system reliability and service visibility.
  • Contributing to automation initiatives using scripting, Infrastructure as Code, and repeatable deployment practices.
  • Participating in observability platform upgrades, patching, performance tuning, and operational support activities.
  • Creating and maintaining runbooks, knowledge articles, dashboards standards, and operational runbooks related documentation.
  • Contributing to continuous improvement of observability practices, monitoring coverage, and operational readiness.
Fuel your passion
  • Have 7+ years - SRE/DevOps experience in enterprise-scale or mission-critical environments
  • Have 5+ years - Cloud / Application / Platform operations and administration (AWS, Azure, hybrid or multi-cloud)
  • Have 5+ years - Automation, CI/CD, and scripting proficiency (Python, Bash, PowerShell, Ruby, or equivalent)
  • Have 5+ years - Exposure to containers and cloud-native platforms such as Docker, Kubernetes, Prometheus, or Grafana.
  • Have 3+ years - Proven experience administering Elastic Observability platforms across the full lifecycle, including deployment, maintenance, upgrades, patching, and capacity scaling.
Preferred qualifications
  • AWS or Azure Associate-level certification, or equivalent practical cloud operations experience.
  • Elastic Certified Engineer or equivalent observability platform certification
  • Familiarity with infrastructure as code (GitHub Actions, CloudFormation, Terraform, Ansible) for repeatable automation
  • Exposure to cloud-native observability frameworks (Open Telemetry, service meshes)
  • Experience documenting runbooks, playbooks, and consumption guides for SMEs
  • Process knowledge such as Agile and/or ITIL
Must have Technical Skills
  • Strong background in observability platforms ( Elastic.io stack preferred : Elasticsearch, Kibana, Logstash, Beats, Elastic APM, and Fleet/Elastic Agent)
  • Telemetry Fundamentals : Strong, practical understanding of fundamental observability concepts, including the collection and analysis of logs, metrics, traces, and synthetic monitoring.
  • Experience with administration of leading observability platforms (Grafana, Graylog, Splunk, Sumo Logic, Tanzu, or open-source equivalents) including lifecycle management ( Kubernetes, Docker, Prometheus, and Grafana installation, patching, upgrades, scaling).
  • Strong knowledge of distributed infrastructure domains (network, servers, VMs, AWS, Azure, databases) from an observability perspective.
  • Proven ability to design and tune scalable ingestion pipelines for diverse, globally distributed data sources.
  • Flexibility to adapt and evolve observability standards per domain needs while ensuring consistency across the enterprise.
  • Operational ownership mindset - accountable for uptime, reliability, and lifecycle management of the hosted observability platform.
  • Incident management and SRE practices: monitoring, alerting, troubleshooting, root cause analysis, and postmortems.
  • Proficiency in automation and scripting (Python, Bash, PowerShell, Ruby, etc.) for ingestion, upgrades, and operational tasks.
  • Familiarity with REST APIs and tools like Postman, plus DevOps constructs (GitHub, Jenkins, CI/CD pipelines, serverless technologies).
  • Configuration Proficiency : Demonstrated proficiency in managing system configurations usi
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Observability Engineer
Observability Engineer

Coforge • Hyderabad, Pune District, Greater Noida

Hybrid
INR 2,400,000 - 4,200,000
AI Observability & Monitoring Engineer
AI Observability & Monitoring Engineer

Elabs Infotech • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Observability Technical Lead
Observability Technical Lead

Jeevan Technologies • Chennai District

On-site
INR 3,000,000 - 6,000,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Senior Engineer / Lead / Architect
Senior Engineer / Lead / Architect

Tranzeal • Hyderabad, Chennai District, Bengaluru

On-site
INR 4,500,000 - 7,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Observability Engineer - Grafana & Prometheus
Senior Observability Engineer - Grafana & Prometheus

Zensar • Pune District, Bengaluru

Hybrid
INR 1,800,000 - 3,200,000
SRE Expert
SRE Expert

HCLTech • Bengaluru

On-site
INR 1,500,000 - 2,400,000