Observability Operations Engineer

IntraEdge

Phoenix (AZ)

On-site

USD 140,000 - 190,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IntraEdge seeks a Senior Observability Operations Engineer to run and upgrade enterprise observability platforms. You will optimize Dynatrace, Splunk, and OpenSearch/Elasticsearch while implementing scalable monitoring, logging and tracing across cloud-native environments.

The role emphasizes automation, AI-assisted operations, and close collaboration with Platform Engineering, SRE and DevOps teams to ensure high availability and rapid incident resolution.

Qualifications

  • Broad experience operating enterprise observability platforms.
  • Proven ability to design and optimize monitoring, logging, tracing and alerting.
  • Strong automation experience across cloud-native stacks and CI/CD.

Responsibilities

  • Administer and optimize Dynatrace, Splunk and OpenSearch/Elasticsearch for enterprise platforms.
  • Design, deploy, configure and maintain monitoring, logging, tracing and alerting solutions.
  • Manage large OpenSearch/Elasticsearch clusters including indexing, tuning, backups and capacity planning.
  • Configure Dynatrace components and AI features for DEM, RUM and APM.
  • Administer Splunk Enterprise components and ITSI; develop dashboards and executive metrics.
  • Support Linux-based infrastructure and Kubernetes environments (OpenShift/Rancher preferred).
  • Automate tasks using Python, Shell, REST APIs, Terraform or Ansible.
  • Participate in incident, problem, change and release management; drive platform upgrades and security governance.
  • Improve reliability through automation, self-healing and AI-assisted operations.

Skills

Dynatrace
Splunk
OpenSearch/Elasticsearch
Kubernetes
Linux
OpenTelemetry
AI/ML observability
Python
Shell scripting
REST APIs
Terraform
Ansible
Docker
OpenShift/Rancher
Incident management
SRE/DevOps collaboration
Automation

Tools

Dynatrace OneAgent
Splunk ITSI
OpenSearch
Kubernetes platforms
Docker
OpenShift
Terraform
Ansible

Job description

We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions. Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is highly desirable.

The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.

Key Responsibilities
  • Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
  • Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
  • Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
  • Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).
  • Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
  • Develop dashboards, alerts, reports, and executive operational metrics.
  • Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
  • Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
  • Perform root cause analysis for production incidents using observability platforms.
  • Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
  • Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible.
  • Participate in incident, problem, change, and release management processes.
  • Drive platform upgrades, patching, security compliance, and operational governance.
  • Improve platform reliability through automation, self-healing, and AI-assisted operations.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Back fill Engineer
Back fill Engineer

JPC TECHNO INC • Phoenix (AZ)

On-site
USD 120,000 - 180,000
Observability Operations Engineer
Observability Operations Engineer

Tata Consultancy Services • Phoenix (AZ)

On-site
USD 100,000 - 120,000
Observability Engineer
Observability Engineer

BCforward • Phoenix (AZ)

Hybrid
USD 120,000 - 140,000
Senior Observability Platform Engineer - AI-Enhanced Ops
Senior Observability Platform Engineer - AI-Enhanced Ops

IntraEdge • Phoenix (AZ)

On-site
USD 140,000 - 190,000
Observability Engineer (Splunk & Dynatrace)
Observability Engineer (Splunk & Dynatrace)

Veriipro • Plano (TX)

On-site
USD 110,000 - 160,000
Senior Observability Engineer
Senior Observability Engineer

Veriipro • Boston (MA)

On-site
USD 140,000 - 190,000
Local_Observability Operations Engineer
Local_Observability Operations Engineer

Themesoft Inc • Phoenix (AZ)

On-site
USD 110,000 - 140,000
Observability Architect
Observability Architect

TechDigital Group • Atlanta (GA)

On-site
USD 120,000 - 150,000
Solution Architect / Team Lead - Observability
Solution Architect / Team Lead - Observability

VOLTO Consulting • Irvine (CA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

TechDigital Group • Tyson (AZ)

On-site
USD 120,000 - 160,000