Observability Engineer

BCforward

Phoenix (AZ)

Hybrid

USD 120,000 - 140,000

Full time

8 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BCforward seeks a Senior Observability Operations Engineer for its Phoenix, AZ hybrid team. The role focuses on managing Dynatrace, Splunk, and OpenSearch/Elasticsearch platforms, with Linux and Kubernetes support and a drive toward AI-assisted operations.

The candidate will automate tasks, optimize monitoring, and design robust logging and alerting. 3 days on-site weekly with flexible shifts and a weekend rotation may apply.

Qualifications

  • Bachelor's degree or higher in CS/IT/Engineering or equivalent.
  • 6–10+ years IT infrastructure or observability operations experience.
  • 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch.
  • Strong troubleshooting and analytical skills.
  • Excellent communication and stakeholder management.

Responsibilities

  • Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
  • Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
  • Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
  • Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).
  • Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
  • Develop dashboards, alerts, reports, and executive operational metrics.
  • Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
  • Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
  • Perform root cause analysis for production incidents using observability platforms.
  • Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
  • Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible.
  • Participate in incident, problem, change, and release management processes.
  • Drive platform upgrades, patching, security compliance, and operational governance.
  • Improve platform reliability through automation, self-healing, and AI-assisted operations.

Skills

Observability Platforms
Dynatrace Administration
OpenSearch Administration
Elasticsearch Administration
Grafana
Kibana
OpenTelemetry
Kafka

Education

Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience

Tools

Docker
OpenShift
Rancher
Terraform
Ansible
Git
REST APIs

Job description

BCforward is seeking a highly motivated and experienced Observability Operations Engineer
Note: Candidate must be local to Phoenix, Arizona.
Job Title: Observability Operations Engineer
Job Location: Phoenix, AZ Hybrid
Duration: Long-term
Pay Rate: $60/hr W2 and $68 CTC
Must Haves:
  • Dynatrace and Splunk, as well as OpenSearch/elastisearch and OpenTelemetry.
  • Kubernetes AI/ ML, Grafana and Sahara.

Role is 20% automating at 80% operations.

3 days a week. Standard shifts are either 9:00 AM to 6:00 PM or 10:00 AM to 6:30 PM to provide coverage.

Candidates also must be willing to work 1 weekend day every 2-3 weeks.

Job Description:

We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions. Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is highly desirable.

The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.

Key Responsibilities
  • Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
  • Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
  • Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
  • Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).
  • Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
  • Develop dashboards, alerts, reports, and executive operational metrics.
  • Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
  • Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
  • Perform root cause analysis for production incidents using observability platforms.
  • Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
  • Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible.
  • Participate in incident, problem, change, and release management processes.
  • Drive platform upgrades, patching, security compliance, and operational governance.
  • Improve platform reliability through automation, self-healing, and AI-assisted operations.
Required Technical Skills
  • Observability Platforms
  • Dynatrace Administration
  • OpenSearch Administration
  • Elasticsearch Administration
  • Grafana
  • Kibana
  • OpenTelemetry
  • Kafka (preferred)
Infrastructure
  • Docker
  • OpenShift or Rancher
  • Networking (TCP/IP, DNS, Load Balancers, Firewalls)
  • System Administration
  • Git
  • Terraform
  • Ansible
  • REST APIs
Scripting
  • Python
  • Bash/Shell
  • PowerShell (preferred)
AI & Automation Skills (Preferred)
  • Experience using Generative AI (ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar) to improve operational efficiency.
  • Knowledge of AIOps platforms and AI-driven observability.
  • Experience with Dynatrace Davis AI for anomaly detection and root cause analysis.
  • Understanding of machine learning concepts for predictive monitoring and intelligent alerting.
  • Experience building AI-assisted operational runbooks and troubleshooting workflows.
  • Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and AI-powered knowledge search is a plus.
  • Experience integrating AI with observability platforms using APIs.
  • Familiarity with LLMs, prompt engineering, and AI-assisted automation.
  • Experience using Python with AI frameworks (LangChain, LangGraph, OpenAI APIs, or similar) is desirable.
  • Exposure to AI-driven incident summarization, log analysis, and automated ticket enrichment.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
    • 6–10+ years of IT infrastructure or observability operations experience.
    • 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch.
    • Strong troubleshooting and analytical skills.
    • Excellent communication and stakeholder management skills.
Preferred Certifications
  • Dynatrace Associate or Professional Certification
  • Elastic Certified Engineer
  • ITIL Foundation
  • AI/ML or Generative AI certification (preferred)
Soft Skills
  • Strong ownership and accountability
  • Excellent problem-solving and analytical thinking
  • Ability to work independently with minimal supervision
  • Strong collaboration across cross-functional teams
  • Continuous learning mindset
  • Ability to thrive in fast-paced production environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Back fill Engineer
Back fill Engineer

JPC TECHNO INC • Phoenix (AZ)

On-site
USD 120,000 - 180,000
Local_Observability Operations Engineer
Local_Observability Operations Engineer

Themesoft Inc • Phoenix (AZ)

On-site
USD 110,000 - 140,000
Observability Operations Engineer
Observability Operations Engineer

Tata Consultancy Services • Phoenix (AZ)

On-site
USD 100,000 - 120,000
Observability Engineer
Observability Engineer

Cognizant • Phoenix (AZ)

On-site
USD 49,000 - 106,000
Medical/Dental/Vision/Life Insurance
Paid holidays + PTO
401(k) plan
+3
Senior Observability & AI Ops Engineer
Senior Observability & AI Ops Engineer

BCforward • Phoenix (AZ)

Hybrid
USD 120,000 - 140,000
Observability Architect
Observability Architect

TechDigital Group • Atlanta (GA)

On-site
USD 120,000 - 150,000
Sr. Observability Engineer
Sr. Observability Engineer

FreedomPay • Select (KY)

On-site
USD 150,000 - 210,000
Observability Engineer (Splunk & Dynatrace)
Observability Engineer (Splunk & Dynatrace)

Veriipro • Plano (TX)

On-site
USD 110,000 - 160,000
Sr Observability Engineer
Sr Observability Engineer

IT Associates • Irvine (CA)

Hybrid
USD 150,000 - 210,000
Observability Engineer
Observability Engineer

VMC Soft Technologies, Inc • Allen (TX)

On-site
USD 90,000 - 120,000