AIOps / Observability Engineer

Cloud Security Web

Phoenix (AZ)

On-site

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cloud Security Web in Phoenix, AZ seeks an experienced AIOps/Observability Engineer to design, implement, and optimize enterprise observability and monitoring across cloud environments. The role emphasizes automation, API integrations, and proactive reliability improvements.

Ideal candidates will have 8+ years in AIOps or SRE, hands-on experience with Dynatrace, Splunk, Datadog, Prometheus, Grafana, and Elastic Stack, and strong scripting abilities in Python or Terraform.

Qualifications

  • 8+ years in AIOps, observability, SRE, or production support.
  • Hands-on with Dynatrace, Splunk, Datadog, Prometheus, Grafana, New Relic, or Elastic Stack.
  • REST APIs and API integrations experience.
  • Cloud platforms including AWS.
  • Scripting/automation using Python, Shell, PowerShell, Ansible, Terraform, or similar.

Responsibilities

  • Design, implement, and maintain enterprise observability and monitoring solutions.
  • Develop and optimize AIOps strategies to improve efficiency and incident resolution time.
  • Build and integrate APIs to connect monitoring, alerting, and automation platforms.
  • Configure and manage cloud-native monitoring across AWS, Azure, or GCP.
  • Automate operational tasks using scripting and IaC tools.
  • Monitor production environments for high availability and performance.
  • Lead incident management, RCA, and problem resolution for critical issues.
  • Collaborate with teams to improve application performance and resiliency.

Skills

AIOps
Observability
SRE
Cloud monitoring

Tools

Dynatrace
Splunk
Datadog
Prometheus
Grafana
New Relic
Elastic Stack

Job description

We are seeking an experienced AIOps / Observability Engineer to join our team in Phoenix, AZ. The ideal candidate will have strong expertise in observability platforms, cloud monitoring, automation, API integrations, and Site Reliability Engineering (SRE). This role requires hands-on experience supporting production environments, improving system reliability, and implementing proactive monitoring solutions. Banking or financial services experience is highly preferred.

Key Responsibilities
  • Design, implement, and maintain enterprise observability and monitoring solutions.
  • Develop and optimize AIOps strategies to improve operational efficiency and reduce incident resolution time.
  • Build and integrate APIs to connect monitoring, alerting, and automation platforms.
  • Configure and manage cloud-native monitoring across AWS, Azure, or Google Cloud Platform.
  • Automate operational tasks using scripting languages and infrastructure-as-code tools.
  • Monitor production environments to ensure high availability, performance, and reliability.
  • Perform incident management, root cause analysis (RCA), and problem resolution for critical production issues.
  • Collaborate with development, infrastructure, and operations teams to improve application performance and resiliency.
  • Create dashboards, alerts, and reports to provide real-time visibility into application and infrastructure health.
  • Implement best practices for observability, logging, tracing, and performance monitoring.
  • Participate in on-call production support and continuous service improvement initiatives.
Required Skills
  • 8+ years of experience in AIOps, Observability, SRE, or Production Support.
  • Strong experience with observability and monitoring platforms such as Dynatrace, Splunk, AppDynamics, Datadog, Prometheus, Grafana, New Relic, or Elastic Stack.
  • Hands-on experience with REST APIs and API integrations.
  • Experience with cloud platforms including AWS
  • Strong scripting and automation experience using Python, Shell, PowerShell, Ansible, Terraform, or similar technologies.
  • Experience with incident management, troubleshooting, and root cause analysis.
  • Strong understanding of application performance monitoring (APM), distributed tracing, logging, and metrics.
  • Experience supporting mission-critical production environments.
  • Excellent analytical, communication, and problem-solving skills.
Preferred Qualifications
  • Experience with CI/CD pipelines and DevOps practices.
  • Knowledge of Kubernetes, Docker, and container observability.
  • Familiarity with ITSM tools such as ServiceNow.
  • Experience implementing AI/ML-driven monitoring and predictive analytics.
Nice to Have
  • Kafka or messaging platform monitoring.
  • Infrastructure as Code (Terraform/CloudFormation).
  • Certification in AWS, Azure, GCP, Dynatrace, Splunk, or SRE.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AIOps & Observability Engineer
Senior AIOps & Observability Engineer

Cloud Security Web • Phoenix (AZ)

On-site
USD 120,000 - 160,000
Observability Engineer
Observability Engineer

BCforward • Phoenix (AZ)

Hybrid
USD 120,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Shya Workforce Solutions • Town of Florida (NY)

On-site
USD 100,000 - 140,000
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Ontrac Solutions • Chicago (IL)

On-site
USD 120,000 - 180,000
Local_Observability Operations Engineer
Local_Observability Operations Engineer

Themesoft Inc • Phoenix (AZ)

On-site
USD 110,000 - 140,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

TechDigital Group • Tyson (AZ)

On-site
USD 120,000 - 160,000
Observability Operations Engineer
Observability Operations Engineer

IntraEdge • Phoenix (AZ)

On-site
USD 140,000 - 190,000
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs

Tech Mirrors • Fort Mill (SC)

Hybrid
USD 140,000 - 190,000