Observability Analyst

State of Oklahoma

Oklahoma

On-site

USD 65,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The State of Oklahoma, Office of Management and Enterprise Services, is seeking an Observability Analyst to monitor, triage, and improve the monitoring and observability capabilities across server, cloud, and network environments. You will build dashboards, alerts, and turn telemetry into early warning signals to help teams find and fix problems before they impact the business.

This hands-on, on-site role in Oklahoma City requires Datadog experience and the ability to translate complex telemetry

Qualifications

  • Bachelor's or higher in IT/CS or related field, or equivalent hands-on experience.
  • 2+ years in infrastructure monitoring, observability, systems administration, or IT operations.
  • Hands-on experience with Datadog or similar observability platforms.
  • Working knowledge of server, cloud (AWS/Azure/GCP), and network fundamentals for instrumentation and troubleshooting.
  • Experience configuring dashboards, alerts, and notification workflows for incident response.

Responsibilities

  • Design, implement, and maintain monitoring and observability tooling across server, cloud, and network environments.
  • Instrument infrastructure and services with metrics, logs, traces, and synthetic checks for full-stack visibility.
  • Build dashboards to provide role-appropriate visibility into system health for engineers, managers, and leadership.
  • Configure alert thresholds and escalation policies to detect issues early while minimizing noise.
  • Integrate monitoring tools with incident management, ticketing, and on-call systems.

Skills

Datadog
Monitoring
Observability
Cloud
IT Operations

Education

Associate's or Bachelor's degree in IT/CS

Tools

Dynatrace
New Relic
Splunk
Prometheus/Grafana

Job description

Agency

090 OFFICE OF MANAGEMENT AND ENTERPRISE SERV

Job Posting Title

Observability Analyst

Supervisory Organization

IS-CS

Job Posting End Date

Refer to the date listed at the top of this posting, if available. Continuous if date is blank.

Job Posting End Date Note

Note: Applications will be accepted until 11:59 PM on the day prior to the posting end date above.

Full/Part-Time

Full time

Job Type

Regular

Compensation

Job Details

  • This is a full-time, 40-hour per week position
  • Support the Information Services Division
  • Position is on-site in Oklahoma City, OK
  • The salary for this position is based on education and experience
Job Description

Position Summary

The Observability Analyst is responsible for day-to-day monitoring, triage, correlation, and first-line coordination with domain teams and the IT Operations Command Center (ITOCC). You will maintain and continuously improve the monitoring and observability capabilities that keep the organization's server, cloud, and network environments healthy. Using Datadog or similar platforms, this role builds meaningful dashboards and alerts and turns raw telemetry into early warning signals that let teams find and fix problems before they impact the business. You will focus on pattern analysis and hunting silent anomalies, feeding findings into the continuous-improvement loop. This is a hands-on technical role for someone who enjoys making complex environments visible, understandable, and measurably more reliable.

Job Details

  • This is a full-time, 40-hour per week position
  • Support the Information Services Division
  • Position is on-site in Oklahoma City, OK
  • The salary for this position is based on education and experience
Position Responsibilities
Monitoring & Observability Engineering
  • Design, deploy, and maintain monitoring and observability tooling (Datadog or similar) across server, cloud, and network environments.
  • Instrument infrastructure, applications, and services with metrics, logs, traces, and synthetic checks to provide full-stack
  • Maintain monitoring and observability tooling (Datadog or similar) across server, cloud, and network environments.
  • Instrument infrastructure, applications, and services with metrics, logs, traces, and synthetic checks to provide full-stack visibility.
  • Build and maintain dashboards that give clear, role-appropriate visibility into system health for engineers, managers, and leadership.
  • Configure and tune alerting thresholds and escalation policies to catch real issues early while minimizing noise and alert fatigue.
  • Integrate monitoring tools with incident management, ticketing, and on-call notification systems (e.g., PagerDuty, ServiceNow, Slack).
IT Health & Continuous Improvement
  • Track and report on key IT health indicators such as uptime, latency, error rates, capacity headroom, and patch/compliance status across multiple environments.
  • Partner with infrastructure, cloud, and network teams to identify recurring issues and drive root-cause fixes rather than repeated firefighting.
  • Support capacity planning by analyzing utilization trends and flagging environments approaching risk thresholds.
  • Contribute to post-incident reviews by providing telemetry, timelines, and health data that clarify what happened and why.
  • Continuously refine monitoring coverage as new systems, services, and cloud resources are added, retiring stale checks and dashboards.
Collaboration & Documentation
  • Work closely with server, cloud, network, and application teams to understand what 'healthy' looks like for each environment and translate that into monitoring coverage.
  • Document monitoring standards, runbooks, and dashboard conventions so coverage stays consistent as the environment grows.
  • Train and support other engineers in interpreting dashboards, alerts, and observability data.
  • Evaluate and recommend improvements or additions to the observability toolset as monitoring needs evolve.
Physical Demands and Work Environment
  • Office-based work involving extensive computer and phone use.
  • Requires long periods of sitting, up to eight hours a day.
  • Possible on-call rotation
  • Work environment is generally quiet, occasional travel may be required.
Education And Experience
  • Associate's or Bachelor's degree in Information Technology, Computer Science, or related field, or equivalent hands-on experience.
  • 2+ years of experience in infrastructure monitoring, observability, systems administration, or IT operations.
  • Hands-on experience with Datadog or a comparable observability platform (e.g., Dynatrace, New Relic, Splunk, Prometheus/Grafana).
  • Working knowledge of server, cloud (AWS, Azure, or GCP), and network fundamentals sufficient to instrument and troubleshoot across environments.
  • Experience configuring dashboards, alerts, and notification workflows that support fast, accurate incident response.
Preferred Qualifications
  • Familiarity with ITSM practices (incident, problem, change management) and on-call/escalation processes.
  • Exposure to AIOps or automated remediation approaches that reduce manual intervention.
  • Experience monitoring environments in regulated industries with attention to security and compliance requirements.
About OMES

The Office of Management and Enterprise Services provides excellent service, expert guidance and continuous improvement in support of our partners’ goals. We are a highly qualified workforce committed to serve those who serve Oklahomans and make government run in the most efficient, innovative manner possible.

Equal Opportunity Employment

OMES is an Equal Opportunity Employer. Reasonable accommodation to individuals with disabilities may be provided upon request.

The State of Oklahoma is an equal opportunity employer and does not discriminate on the basis of genetic information, race, religion, color, sex, age, national origin, or disability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer
Observability Engineer

State of Oklahoma • Oklahoma

On-site
USD 90,000 - 120,000
Director of Network Operations
Director of Network Operations

State of Oklahoma • Oklahoma City (OK)

On-site
USD 126,000 - 154,000
Vacation and sick leave
Holidays
Comprehensive benefits
Observability Engineer - Build Dashboards & Detect Anomalies
Observability Engineer - Build Dashboards & Detect Anomalies

State of Oklahoma • Oklahoma

On-site
USD 65,000 - 90,000
Server Operations Support
Server Operations Support

Oklahoma AG • Oklahoma City (OK)

On-site
USD 54,000 - 66,000
Generous leave including 15 days ofvac
Comprehensive benefits package with an
insurance premiums for employees and+
Director of Network Operations
Director of Network Operations

090 OFFICE OF MANAGEMENT AND ENTERPRISE SERV • Oklahoma City (OK)

On-site
USD 110,000 - 140,000
Generous leave and holidays
Comprehensive benefits package
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)

Data Analysis Incorporated • Los Angeles (CA)

On-site
USD 115,000 - 125,000
10% yearly bonus
Dynamic workplace environment
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)

O'Neil Digital Solutions, LLC • Los Angeles (CA)

On-site
USD 115,000 - 125,000
10% annual bonus target
Computer Support Technician I
Computer Support Technician I

090 OFFICE OF MANAGEMENT AND ENTERPRISE SERV • Mangum (OK)

On-site
USD 51,000 - 63,000
Computer Support Technician
Computer Support Technician

State of Oklahoma • Oklahoma

On-site
USD 48,000 - 57,000
15 days of vacation
15 days of sick leave
11 paid holidays annually
Computer Support Technician I
Computer Support Technician I

State of Oklahoma • Oklahoma

On-site
USD 34,000 - 57,000
Generous leave
Comprehensive benefits