AI-Driven Observability & SRE Engineer

World-Wide-Technology

United States

Remote

USD 90,000 - 112,000

Full time

13 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health & Wellness
Profit Sharing
401k Matching
PTO & Holidays
Parental Leave
Tuition Reimbursement
Employee Discounts

Job summary

World Wide Technology (WWT) is seeking an Engineer to join the Observability & Site Reliability Engineering (SRE) team. The role focuses on designing, building, and operating systems that provide health visibility into IT services, applying SRE practices to boost reliability, performance, and resilience.

The ideal candidate will instrument, measure, and improve systems with an AI-first mindset, using automation and data-driven practices to accelerate response and reduce manual effort.

Qualifications

  • 5+ years of professional experience in IT Operations.
  • Experience in Metrics, Monitoring and Alerting, including tools like Prometheus and Grafana, Big Panda, Splunk, etc.
  • Knowledge of programming languages Python, Go, or equivalent.
  • Knowledge of system frameworks including Git and GitHub.
  • Critical thinker with excellent written and verbal communication skills.
  • Team‑oriented individual with very strong work ethic.
  • Familiarity with Linux, preferably administrative knowledge.
  • Understanding of container technologies (Docker, Podman, etc.).
  • Understanding of SRE concepts such as SLIs, SLOs, incident response, capacity planning, and reliability automation.
  • Practical understanding of AI‑enabled tools, automation patterns, and responsible AI use to improve operational efficiency.
  • Willingness to participate in an on‑call rotation and drive incidents through triage, mitigation, and resolution.
  • Experience with Infrastructure as Code (Terraform, Ansible, or similar) and CI/CD pipelines.
  • Understanding of distributed systems fundamentals: fault tolerance, redundancy, load balancing, and caching.

Responsibilities

  • Collection and strategic application of metrics to drive organizational decisions.
  • Providing a holistic view of system health using observability practices.
  • APM, RUM, and Synthetic Transaction monitoring.
  • Driving reliability through monitoring, alerting, observability, and SRE practices.
  • Applying AI‑first thinking to automate workflows, correlate alerts, surface insights, and speed incident response.
  • Improving reliability through SLOs, SLIs, error budgets, incident reviews, and continuous improvement.
  • Defining and maintaining on‑call rotations, escalation paths, and incident response processes, including participating in on‑call coverage.
  • Reducing operational toil through automation, self‑healing systems, and infrastructure as code.
  • Capacity planning and performance engineering to ensure systems scale reliably under load.
  • Partnering with development teams on production readiness reviews, architecture reviews, and resilience/chaos testing to prevent incidents before they happen.

Skills

IT Operations
Metrics & Monitoring
Programming: Python/Go
Linux administration
On-call & incident response
CI/CD pipelines
Terraform/Ansible (IaC)

Tools

Prometheus
Grafana
BigPanda
Splunk
CI/CD tooling
Terraform
Ansible
Docker

Job description

World Wide Technology (WWT) is seeking an Engineer to join the Observability & Site Reliability Engineering (SRE) team. The role focuses on designing, building, and operating systems that provide health visibility into IT services, applying SRE practices to boost reliability, performance, and resilience.

The ideal candidate will instrument, measure, and improve systems with an AI-first mindset, using automation and data-driven practices to accelerate response and reduce manual effort.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI-First Observability and SRE Engineer
AI-First Observability and SRE Engineer

World Wide Technology • New Home (MO)

On-site
USD 90,000 - 112,000
Health benefits
401k matching
Paid time off
+3
Observability & SRE Engineer
Observability & SRE Engineer

World Wide Technology • New Home (MO)

On-site
USD 90,000 - 112,000
Health benefits
401k matching
Paid time off
+3
Staff SRE: AI‑Driven Reliability Leader
Staff SRE: AI‑Driven Reliability Leader

WEX Inc. • United States

On-site
USD 121,000 - 151,000
Health insurance
Retirement savings plan
Paid time off
+1
SRE Architect: AI-Powered Reliability Strategy
SRE Architect: AI-Powered Reliability Strategy

WEX, Inc. • Chicago (IL)

On-site
USD 200,000 - 251,000
Health insurance
Retirement savings plan
Paid time off
+1
Senior SRE Lead - AI-Powered Reliability & Scale
Senior SRE Lead - AI-Powered Reliability & Scale

WEX • San Francisco (CA)

On-site
USD 121,000 - 151,000
Health insurance
Dental insurance
Vision insurance
+7
Senior Agentic Ops & Observability Architect (Remote)
Senior Agentic Ops & Observability Architect (Remote)

World Wide Technology, Inc. • Northern (KY)

Hybrid
USD 140,000 - 170,000
Health and Wellbeing benefits
401k with company matching
Paid time off and parental leave
Staff SRE: AI-Driven Reliability & Platform Leader
Staff SRE: AI-Driven Reliability & Platform Leader

WEX, Inc. • San Francisco (CA)

On-site
USD 121,000 - 151,000
AI-Driven SRE Architect for Enterprise Reliability
AI-Driven SRE Architect for Enterprise Reliability

WEX, Inc. • Dallas (TX)

On-site
USD 200,000 - 251,000
Staff SRE & AI-Driven Reliability Leader
Staff SRE & AI-Driven Reliability Leader

WEX • Chicago (IL)

On-site
USD 121,000 - 151,000
Health insurance
Dental insurance
Vision insurance
+7
AI‑Powered SRE Architect—Global Reliability Strategy
AI‑Powered SRE Architect—Global Reliability Strategy

WEX, Inc. • Seattle (WA)

On-site
USD 200,000 - 251,000
Health, dental, vision insurance
Retirement plan
Paid time off
+1