Observability & SRE Engineer

World Wide Technology

New Home (MO)

On-site

USD 90,000 - 112,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
401k matching
Paid time off
Parental leave
Wellness program
Employee discounts

Job summary

World Wide Technology (WWT) seeks an Engineer for the Observability & Site Reliability Engineering (SRE) team. The engineer will design, build, and operate systems that provide visibility into IT service health while applying core SRE practices to boost reliability and performance.

The ideal candidate embraces instrumenting, measuring, and improving systems with an AI-first mindset, using automation and data-driven practices to accelerate response and reduce manual effort.

Qualifications

  • 5+ years of professional experience in IT Operations.
  • Experience with Metrics, Monitoring and Alerting tools such as Prometheus and Grafana.
  • Knowledge of Python, Go, or equivalent.
  • Knowledge of Git and GitHub.
  • Strong written and verbal communication skills.
  • Familiarity with Linux and container tech (Docker, Podman).
  • Understanding of SRE concepts (SLIs, SLOs, incident response).
  • Practical understanding of AI-enabled tooling and automation for reliability.
  • Experience with Infrastructure as Code (Terraform, Ansible) and CI/CD pipelines.
  • Understanding of distributed systems fundamentals.

Responsibilities

  • Collect and apply metrics to drive decisions across the org.
  • Maintain a holistic view of system health with observability.
  • Monitor APM, RUM, and synthetic transactions.
  • Drive reliability via monitoring, alerting, and SRE practices.
  • Apply AI-first thinking to automate workflows and speed incident response.
  • Improve reliability with SLOs, SLIs, error budgets, and post-incident reviews.
  • Define on-call rotations, escalation paths, and incident processes.
  • Reduce toil with automation and infrastructure as code.
  • Plan capacity and performance for scalable systems.
  • Collaborate with development for production readiness and resilience tests.

Skills

IT Operations
Monitoring & Alerting
Python
Go
Linux
SRE concepts
communication skills

Tools

Prometheus
Grafana
BigPanda
Splunk
Git
GitHub
Terraform
Ansible

Job description

World Wide Technology (WWT) strives to make a new world happen. WWT's work benefits clients and partners as much as it does its people and community across the globe.

Founded in 1990, WWT brings together strategy, deep technical expertise and world-class partnerships to help public and private sector organizations design, build and scale intelligent AI, digital, cybersecurity, cloud and infrastructure solutions. Through its Advanced Technology Center (ATC)—a collaborative ecosystem featuring state-of-the‑art hardware and software—WWT enables clients and partners to conceptualize, test and validate innovative technology and then deploy solutions at scale using its global integration and distribution capabilities.

With more than 14,000 team members and over 60 locations globally, WWT's culture—grounded in core values and leadership philosophies—has been recognized by Fortune® and Great Place to Work for its commitment to innovation, trust and creating a great place to work for all. WWT provides products and services to large enterprise, global service provider and public sector clients in up to 130 countries across six continents. Softchoice, a World Wide Technology company, supports commercial and SMB markets in the U.S. and Canada.

What will you be doing?

World Wide Technology (WWT) is seeking an Engineer to join the Observability & Site Reliability Engineering (SRE) team. In this role, the Engineer will design, build, and operate the systems that give WWT visibility into the health of its IT services, while also applying core SRE practices to improve their reliability, performance, and resilience.

The ideal candidate is a passionate technologist who enjoys instrumenting, measuring, and improving systems to strengthen observability, reliability, and operational decision‑making. They bring an AI‑first mindset, using AI, automation, and data‑driven practices responsibly to improve reliability, accelerate response, and reduce manual effort.

The Observability & SRE team brings together infrastructure, operations, automation, and reliability engineering skills to build consumable services and platforms that improve visibility, confidence, and resilience across WWT’s IT systems. This is an opportunity for someone looking to grow technically while helping modernize IT through AI‑enabled operations and reliability‑focused engineering.

Responsibilities:
  • Collection and strategic application of metrics to drive organizational decisions
  • Providing a holistic view of system health using observability practices
  • APM, RUM, and Synthetic Transaction monitoring
  • Driving reliability through monitoring, alerting, observability, and SRE practices
  • Applying AI‑first thinking to automate workflows, correlate alerts, surface insights, and speed incident response
  • Improving reliability through SLOs, SLIs, error budgets, incident reviews, and continuous improvement
  • Defining and maintaining on‑call rotations, escalation paths, and incident response processes, including participating in on‑call coverage
  • Reducing operational toil through automation, self‑healing systems, and infrastructure as code
  • Capacity planning and performance engineering to ensure systems scale reliably under load
  • Partnering with development teams on production readiness reviews, architecture reviews, and resilience/chaos testing to prevent incidents before they happen

,

Qualifications:
Minimum Qualifications
  • 5+ years of professional experience in IT Operations
  • Experience in Metrics, Monitoring and Alerting, including tools like Prometheus and Grafana, Big Panda, Splunk, etc.
  • Knowledge of programming languages Python, Go, or equivalent
  • Knowledge of system frameworks including Git and GitHub
  • Critical thinker with excellent written and verbal communication skills
  • Team‑oriented individual with very strong work ethic
  • Familiarity with Linux, preferably administrative knowledge
  • Understanding of container technologies (Docker, Podman, etc.)
  • Understanding of SRE concepts such as SLIs, SLOs, incident response, root cause analysis, capacity planning, and reliability automation
  • Practical understanding of AI‑enabled tools, automation patterns, and responsible AI use to improve operational efficiency
  • Willingness to participate in an on‑call rotation and drive incidents through triage, mitigation, and resolution
  • Experience with Infrastructure as Code (Terraform, Ansible, or similar) and CI/CD pipelines
  • Understanding of distributed systems fundamentals: fault tolerance, redundancy, load balancing, and caching
Preferred Qualifications
  • Kubernetes experience
  • Experience working with Agile methodology
  • Test‑First development mindset with functional, end‑to‑end and regression testing experience
  • Experience with AI‑assisted engineering, AIOps, or machine learning techniques for monitoring, alerting, or incident response
  • Hands‑on SRE experience with SLOs, error budgets, blameless reviews, toil reduction, and production readiness
  • Experience with chaos engineering or resilience/failure‑injection testing
  • Experience authoring and maintaining runbooks, playbooks, and production readiness review checklists
  • Public cloud platform experience (AWS, Azure, or GCP)
  • Familiarity with distributed tracing and log aggregation tooling (e.g., OpenTelemetry, ELK/Splunk)

Certain states and localities require employers to post a reasonable estimate of the salary range. A reasonable estimate of the current base pay range for this position is $90,200 to $112,000 annually. Actual salary will be based on a variety of factors, including shift, location, experience, skill set, performance, licensure and certification, and business needs. The range for this position in other geographic locations may differ. Certain positions may also be eligible for variable incentive compensation, such as bonuses or commissions, that are not included in the base pay.

The well‑being of WWT employees is essential. When it comes to our benefits package, WWT has one of the best. We offer the following benefits to all full‑time employees:

  • Health and Wellbeing: Health (Medical & Prescription), Dental, and Vision Care, Onsite Health Centers (MO & IL), Employee Assistance Program, Wellness program
  • Financial Benefits: Competitive Pay, Profit Sharing, 401k Plan with Company Matching, Life and Disability Insurance, Flexible Spending Accounts, Tuition Reimbursement
  • Paid Time Off: PTO & Holidays, Parental Leave, Medical Leave, Military Leave, Bereavement, Day of Caring
  • Additional Perks: Family Planning Benefits, Nursing Mothers Benefits, Voluntary Legal, Voluntary Supplemental Accident/Illness/Hospital, Voluntary ID Theft, Pet Insurance, Employee Discount Program

Note: This is not an all‑encompassing list and should not be used as a complete description of the plan’s benefits. For more information, see our US Benefits Website

We strive to create an environment where all employees are empowered to succeed based on their skills, performance, and dedication. Our goal is to cultivate a culture of belonging that encourages innovation, collaboration, and respect for all team members, ensuring that WWT remains a great place to work for all!

If you require accessibility accommodation(s) or adjustment during any stage of the hiring process, please let your WWT Recruiter know. The recruiter will work with you to understand your needs and help ensure an accessible experience throughout the interview process. World Wide Technology is an Equal Opportunity Employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

World Wide Technology • United States

On-site
USD 108,000 - 136,000
Health and Wellbeing benefits
Competitive Pay + 401k with matching
Profit Sharing
+3
Analyst - Management Consulting
Analyst - Management Consulting

World Wide Technology • New York (NY)

On-site
USD 80,000 - 90,000
Health Insurance
Dental
Vision
+10
Team Lead - Engineering Support and Automation
Team Lead - Engineering Support and Automation

World Wide Technology • Maryland Heights (MO)

On-site
USD 114,000 - 143,000
Health and Wellness Benefits
401k with company matching
Paid time off
+1
Manager, Development — AI Execution Platforms
Manager, Development — AI Execution Platforms

World Wide Technology • Maryland Heights (MO)

On-site
USD 150,000 - 170,000
Health insurance
401k with company matching
Paid time off
+3
Senior Engineering Manager – AI Advanced Development
Senior Engineering Manager – AI Advanced Development

World Wide Technology • Maryland Heights (MO)

On-site
USD 140,000 - 210,000
Health benefits
Dental & Vision
Onsite health centers (MO & IL)
+10
Analyst - Management Consulting
Analyst - Management Consulting

World Wide Technology • Maryland Heights (MO)

On-site
USD 80,000 - 90,000
Health and Wellbeing
Profit Sharing
401k Plan with Company Matching
+5
Customer Experience Operations Analyst
Customer Experience Operations Analyst

World Wide Technology, Inc. • Northern (KY)

Hybrid
USD 65,000 - 80,000
Health and Wellbeing
Competitive Pay
Profit Sharing
+2
Consulting Systems Engineer (Southern California) - Network
Consulting Systems Engineer (Southern California) - Network

World Wide Technology • California (MO)

On-site
USD 150,000 - 200,000
Health and Wellbeing
401k Plan with Company Matching
Paid Time Off
Consulting Systems Engineer (San Francisco, CA) - Hyperscaler Focused
Consulting Systems Engineer (San Francisco, CA) - Hyperscaler Focused

World Wide Technology • California (MO)

On-site
USD 150,000 - 210,000
Health insurance
401k with company matching
Paid time off
Team Lead, Security Operations Center (SOC) - 3rd Shift
Team Lead, Security Operations Center (SOC) - 3rd Shift

World Wide Technology • Maryland Heights (MO)

On-site
USD 122,000 - 152,000