SRE: AI‑Driven Reliability & Observability

Socket.dev

Minnesota

On-site

USD 71,000 - 131,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid Work Model
Career Development
Mental Health Days
401k Plan

Job summary

Thomson Reuters is fortifying its Site Reliability Engineering capability to build, operate, and improve reliable production services. You will work with observability platforms, automation, and AI-enabled tools to improve the quality and availability of context used during incidents.

This hands-on role emphasizes learning across platforms, improving runbooks and telemetry, and collaborating with Product Engineering and platform teams to reduce toil and enhance reliability.

Qualifications

  • 3+ years in Site Reliability Engineering, DevOps, cloud infrastructure, or related field.
  • Experience with production telemetry (logs, metrics, traces, dashboards).
  • Experience troubleshooting production issues and incident response.
  • Experience with scripting or programming (Python, Bash, JavaScript, Go, etc.).
  • Ability to document operational processes and runbooks clearly.
  • Familiarity with AI-enabled coding or incident-management tools.

Responsibilities

  • Maintain SRE tooling: dashboards, alerts, runbooks, and telemetry baselines.
  • Investigate service health using logs, metrics, traces, and alerts.
  • Participate in incident response and post-incident reviews.
  • Execute runbooks within change-management processes.
  • Document facts, observations, and actions during incidents and handoffs.
  • Contribute to automation and deployment telemetry enhancements.
  • Review and validate AI-generated operational artifacts.

Skills

Site Reliability Eng
Cloud infrastructure
Observability & monitoring
Incident response
Scripting (Python, Bash)
Communication

Tools

Datadog
Dynatrace
New Relic
Grafana
Prometheus

Job description

Thomson Reuters is fortifying its Site Reliability Engineering capability to build, operate, and improve reliable production services. You will work with observability platforms, automation, and AI-enabled tools to improve the quality and availability of context used during incidents.

This hands-on role emphasizes learning across platforms, improving runbooks and telemetry, and collaborating with Product Engineering and platform teams to reduce toil and enhance reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: AI-Enhanced Observability
Site Reliability Engineer: AI-Enhanced Observability

Refinitiv • Eagan (MN)

Hybrid
USD 71,000 - 131,000
Global SRE Director - Hybrid, AI & Observability
Global SRE Director - Hybrid, AI & Observability

Socket.dev • Town of Texas (WI)

Hybrid
USD 159,000 - 295,000
Hybrid Work Model
Grow My Way
Mental Health Days
+3
Director, Global SRE & Reliability Leadership
Director, Global SRE & Reliability Leadership

Thomson Reuters • Frisco (TX)

Hybrid
USD 159,000 - 295,000
Hybrid Work Model
Flex My Way policies
Career Development & Growth
+5
SRE Engineer: AI‑Driven Reliability & Automation
SRE Engineer: AI‑Driven Reliability & Automation

ICE • Atlanta (GA)

On-site
USD 140,000 - 210,000
SRE for AI Platform: Reliability at Scale
SRE for AI Platform: Reliability at Scale

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior Applied AI SRE: Reliability & Observability Lead
Senior Applied AI SRE: Reliability & Observability Lead

PowerToFly • Tennessee

On-site
USD 120,000 - 190,000
Staff SRE: Platform Reliability & AI Observability
Staff SRE: Platform Reliability & AI Observability

Anduril Industries, Inc. • Costa Mesa (CA), Northern (KY)

Hybrid
USD 191,000 - 253,000
Associate SRE: Automation & Observability
Associate SRE: Automation & Observability

Calabrio • United States

On-site
USD 90,000 - 130,000
Senior SRE: Reliability & Observability Lead
Senior SRE: Reliability & Observability Lead

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1