Site Reliability Engineer (SRE) I

Thomson Reuters

Eagan, Northern (MN, KY)

Hybrid

USD 110,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Flexible work policy

Job summary

Thomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate, and improve reliable production services. You will focus on observability, automation, and AI-enabled tooling to enhance incident response and reduce toil.

Join a hands-on, collaborative team that spans product engineering, platform teams, and operations. You will maintain runbooks, dashboards, and documentation while contributing to automation and reliability

Qualifications

  • 3+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, platform engineering, systems engineering, production operations, software engineering, or a related technical field.
  • Experience working with production telemetry, including logs, metrics, traces, dashboards, monitoring platforms, or alerting systems.
  • Experience writing, maintaining, or improving operational documentation, runbooks, knowledge articles, or support procedures.

Responsibilities

  • Support and maintain SRE operational tooling, including dashboards, alerts, runbooks, service documentation, telemetry baselines, deployment visibility, and dependency information.
  • Use observability tools to investigate service-health issues, identify trends, and support incident response.
  • Participate in incident response by gathering relevant context, reviewing changes, following runbooks, documenting findings, and coordinating follow-up actions.

Skills

SRE experience
Cloud infrastructure
Observability/Monitoring
Incident response
Scripting

Tools

Datadog
Dynatrace
Prometheus

Job description

## Job DescriptionThomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate, and improve reliable production services.The **Site Reliability Engineer** will support the tools, processes, and operational practices that help teams detect, investigate, respond to, and prevent production reliability issues. You will work with observability platforms, operational documentation, automation, deployment information, and AI-enabled tools to improve the quality and availability of context used during day-to-day operations and incidents.This is a hands-on engineering role for someone who enjoys learning how complex systems work, improving operational readiness, and contributing practical solutions to production challenges. You will work closely with experienced SREs, Product Engineering teams, platform teams, and operations partners to maintain reliable services and reduce operational toil.Rather than expecting you to know every architecture on day one, this role will help you build familiarity across products and platforms through maintained documentation, dashboards, runbooks, telemetry, deployment data, and operational tooling. When you identify gaps in that context, you will help improve the systems and processes that keep it current.You will also use **AI-enabled engineering and investigation tools** responsibly to accelerate analysis, documentation, and operational workflows. You will apply technical judgment, validate outputs, and escalate when additional expertise or review is needed.## Key Responsibilities* Support and maintain SRE operational tooling, including dashboards, alerts, runbooks, service documentation, telemetry baselines, deployment visibility, and dependency information.* Use observability tools - including logs, metrics, traces, dashboards, and alerts - to investigate service-health issues, identify trends, and support incident response.* Participate in incident response by gathering relevant context, reviewing recent changes, following established runbooks, documenting findings, and helping coordinate technical follow-up actions.* Execute approved runbooks and mitigation procedures within established escalation, change-management, and decision-making processes.* Clearly document facts, observations, hypotheses, actions, and open questions during incidents, handoffs, and operational reviews.* Help improve the accuracy, completeness, and freshness of operational context used by engineering and operations teams during incidents and routine production support.* Contribute to automation and integration work that keeps operational information current, such as CI/CD notifications, deployment telemetry, change-correlation data, service ownership records, and monitoring configuration.* Review and validate operational artifacts, including runbooks, diagrams, dashboards, alerts, and AI-generated documentation, with guidance from senior engineers and service owners.* Assist with root-cause analysis, post-incident reviews, and follow-up work by identifying gaps in monitoring, documentation, automation, instrumentation, or operational processes.* Treat missing runbooks, outdated documentation, incomplete telemetry, and unclear service ownership as improvement opportunities; partner with the appropriate teams to help resolve those gaps.* Contribute to service-health and error-reduction initiatives using available SLO, error budget, incident, alerting, and operational data.* Partner with Product Engineering and platform teams to identify reliability and observability improvements, including monitoring gaps, alert quality, deployment visibility, capacity concerns, and failure-mode coverage.* Contribute directly to code, scripts, infrastructure configuration, dashboards, alerts, automation, and documentation that improve service reliability and reduce manual operational work.* Use AI-enabled coding and investigation tools to accelerate log review, documentation updates, runbook drafting, incident summarization, and hypothesis generation, while validating results before relying on them.* Provide actionable feedback when AI-enabled operational tools produce incomplete, inaccurate, or insufficiently supported outputs.* Participate in design reviews, sprint planning, and operational-readiness discussions, helping ensure reliability and observability considerations are addressed before production deployment.* Support blameless post-incident reviews focused on learning, systemic improvement, and preventing recurring issues.## Required Qualifications* 3+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, platform engineering, systems engineering, production operations, software engineering, or a related technical field.* Working knowledge of at least two of the following areas: cloud infrastructure, distributed systems, observability and monitoring, networking, databases, CI/CD, containers, or infrastructure automation.* Experience working with production telemetry, including logs, metrics, traces, dashboards, monitoring platforms, or alerting systems.* Experience troubleshooting production issues, participating in incident response, or supporting business-critical applications and services.* Experience writing, maintaining, or improving operational documentation, runbooks, knowledge articles, or support procedures.* Experience with scripting or programming in one or more languages, such as Python, Bash, PowerShell, JavaScript, Java, Go, or a comparable language.* Willingness and ability to contribute to code, infrastructure configuration, dashboards, monitoring rules, alerts, automation, or documentation.* Ability to communicate clearly about technical findings, operational risks, and next steps with engineers and operational stakeholders.* Ability to work effectively in a collaborative environment, ask for help when needed, and learn from more experienced engineers.* Familiarity with AI-enabled coding, documentation, investigation, or operational-analysis tools, along with an understanding that outputs must be reviewed and validated.## Preferred Qualifications* Experience with cloud platforms, Kubernetes, containers, CI/CD tooling, infrastructure-as-code, or configuration-management tools.* Experience with observability platforms such as Datadog, Dynatrace, New Relic, Splunk, Grafana, Prometheus, Elastic, or similar technologies.* Experience defining or working with Service Level Objectives, Service Level Indicators, error budgets, service-health metrics, or incident-management processes.* Experience supporting 24/7 production environments or participating in an on-call rotation.* Experience improving dashboards, alerts, runbooks, deployment visibility, change correlation, service documentation, or operational workflows.* Familiarity with AI agent workflows, AI-assisted root-cause analysis, or AI-enabled incident-management tools.* Experience contributing to automation that reduces repetitive operational work and improves response consistency.* Experience participating in blameless post-incident reviews and helping drive corrective actions to completion.* Relevant certifications in cloud infrastructure, Kubernetes, DevOps, SRE, observability, or incident management are beneficial but not required.#LI-LP2* **Hybrid Work Model:** We've adopted a flexible hybrid working environment for our office-based roles while delivering a seamless experience that is digitally and physically connected.* **Flexibility & Work-Life Balance:** Flex My Way is a set of supportive workplace policies designed to help manage personal and professional responsibilities, whether caring for family, giving back to the community, or finding time to refresh and reset. This builds upon our flexible work arrangements, including work from anywhere for up to 8 weeks per year, empowering employees to achieve a better work-life balance.*
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) I
Site Reliability Engineer (SRE) I

Reuters • Eagan (MN), Northern (KY)

Hybrid
USD 120,000 - 150,000
Hybrid work model
Professional development opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

London Stock Exchange Group • Raleigh (NC)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer I
Senior Site Reliability Engineer I

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks, Inc. • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay including Bonus
401(k) Plan with Company Match
Paid Vacation and Sick Days
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1