Site Reliability Engineer (SRE) I

Socket.dev

Minnesota

Hybrid

USD 71,000 - 131,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid Work Model
Career Development
Mental Health Days
401k Plan

Job summary

Thomson Reuters is fortifying its Site Reliability Engineering capability to build, operate, and improve reliable production services. You will work with observability platforms, automation, and AI-enabled tools to improve the quality and availability of context used during incidents.

This hands-on role emphasizes learning across platforms, improving runbooks and telemetry, and collaborating with Product Engineering and platform teams to reduce toil and enhance reliability.

Qualifications

  • 3+ years in Site Reliability Engineering, DevOps, cloud infrastructure, or related field.
  • Experience with production telemetry (logs, metrics, traces, dashboards).
  • Experience troubleshooting production issues and incident response.
  • Experience with scripting or programming (Python, Bash, JavaScript, Go, etc.).
  • Ability to document operational processes and runbooks clearly.
  • Familiarity with AI-enabled coding or incident-management tools.

Responsibilities

  • Maintain SRE tooling: dashboards, alerts, runbooks, and telemetry baselines.
  • Investigate service health using logs, metrics, traces, and alerts.
  • Participate in incident response and post-incident reviews.
  • Execute runbooks within change-management processes.
  • Document facts, observations, and actions during incidents and handoffs.
  • Contribute to automation and deployment telemetry enhancements.
  • Review and validate AI-generated operational artifacts.

Skills

Site Reliability Eng
Cloud infrastructure
Observability & monitoring
Incident response
Scripting (Python, Bash)
Communication

Tools

Datadog
Dynatrace
New Relic
Grafana
Prometheus

Job description

Job Description

Thomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate, and improve reliable production services.

TheSite Reliability Engineerwill support the tools, processes, and operational practices that help teams detect, investigate, respond to, and prevent production reliability issues. You will work with observability platforms, operational documentation, automation, deployment information, and AI-enabled tools to improve the quality and availability of context used during day-to-day operations and incidents.

This is a hands‑on engineering role for someone who enjoys learning how complex systems work, improving operational readiness, and contributing practical solutions to production challenges. You will work closely with experienced SREs, Product Engineering teams, platform teams, and operations partners to maintain reliable services and reduce operational toil.

Rather than expecting you to know every architecture on day one, this role will help you build familiarity across products and platforms through maintained documentation, dashboards, runbooks, telemetry, deployment data, and operational tooling. When you identify gaps in that context, you will help improve the systems and processes that keep it current.

You will also use AI‑enabled engineering and investigation tools responsibly to accelerate analysis, documentation, and operational workflows. You will apply technical judgment, validate outputs, and accelerate when additional expertise or review is needed.

Key Responsibilities
  • Support and maintain SRE operational tooling, including dashboards, alerts, runbooks, service documentation, telemetry baselines, deployment visibility, and dependency information.

  • Use observability tools—including logs, metrics, traces, dashboards, and alerts—to investigate service‑health issues, identify trends, and support incident response.

  • Participate in incident response by gathering relevant context, reviewing recent changes, following established runbooks, documenting findings, and helping coordinate technical follow‑up actions.

  • Execute approved runbooks and mitigation procedures within established escalation, change‑management, and decision‑making processes.

  • Clearly document facts, observations, hypotheses, actions, and open questions during incidents, handoffs, and operational reviews.

  • Help improve the accuracy, completeness, and freshness of operational context used by engineering and operations teams during incidents and routine production support.

  • Contribute to automation and integration work that keeps operational information current, such as CI/CD notifications, deployment telemetry, change‑correlation data, service ownership records, and monitoring configuration.

  • Review and validate operational artifacts, including runbooks, diagrams, dashboards, alerts, and AI‑generated documentation, with guidance from senior engineers and service owners.

  • Assist with root‑cause analysis, post‑incident reviews, and follow‑up work by identifying gaps in monitoring, documentation, automation, instrumentation, or operational processes.

  • Treat missing runbooks, outdated documentation, incomplete telemetry, and unclear service ownership as improvement opportunities; partner with the appropriate teams to help resolve those gaps.

  • Contribute to service‑health and error‑reduction initiatives using available SLO, error budget, incident, alerting, and operational data.

  • Partner with Product Engineering and platform teams to identify reliability and observability improvements, including monitoring gaps, alert quality, deployment visibility, capacity concerns, and failure‑mode coverage.

  • Contribute directly to code, scripts, infrastructure configuration, dashboards, alerts, automation, and documentation that improve service reliability and reduce manual operational work.

  • Use AI‑enabled coding and investigation tools to accelerate log review, documentation updates, runbook drafting, incident summarization, and hypothesis generation, while validating results before relying on them.

  • Provide actionable feedback when AI‑enabled operational tools produce incomplete, inaccurate, or insufficiently supported outputs.

  • Participate in design reviews, sprint planning, and operational‑readiness discussions, helping ensure reliability and observability considerations are addressed before production deployment.

  • Support blameless post‑incident reviews focused on learning, systemic improvement, and preventing recurring issues.

Required Qualifications
  • 3+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, platform engineering, systems engineering, production operations, software engineering, or a related technical field.

  • Working knowledge of at least two of the following areas: cloud infrastructure, distributed systems, observability and monitoring, networking, databases, CI/CD, containers, or infrastructure automation.

  • Experience working with production telemetry, including logs, metrics, traces, dashboards, monitoring platforms, or alerting systems.

  • Experience troubleshooting production issues, participating in incident response, or supporting business‑critical applications and services.

  • Experience writing, maintaining, or improving operational documentation, runbooks, knowledge articles, or support procedures.

  • Experience with scripting or programming in one or more languages, such as Python, Bash, PowerShell, JavaScript, Java, Go, or a comparable language.

  • Willingness and ability to contribute to code, infrastructure configuration, dashboards, monitoring rules, alerts, automation, or documentation.

  • Ability to communicate clearly about technical findings, operational risks, and next steps with engineers and operational stakeholders.

  • Ability to work effectively in a collaborative environment, ask for help when needed, and learn from more experienced engineers.

  • Familiarity with AI‑enabled coding, documentation, investigation, or operational‑analysis tools, along with an understanding that outputs must be reviewed and validated.

Preferred Qualifications
  • Experience with cloud platforms, Kubernetes, containers, CI/CD tooling, infrastructure‑as‑code, or configuration‑management tools.

  • Experience with observability platforms such as Datadog, Dynatrace, New Relic, Splunk, Grafana, Prometheus, Elastic, or similar technologies.

  • Experience defining or working with Service Level Objectives, Service Level Indicators, error budgets, service‑health metrics, or incident‑management processes.

  • Experience supporting 24/7 production environments or participating in an on‑call rotation.

  • Experience improving dashboards, alerts, runbooks, deployment visibility, change correlation, service documentation, or operational workflows.

  • Familiarity with AI agent workflows, AI‑assisted root‑cause analysis, or AI‑enabled incident‑management tools.

  • Experience contributing to automation that reduces repetitive operational work and improves response consistency.

  • Experience participating in blameless post‑incident reviews and helping drive corrective actions to completion.

  • Relevant certifications in cloud infrastructure, Kubernetes, DevOps, SRE, observability, or incident management are beneficial but not required.

#LI-LP2

What's in it For You?

  • Hybrid Work Model: We've adopted a flexible hybrid working environment for our office‑based roles while delivering a seamless experience that is digitally and physically connected.
  • Flexibility & Work-Life Balance: Flex My Way is a set of supportive workplace policies designed to help manage personal and professional responsibilities, whether caring for family, giving back to the community, or finding time to refresh and reset. This builds upon our flexible work arrangements, including work from anywhere for up to 8 weeks per year, empowering employees to achieve a better work‑life balance.
  • Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrow’s challenges and deliver real‑world solutions. Our Grow My Way programming and skills‑first approach ensures you have the tools and knowledge to grow, lead, and thrive in an AI‑enabled future.
  • Industry Competitive Benefits: We offer comprehensive benefit plans to include flexible vacation, two company‑wide Mental Health Days off, access to the Headspace app, retirement savings, tuition reimbursement, employee incentive programs, and resources for mental, physical, and financial wellbeing.
  • Culture: Globally recognized, award‑winning reputation for inclusion and belonging, flexibility, work‑life balance, and more. We live by our values: Obsess over our Customers, Compete to Win, Challenge (Y)our Thinking, Act Fast / Learn Fast, and Stronger Together.
  • Social Impact: Make an impact in your community with our Social Impact Institute. We offer employees two paid volunteer days off annually and opportunities to get involved with pro‑bono consulting projects and Environmental, Social, and Governance (ESG) initiatives.
  • Making a Real‑World Impact:We are one of the few companies globally that helps its customers pursue justice, truth, and transparency. Together, with the professionals and institutions we serve, we help uphold the rule of law, turn the wheels of commerce, catch bad actors, report the facts, and provide trusted, unbiased information to people all over the world.

In the United States, Thomson Reuters offers a comprehensive benefits package to our employees. Our benefit package includes market competitive health, dental, vision, disability, and life insurance programs, as well as a competitive 401k plan with company match. In addition, Thomson Reuters offers market leading work life benefits with competitive vacation, sick and safe paid time off, paid holidays (including two company mental health days off), parental leave, sabbatical leave. These benefits meet or exceeds the requirements of paid time off in accordance with any applicable state or municipal laws. Finally, Thomson Reuters offers the following additional benefits: optional hospital, accident and sickness insurance paid 100% by the employee; optional life and AD&D insurance paid 100% by the employee; Flexible Spending and Health Savings Accounts; fitness reimbursement; access to Employee Assistance Program; Group Legal Identity Theft Protection benefit paid 100% by employee; access to 529 Plan; commuter benefits; Adoption & Surrogacy Assistance; Tuition Reimbursement; and access to Employee Stock Purchase Plan. Thomson Reuters complies with local laws that require upfront disclosure of the expected pay range for a position. The base compensation range varies across locations. For any eligible US locations, unless otherwise noted, the base compensation range for this role is $70,800 USD - $131,400 USD. Base pay is positioned within the range based on several factors including an individual’s knowledge, skills and experience with consideration given to internal equity. Base pay is one part of a comprehensive Total Reward program which also includes flexible and supportive benefits and other wellbeing programs. This role may also be eligible for an Annual Bonus based on a combination of enterprise and individual performance.

About Us

Thomson Reuters informs the way forward by bringing together the trusted content and technology that people and organizations need to make the right decisions. We serve professionals across legal, tax, accounting, compliance, government, and media. Our products combine highly specialized software and insights to empower professionals with the data, intelligence and solutions needed to make …

We are powered by the talents of 26,000 employees across more than 70 countries, where everyone has a chance to contribute and grow professionally in flexible work environments. At a time when objectivity, accuracy, fairness, and transparency are under attack, we consider it our duty to pursue them. Sound exciting? Join us and help shape the industries that move society forward.

As a global business, we rely on the unique backgrounds, perspectives, and experiences of all employees to deliver on our business goals. To ensure we can do that, we seek talented, qualified employees in all our operations around the world regardless of race, color, sex/gender, including pregnancy, gender identity and expression, national origin, religion, sexual orientation, disability, age, marital status, citizen status, veteran status, or any other protected classification under applicable law. Thomson Reuters is proud to be an Equal Employment Opportunity Employer providing a drug‑free workplace.

Thomson Reuters makes reasonable accommodations for applicants with disabilities, including veterans with disabilities, and for sincerely held religious beliefs in accordance with applicable law. If you reside in the United States and require an accommodation in the recruiting process, you may contact our Human Resources Department at HR.Leave-Expert@thomsonreuters.com. Disability accommodations in the recruiting process may include things like a sign language interpreter, making interview rooms accessible, providing assistive technology, or other relevant accommodations. Please note this email is not intended for general recruitment questions and we will promptly respond to inquiries regarding accommodations. More information on requesting an accommodation here.

Learn more on how to protect yourself from fraudulent job postings here.

More information about Thomson Reuters can be found on thomsonreuters.com

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) I
Site Reliability Engineer (SRE) I

Refinitiv • Eagan (MN)

Hybrid
USD 71,000 - 131,000
Director, Forward Deployed Solutions (AI)
Director, Forward Deployed Solutions (AI)

Socket.dev • Town of Texas (WI)

Hybrid
USD 137,000 - 255,000
Hybrid work model
Competitive compensation
Mental health days
Account Specialist
Account Specialist

PowerToFly • Ann Arbor (MI)

Hybrid
USD 106,000 - 198,000
Data & Analytics Business Partner, Reuters
Data & Analytics Business Partner, Reuters

Refinitiv • New York (NY)

Hybrid
USD 131,000 - 219,000
Hybrid work model
Mental health days
Headspace access
+4
Senior Account Specialist
Senior Account Specialist

Socket.dev • Town of Texas (WI)

Hybrid
USD 106,000 - 198,000
Hybrid Work Model
Flex My Way
Grow My Way
+2
Program Manager, Social Impact, Government Affairs & ESG
Program Manager, Social Impact, Government Affairs & ESG

Refinitiv • Eagan (MN)

Hybrid
USD 74,000 - 138,000
Hybrid work model
Flexible work policies
Career development
Lead Platform Engineer
Lead Platform Engineer

Thomson Reuters • Frisco (TX)

Hybrid
USD 118,000 - 220,000
Hybrid Work Model
Work-Life Balance
Career Growth
+4
Cloud Engineer - Professional Services
Cloud Engineer - Professional Services

Socket.dev • Town of Texas (WI)

On-site
USD 82,000 - 152,000
Flexible vacation
Mental health days
Tuition reimbursement
+1
Director, Revenue Operations
Director, Revenue Operations

Thomson Reuters • Eagan (MN)

Hybrid
USD 137,000 - 255,000
Hybrid work model
Flexible vacation
Tuition reimbursement
Director, Forward Deployed Solutions (AI)
Director, Forward Deployed Solutions (AI)

Thomson Reuters • New York (NY)

On-site
USD 158,000 - 294,000
Hybrid Work Model
Tuition Reimbursement
Mental Health & Wellness benefits
+1