Senior Site Reliability Engineer

Precisely International Jobs

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Precisely is seeking a Senior Site Reliability Engineer to enhance reliability, performance, and scalability across CCX, RCX, and HMS. You will bridge software engineering and operations, building automation, observability tooling, and reliability standards to ensure platform availability.

The role involves mentoring engineers, leading complex reliability work, and contributing to release triage and incident response across AWS and on‑prem environments.

Qualifications

  • Bachelor’s degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
  • 5+ years of systems or infrastructure engineering in an enterprise production environment.
  • High-level competency in one or more infrastructure domains (cloud, IaC, observability).
  • Proficient with Linux (RHEL/Oracle) across on‑prem and cloud environments.
  • Terraform and Ansible proficiency; experience with AWS workloads.
  • Scripting in Python or Bash for automation.
  • Strong monitoring and alerting design experience (Datadog preferred).
  • Understanding of TCP/IP, DNS, load balancing, and distributed systems.
  • CI/CD design and deployment automation experience.
  • Ability to define SLOs and lead ORRs; cross‑functional collaboration.
  • Willingness to participate in on-call rotations.

Responsibilities

  • Define and maintain reliability standards across CCX, RCX, and HMS (SLOs, SLIs, error budgets).
  • Set standards for alerting, logging, and tracing; partner with engineering to meet them.
  • Act as platform expert across on‑prem, private cloud, and AWS environments.
  • Build and maintain IaC, deployment automation, monitoring, alerting, and observability tooling.
  • Embed reliability and operability considerations into service design and delivery.
  • Lead ORRs and validate production readiness and disaster recovery.
  • Contribute to release triage, deployment coordination, and change reviews.
  • Lead incident response, incident commander when needed, and communicate status clearly.
  • Write root cause analyses and preventive automation to reduce toil and MTTR.
  • Maintain runbooks, reliability backlogs, and knowledge resources.

Skills

5+ years experience
SRE fundamentals
Linux proficiency
Terraform
Ansible
Datadog
Python/Bash scripting
CI/CD
Networking basics
Incident management

Education

Bachelor's degree in CS/IS/Engineering

Tools

Terraform
Ansible
Datadog
AWS (EC2, ECS, S3, VPC, IAM, CloudWatch)

Job description

At Precisely, we’re not just building software — we’re shaping the future of data integrity.

As a global leader in data quality, data enrichment, and location intelligence, Precisely helps thousands of the world’s most trusted brands make confident decisions with data they can rely on. We’re an AI-first organization, which means artificial intelligence isn’t a buzzword here — it’s woven into how we build products, how we work, and how we think about solving complex problems for our customers. When you join Precisely, you join a team of curious, driven innovators who believe that better data makes the world run better. If you’re ready to do meaningful work at the intersection of AI and data — and help define what’s possible — we’d love for you to apply!


This position is 100% remote within the U.S., with a preference for candidates located in the Mountain or Pacific Time zones.

Overview:

The Site Reliability Engineer (SRE) is responsible for the reliability, performance, and scalability of Precisely’s infrastructure platforms across CEDAR (CCX) — an on-premises, private cloud managed services environment; RapidCX (RCX) — an AWS cloud environment for SaaS-delivered customer communications management; and Hosted Managed Services (HMS) — an AWS cloud environment supporting managed client deployments.


This role bridges software engineering and systems operations, building automation, observability tooling, and reliability standards to ensure platform availability and operational excellence. SREs are enabling partners: they set reliability standards, define what ‘reliable’ looks like for each service, validate production readiness, and coach engineering teams on operational best practices. Engineering teams own the reliability outcomes of the services they build; the SRE ensures they have the standards, tooling, and guidance to meet them.


As a Senior SRE, this role is regarded as a platform expert across CCX, RCX, and HMS, taking on the majority of complex reliability engineering work, mentoring less experienced engineers, and contributing to release triage and coordination efforts.


What you will do:

  • Define and maintain reliability standards across CCX, RCX, and HMS, including SLOs, SLIs, and error budgets.
  • Define standards for actionable alerting, meaningful logging, and diagnostic tracing; partner with engineering teams to ensure those standards are met.
  • Serve as a platform expert and lead complex reliability engineering work across on-premises, private cloud, and AWS environments.
  • Build and maintain infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling using Terraform, Ansible, Datadog, and scripting languages.
  • Partner with engineering teams to embed reliability, operability, scalability, backup, recovery, and failure‑mode considerations into service design and delivery.
  • Lead Operational Readiness Reviews and validate production and disaster recovery readiness for qualifying changes.
  • Contribute to release triage, deployment coordination, maintenance windows, and change and security reviews.
  • Lead response to complex P1/P2 incidents, serve as incident commander when required, and communicate status clearly to technical and non‑technical stakeholders.
  • Author root cause analyses and post‑incident reports, identify recurring patterns, and implement preventive automation to reduce manual toil and improve MTTR.
  • Maintain operational runbooks, reliability backlogs, and team knowledge resources covering monitoring gaps, recurring incidents, and operational risks.
  • Use Precisely‑provided AI tools for automation, incident analysis, troubleshooting, runbook creation, solution testing, and architecture documentation.
  • Ensure infrastructure configurations meet security, compliance, vulnerability‑remediation, and data‑protection requirements.
  • Mentor Associate SREs and SREs, coach engineering teams on production diagnostics and operational practices, and contribute to cross‑team design reviews.
  • Lead project‑based work, coordinate on‑call coverage, and model strong operational discipline and follow‑through.
  • Participate in a rotating on‑call schedule covering after‑hours, weekends, and holidays, including critical production changes and escalations within defined SLA windows.

What we are looking for:

Required:

  • Educational requirements (equivalent work experience will be accepted in place of the education requirement): Bachelor’s degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
  • 5+ years of systems or infrastructure engineering experience in an enterprise production environment.
  • High‑level, developing subject‑matter‑expert competency in one or more complex infrastructure domains (e.g., cloud platform engineering, IaC automation, observability, or on‑premises virtualization).
  • Advanced proficiency with Linux (RHEL/Oracle Linux) across multiple environments (on‑premises and cloud, not just multi‑site).
  • Proficient with Terraform for infrastructure‑as‑code; experienced with Ansible role development beyond basic hands‑on use.
  • Experience deploying and managing workloads in AWS at an intermediate level: EC2, ECS, S3, VPC, IAM, CloudWatch, Auto Scaling.
  • Proficiency with at least one scripting language (Python, Bash) for automation development.
  • Demonstrated experience designing monitoring system architecture and alerting strategy — not just operating existing dashboards (Datadog preferred).
  • Solid understanding of TCP/IP networking, DNS, load balancing, and distributed systems.
  • Experience designing and managing CI/CD pipelines and deployment automation standards.
  • Strong analytical skills; demonstrated experience authoring root cause analyses for complex incidents and identifying systemic, preventive fixes from recurring incident patterns.
  • Ability to define SLOs and lead Operational Readiness Reviews (ORRs); comfortable partnering with engineering teams on production readiness.
  • Demonstrated ability to work cross‑functionally with engineering teams on reliability standards and observability requirements.
  • Experience participating in change advisory processes and contributing to capacity and reliability planning.
  • Travel is required: No — approximately 0%.

AI Skills/Knowledge:

Active, proficient use of Precisely‑provided AI tools (GitHub Copilot, Claude, or equivalent) for complex automation, solution testing, and architecture documentation is a required baseline for this role — not a differentiator — and includes mentoring others on AI‑assisted engineering practices.


  • Apply AI tools for complex automation, incident analysis, and runbook authoring.
  • Mentor other engineers on AI‑assisted engineering practices.
  • Maintain fluency with Precisely‑approved AI coding assistants as a baseline expectation.

Preferred Skills (a plus but not required):

  • Experience with containerization and orchestration (Docker, ECS, Kubernetes).
  • Familiarity with GitOps workflows and source control best practices (Git, GitLab).
  • Knowledge of enterprise virtualization platforms in a hybrid cloud context.
  • Understanding of change management and ITIL operational practices.
  • Experience with enterprise security tooling (Qualys, CrowdStrike, Rapid7).
  • AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification.
  • Wireshark and protocol analysis experience.
  • Prior experience mentoring engineers or leading on‑call rotations for a team.

#LI-KM1 #LI-Remote


Application and Interview Impersonation Notice:

Impersonating another individual when applying for employment, and/or participating in an interview process to assist another individual in obtaining employment, with Precisely Software Incorporated (“Precisely”) is unlawful. If Precisely identifies such fraudulent conduct, then as applicable and to the extent permitted by law, the application will be rejected, an offer (if made) will be rescinded, or the employment will be terminated, and legal action may be taken against the impersonators.


The personal data that you provide as a part of this job application will be handled in accordance with relevant laws. For more information about how Precisely handles the personal data of job applicants, please see the Precisely Candidate Privacy Notice

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Engineer
Senior Platform Engineer

Precisely • United States

On-site
USD 140,000 - 190,000
Senior Software Engineer (Full Stack - Node.js, React, MongoDB)
Senior Software Engineer (Full Stack - Node.js, React, MongoDB)

Precisely International Jobs • United States

Remote
USD 140,000 - 180,000
Support Engineer I
Support Engineer I

Precisely International Jobs • United States

Remote
USD 65,000 - 90,000
Senior Staff Software Engineer – SRE & AIOps
Senior Staff Software Engineer – SRE & AIOps

ServiceNow • California (MO)

Hybrid
USD 191,000 - 334,000
Senior Software Engineer (Full Stack - Node.js, React, MongoDB)
Senior Software Engineer (Full Stack - Node.js, React, MongoDB)

Precisely • Virginia (MN)

On-site
USD 120,000 - 160,000
Data Analyst
Data Analyst

Precisely International Jobs • United States

Remote
USD 90,000 - 130,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Platform Engineer
Platform Engineer

Precisely International Jobs • United States

Remote
USD 120,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • United States

On-site
USD 145,000 - 200,000
Remote-first