Site Reliability Engineer - Observability

United States Digital Space LLC

United States

Hybrid

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model

Job summary

Priceline's Technology team seeks a seasoned Observability/SRE engineer to strengthen reliability across infrastructure and applications, including Kubernetes environments. You will optimize telemetry pipelines, instrument services with OpenTelemetry, and manage platforms such as Splunk, New Relic, and Grafana, ensuring logs, metrics, and traces meet reliability targets.

Collaborating with product and platform teams, you will drive SLO adoption, alerting improvements, and scalable monitoring

Qualifications

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • 4+ years in Observability, SRE, DevOps, or platform engineering for production systems.
  • Strong understanding of MELT, SLIs/SLOs, alert tuning and root cause analysis.
  • Hands-on with at least one modern observability/APM platform (Splunk, New Relic, Grafana).
  • Experience building dashboards/alerts and OpenTelemetry instrumentation.
  • Familiarity with Kubernetes and cloud-native environments; CI/CD concepts.

Responsibilities

  • Support and evolve end-to-end observability for OpenTelemetry signals across infrastructure and Kubernetes.
  • Administer core observability platforms (Splunk, New Relic, Grafana, ClickHouse, Lightrun) including upgrades and SLAs.
  • Drive standardized instrumentation practices and superior logging/metrics/tracing schemas.
  • Collaborate with product and engineering to improve production visibility and SLO-driven reliability.
  • Optimize telemetry pipelines for performance, quality, scalability and cost.
  • Contribute to AI-enabled observability capabilities and incident investigations with postmortems.

Skills

Observability
SRE fundamentals
Kubernetes
CI/CD
Automation scripts
Telemetry pipelines
Dashboarding & alerts
OpenTelemetry
Incident response
Certifications

Education

Bachelor’s degree in Computer Science

Tools

Splunk
New Relic
Grafana
PagerDuty
Terraform

Job description

This role is eligible for our hybrid work model: Two days in-office.

Our Technology team is the backbone of our company: constantly creating, testing, learning and iterating to better meet the needs of our customers. If you thrive in a fast-paced, ideas-led environment, you’re in the right place.

Responsibilities
  • Support and evolve end-to-end observability solutions for collecting, shipping, storing, and querying OpenTelemetry signals (metrics, logs, and traces) across infrastructure, containers, and Kubernetes environments, while influencing architectural decisions for scalability and long-term sustainability.
  • Administer and operate core observability platforms (Splunk, New Relic, ClickHouse, Grafana, Lightrun), including onboarding, access management, configuration, upgrades, and ensuring platform reliability, performance, and SLAs.
  • Drive the adoption and standardization of instrumentation practices across services, establishing consistent logging, metrics, and distributed tracing standards, schemas, and conventions.
  • Partner with product, platform, and engineering teams to enhance production visibility, support SLO-driven reliability practices, and act as a subject matter expert for observability.
  • Optimize telemetry pipelines for performance, data quality, scalability, and cost efficiency, including implementing strategies such as sampling, filtering, and data lifecycle management.
  • Contribute to advancing the observability platform toward intelligent and AI-enabled capabilities, exploring MCP-based and other solutions to improve signal quality, incident triage, and operational efficiency.

Define and support observability governance standards, driving consistency and adoption through documentation, tooling, and enablement. Lead complex incident investigations and postmortems, identifying observability gaps and driving improvements to reduce MTTR and MTTD while improving alert quality and signal-to-noise ratio.

Qualifications
  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • 4+ years of experience in Observability, SRE, DevOps, or platform engineering roles supporting production systems.
  • Strong understanding of APM and SRE fundamentals, including MELT (Metrics, Events, Logs, Traces), latency analysis, error rate monitoring, service dependency mapping, SLIs/SLOs, alert tuning, and root cause analysis, with demonstrated application in large-scale distributed systems.
  • Hands‑on experience administering at least one modern observability/APM platform (e.g., Splunk, New Relic, Grafana), with practical exposure to metrics, logs, distributed tracing, and platform configuration. Experience supporting full‑stack observability coverage across infrastructure, application, and browser monitoring layers, including operating platforms at scale.
  • Experience building dashboards and actionable alerts, including configuring alert workflows and integrations with incident management tools such as PagerDuty. Experience implementing or supporting OpenTelemetry-based instrumentation and improving telemetry quality across services, with a focus on reducing alert fatigue and improving signal-to-noise ratio.
  • Familiarity with Kubernetes and cloud‑native environments – an understanding of how applications are deployed, monitored, and scaled, including troubleshooting complex production issues in distributed environments.
  • Experience managing telemetry pipelines and agents (e.g., collectors, forwarders, sidecars), including onboarding services and troubleshooting ingestion issues, and optimizing pipelines for scale and efficiency.
  • Working knowledge of scripting or automation (e.g., Shell, Python) and CI/CD concepts. Experience or familiarity with infrastructure‑as‑code tools such as Terraform for managing platform configurations and integrations is a plus.
  • Comfortable collaborating with engineering teams to improve monitoring standards, instrumentation quality, and overall production visibility, with proven ability to influence.
  • Ability to analyze trade‑offs between observability depth, performance, and cost, and make recommendations aligned with business and engineering priorities.
  • Experience leading or contributing to incident investigations and postmortems, identifying observability gaps and driving continuous improvement.
  • Relevant certifications such as New Relic APM Professional, Reliability Engineer – Professional, Splunk Admin, or GCP Associate Cloud Engineer are a plus.
  • Demonstrated history of living the values important to Priceline: Customer, Innovation, Team, Accountability and Trust.
Equal Opportunity Employment

Priceline is a proud equal opportunity employer. We embrace and celebrate the unique lenses through which our employees see the world. We’d love for you to join us and help shape what makes our team extraordinary.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid SRE – Observability & AI-Driven Ops
Hybrid SRE – Observability & AI-Driven Ops

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 180,000
Hybrid work model
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
Observability Engineer Site Reliability Engineer
Observability Engineer Site Reliability Engineer

Ontrac Solutions • United States

On-site
USD 120,000 - 180,000
SRE/Observability Engineer
SRE/Observability Engineer

BlueSky Resource Solutions • United States

Remote
USD 100,000 - 130,000
Observability Engineer Site Reliability Engineer
Observability Engineer Site Reliability Engineer

Ontrac Solutions • Arizona

Hybrid
USD 140,000 - 190,000
Verification cost reimbursement
Senior Site Reliability Engineer, Observability New York, NY, United States
Senior Site Reliability Engineer, Observability New York, NY, United States

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Staff SRE - Observability
Staff SRE - Observability

Focused • Chicago (IL)

Hybrid
USD 160,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Infosys • Richardson (TX)

On-site
USD 80,000 - 120,000
Senior Vice President, Site Reliability Automation Engineer
Senior Vice President, Site Reliability Automation Engineer

BNY Mellon • Town of Florida (NY)

Hybrid
USD 100,000 - 130,000
Principal Enterprise Architect, Infrastructure & Operations
Principal Enterprise Architect, Infrastructure & Operations

Priceline.com LLC • New York (NY)

Hybrid
USD 175,000 - 220,000