Hybrid Work Model: Two days in‑office per week (Tuesday, Wednesday, or Thursday), with remaining days remote or in office.
Job Overview
Our Technology team is at the core of Priceline, creating, testing, learning and iterating to meet customer needs. As Priceline scales globally, reliable production visibility is critical. This role supports and evolves end‑to‑end observability solutions for OpenTelemetry signals across infrastructure, containers and Kubernetes environments.
Responsibilities
- Support and evolve end‑to‑end observability solutions for collecting, shipping, storing, and querying OpenTelemetry signals (metrics, logs, and traces) across infrastructure, containers, and Kubernetes environments, while influencing architectural decisions for scalability and long‑term sustainability.
- Administer and operate core observability platforms (Splunk, New Relic, ClickHouse, Grafana, Lightrun), including onboarding, access management, configuration, upgrades, and ensuring platform reliability, performance, and SLAs.
- Drive the adoption and standardization of instrumentation practices across services, establishing consistent logging, metrics, and distributed tracing standards, schemas, and conventions.
- Partner with product, platform, and engineering teams to enhance production visibility, support SLO‑driven reliability practices, and act as a subject‑matter expert for observability.
- Optimize telemetry pipelines for performance, data quality, scalability and cost efficiency, including implementing strategies such as sampling, filtering, and data lifecycle management.
- Define and support observability governance standards, driving consistency and adoption through documentation, tooling, and enablement.
- Lead complex incident investigations and post‑mortems, identifying observability gaps and driving improvements to reduce MTTR and MTTD while improving alert quality and signal‑to‑noise ratio.
- Contribute to advancing the observability platform toward intelligent and AI‑enabled capabilities, exploring MCP‑based and other solutions to improve signal quality, incident triage and operational efficiency.
Qualifications
- Bachelor’s degree in Computer Science or equivalent practical experience.
- 4+ years of experience in Observability, SRE, DevOps, or platform engineering roles supporting production systems.
- Strong understanding of APM and SRE fundamentals, including MELT (Metrics, Events, Logs, Traces), latency analysis, error‑rate monitoring, service dependency mapping, SLIs/SLOs, alert tuning and root‑cause analysis, demonstrated in large‑scale distributed systems.
- Hands‑on experience administering at least one modern observability/APM platform (e.g., Splunk, New Relic, Grafana), with practical exposure to metrics, logs, distributed tracing and platform configuration.
- Experience supporting full‑stack observability coverage across infrastructure, application and browser layers, including operating platforms at scale.
- Experience building dashboards and actionable alerts, configuring alert workflows and integrations with incident management tools such as PagerDuty.
- Experience implementing or supporting OpenTelemetry‑based instrumentation and improving telemetry quality across services, focusing on reducing alert fatigue and improving signal‑to‑noise ratio.
- Familiarity with Kubernetes and cloud‑native environments – understanding of application deployment, monitoring, scaling and troubleshooting complex production issues.
- Experience managing telemetry pipelines and agents (e.g., collectors, forwarders, sidecars), onboarding services, troubleshooting ingestion issues, and optimizing pipelines for scale and efficiency.
- Working knowledge of scripting or automation (Shell, Python) and CI/CD concepts; familiarity with infrastructure‑as‑code tools such as Terraform is a plus.
- Comfortable collaborating with engineering teams to improve monitoring standards, instrumentation quality, and overall production visibility, with proven ability to influence.
- Ability to analyze trade‑offs between observability depth, performance, and cost, and make recommendations aligned with business and engineering priorities.
- Experience leading or contributing to incident investigations and post‑mortems, identifying observability gaps and driving continuous improvement.
- Relevant certifications such as New Relic APM Professional, Reliability Engineer – Professional, Splunk Admin or GCP Associate Cloud Engineer are a plus.
- Demonstrated history of living Priceline values (Customer, Innovation, Team, Accountability and Trust).
Benefits
- Health and wellness coverage including medical, dental, vision and mental health resources.
- Generous time off: PTO, holidays, a company‑wide Priceline Pause reset week and paid volunteer days.
- Work/life support: work up to 4 weeks per year from anywhere, parental leave, dependent care, family support, Summer Fridays and office perks like stocked kitchens and catered meals.
- Financial security programs: retirement plans with company contributions, life and disability coverage, tax‑advantaged accounts.
- Signature travel perks: employee‑only discounts on hotels and flights, VIP deals and Big Deal Bucks credits.
- Additional perks and discounts: travel and partner discounts, tuition support, legal support and pet benefits.
- People‑first culture with Employee Resource Groups, social events, recognition programs and service awards.
Priceline is a proud equal opportunity employer. We embrace and celebrate the unique lenses through which our employees see the world.