Senior Platform Engineer

Ww

United States

On-site

USD 200,000 - 215,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ww is expanding its Platform Engineering team to design and implement a modern observability stack across hundreds of services. You will architect logging, metrics, and tracing patterns, reduce toil, surface actionable signals, and help turn chaos into clarity.

Reporting to Head of Platform Engineering as an individual contributor, you will set standards, guide incident management, and drive automation. Strong preference for hands-on leadership and bias for action.

Qualifications

  • 5+ years in SRE or platform engineering
  • Deep expertise in at least one observability domain
  • Experience with agent configuration and instrumentation
  • Track record of building monitoring/alerting systems
  • Comfortable making opinionated architectural decisions with incomplete information
  • Strong systems thinking and tradeoff awareness

Responsibilities

  • Design and deploy a logging, metrics, and APM stack that scales with the organization
  • Establish agent configuration best practices (New Relic, OpenTelemetry, or alternatives)
  • Create structured logging patterns and cardinality management strategy
  • Implement distributed tracing to understand latency and failure paths across services
  • Evaluate and select an incident management platform
  • Build alerting policies that surface real signals without noise
  • Develop incident triage, escalation, and communication workflows
  • Create runbooks and post-incident analysis templates
  • Identify manual work in incident response and automate it away
  • Build dashboards, alerts, and self-service tools for troubleshooting
  • Codify monitoring and alerting as IaC
  • Set SRE practices and standards for the organization
  • Mentor Platform Engineering team on observability concepts and tooling
  • Drive adoption: make it easy for service owners to instrument their code

Skills

SRE experience
Observability domain
Agent configuration
Systems thinking
Architectural decisions

Tools

New Relic
Datadog
OpenTelemetry
Prometheus
ELK
Grafana
Jaeger

Job description

Opportunity

We are expanding our Platform Engineering team to build out observability and incident response from the ground up. Our cloud architecture supports hundreds of services, but visibility into system health, performance, and failures is largely manual today. We need someone to architect and implement a modern observability stack—not just monitor it.

This is a chance to make an immediate impact: you will design logging, metrics, and tracing patterns that the entire engineering organization depends on, reduce toil, surface actionable signals, and turn chaos into clarity.

You’ll report to our Head of Platform Engineering as an individual contributor and will have the autonomy to move quickly, make opinionated decisions, and set standards for how we operate.

What Success Looks Like
  • Phase 1: Audit & Plan – Map current gaps, propose tooling & rollout strategy, get buy‑in.
  • Phase 2: Foundation – Select tools, document decisions, pilot agents on 2‑3 services.
  • Phase 3: Standardize – Structured logging patterns, baseline dashboards, incident routing, alerting.
  • Phase 4: Prove It – Demonstrate reduced MTTD and MTTR, fewer manual escalations, documented runbooks, team trained on tooling.
Key Responsibilities

Observability Architecture & Implementation

  • Design and deploy a logging, metrics, and APM stack that scales with the organization
  • Establish agent configuration best practices (New Relic, OpenTelemetry, or alternatives)
  • Create structured logging patterns and cardinality management strategy
  • Implement distributed tracing to understand latency and failure paths across services

Incident Management & Response

  • Evaluate and select an incident management platform (we’re not locked into any vendor)
  • Build alerting policies that surface real signals without drowning teams in noise
  • Develop incident triage, escalation, and communication workflows
  • Create runbooks and post‑incident analysis templates

Automation & Toil Reduction

  • Identify manual work in today's incident response and automate it away
  • Build dashboards, alerts, and self‑service tools so teams don’t need platform team involvement for basic troubleshooting
  • Codify monitoring and alerting as infrastructure (IaC)

Technical Leadership

  • Set SRE practices and standards for the organization
  • Mentor Platform Engineering team on observability concepts and tooling
  • Drive adoption: make it easy for service owners to instrument their code correctly
About You

Required

  • 5+ years in SRE, platform engineering, or similar infrastructure role
  • Deep expertise in at least one observability domain (metrics, logging, tracing, or APM)
  • Experience with agent configuration and instrumentation (New Relic, Datadog, OpenTelemetry, or equivalent)
  • Track record of building (not just maintaining) monitoring/alerting systems
  • Comfortable making opinionated architectural decisions with incomplete information
  • Strong systems thinking—you understand tradeoffs between managed vs. open‑source, complexity vs. coverage

Strongly Preferred

  • AWS and containerized environments (EKS, Docker)
  • Infrastructure‑as‑Code (Terraform, CloudFormation)
  • Kafka or async message queues (understanding log/metric pipeline architecture)
  • RDS experience (database monitoring, slow query logs)
  • Open‑source observability stack (Prometheus, ELK, Grafana, Jaeger)

Cultural Fit

  • You move fast and are comfortable with ambiguity—greenfield projects are exciting, not daunting
  • You’re opinionated but pragmatic—you’ll push back on bad ideas and compromise when needed
  • You think like a builder, not a ticket‑taker
  • You communicate clearly: you can explain technical decisions to non‑technical stakeholders
Salary and Benefits

Base salary may vary depending on skills, experience, and location. This role is also eligible for a comprehensive benefits package and annual bonus program.

$200,000 — $215,000 USD

Equal Opportunity Employment

We are proud to be an equal opportunity employer and we do not discriminate on the basis of sex, race, color, creed, national origin, marital status, age, religion, sexual orientation, gender identity, gender expression, veteran status, or disability. By agreeing to participate in our process, you agree that any information we collect is subject to our Privacy Policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director of Platform Engineering
Director of Platform Engineering

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 120,000 - 240,000
Competitive compensation
Equity
Comprehensive benefits
Operational Data & Observability Engineer
Operational Data & Observability Engineer

Nscale • United States

On-site
USD 145,000 - 180,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Senior Platform Engineer (Observability & Telemetry)
Senior Platform Engineer (Observability & Telemetry)

Ports North • Baltimore (MD)

On-site
Senior Platform Engineer
Senior Platform Engineer

District-Partners • Bethesda (MD)

Hybrid
USD 160,000 - 215,000
Annual performance bonus
Equity participation
Opportunity for career advancement
Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Senior Platform Observability Engineer
Senior Platform Observability Engineer

Ww • United States

On-site
USD 200,000 - 215,000
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer, Observability New York, NY, United States
Senior Site Reliability Engineer, Observability New York, NY, United States

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Senior Platform Engineer
Senior Platform Engineer

Gridware Technologies Inc. • San Francisco (CA)

On-site
USD 190,000 - 210,000
Health, Dental & Vision
Paid parental leave
Alternating day off
+3
Principal Platform Engineer
Principal Platform Engineer

Flexential • United States

On-site
USD 180,000 - 210,000
Medical, Telehealth, Dental and Vision
HSA and FSA
AD&D
+2