Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Workday seeks a Principal Distributed Systems Engineer to lead the technical vision and architecture of the Observability Platform, focusing on distributed tracing. You will design a multi-petabyte-scale platform using ClickHouse and Grafana Tempo, with a big-data pipeline (Kafka, Spark/Flink, Iceberg, S3) on AWS and own end-to-end ingestion, storage, and query.
You will collaborate with ML/AI teams to push Observability AI forward, drive performance optimization, and ensure high availability
We are seeking a Principal Distributed Systems Engineer to lead the technical vision and architecture of Workday's Observability Platform, with a focus on distributed tracing. This role involves designing and building a multi-petabyte-scale platform using ClickHouse and Grafana Tempo, supported by a big-data pipeline including Kafka, Spark/Flink, Iceberg, and S3, all running on AWS. You will own the end-to-end infrastructure for ingestion, storage, and query, ensuring low-latency performance and high availability.
The ideal candidate has 14+ years of software development experience, 6+ years in complex distributed systems, and expertise in Go, Python, or Java. A bachelor's degree in a relevant field is required, with a master's strongly preferred.
The role offers a competitive salary range of $222,900-$334,300 annually, with flexibility in work arrangements.