Senior Software Engineer, Observability Delivery
United States
Engineering
Individual Contributor
Yes
5816
Full Time
About GitHub
GitHub is the world’s leading platform for agentic software development — powered by Copilot to build, scale, and deliver secure software. Over 180 million developers, including more than 90% of the Fortune 100 companies, use GitHub to collaborate, and more than 77,000 organisations have adopted GitHub Copilot.
Locations
In this role you can work from Remote, United States
Overview
GitHub’s Observability Delivery team builds and operates the company-wide pipelines for logs, metrics, traces, and exceptions that engineers rely on to monitor and diagnose their services. These systems handle telemetry from across GitHub at high scale. Keeping them reliable and efficient as demand grows is a constant engineering challenge.
We’re looking for a Senior Software Engineer who enjoys solving challenges across software and infrastructure. You’ll write and operate services, deploy and configure observability tools, and make architectural decisions about how telemetry moves through our systems. You’ll lead work to improve pipeline capacity, reliability, and efficiency while partnering with service teams throughout GitHub. If you enjoy making critical, large-scale infrastructure dependable and easier to operate, we’d also love to hear from you.
Responsibilities
- Design, build, and operate high-scale pipelines for logs, metrics, traces, and exceptions, balancing capacity, reliability, performance, and cost.
- Lead cross-team work to identify and remove scaling bottlenecks, improve how telemetry is collected and processed, and safely evolve critical production infrastructure.
- Write and maintain production software, primarily in Go and Ruby, while configuring and integrating open-source and commercial observability tools.
- Work across cloud infrastructure, Kubernetes, virtual machines, networking, and service connectivity to make distributed systems dependable and operable.
- Partner with service teams to understand their observability needs and improve the tools they use to diagnose and maintain their services.
- Own system health through monitoring, incident response, on-call participation, and improvements informed by operational experience.
- Provide technical leadership through design proposals, reviews, mentoring, and collaboration across teams.
Qualifications
- 6+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
- OR Associate's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 5+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
- OR Bachelor's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 4+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
- OR Master's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 2+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
- OR Doctorate in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field
- OR equivalent experience.
- 3+ years building and operating production services or infrastructure in a cloud or large-scale distributed environment.
- 2+ years deploying, configuring, or troubleshooting production workloads with Kubernetes or comparable container orchestration.
- 2+ years developing production software in Go, Ruby, or a comparable general-purpose language.
Preferred Qualifications
- Experience building or operating shared platforms used by other engineering teams.
- Experience with logging, metrics, or distributed tracing systems, including OpenTelemetry concepts or tooling.
- Experience operating telemetry pipelines or platforms such as Datadog or comparable tools.
- Experience with Azure or another major cloud platform, alongside hybrid or self-managed infrastructure.
- Familiarity with networking, service connectivity, and distributed-systems operations.
- Experience leading ambiguous technical work across teams and mentoring engineers through design and code review.
Compensation Range
The base salary range for this job is USD $124,000.00 - USD $329,200.00 /Yr.