Tech Lead Manager, Site Reliability

Jobgether

Pittsburgh (Allegheny County)

Hybrid

USD 146,000 - 183,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Annual bonus
Stock options
Remote-friendly
Health insurance (employer-paid)
Paid time off
Home office stipend
Parental leave
401(k)

Job summary

Jobgether seeks a technical leader to combine hands-on software engineering, Site Reliability Engineering, and people leadership within a high-ownership team in the United States. You will lead 3–4 SREs while remaining deeply involved in coding, architecture, and incident response to improve reliability and delivery speed.

You will drive modernization, observability, and platform tooling, balancing technical debt with delivery needs.

Qualifications

  • 4+ years of professional software engineering with hands-on coding.
  • 3+ years of platform engineering with DevOps/SRE principles.
  • Experience owning technical direction and people management.
  • Ability to be the final technical decision-maker without a separate Tech Lead.
  • Experience with distributed systems and modern infra (AWS, Kubernetes, Vault, Grafana, PostgreSQL, Kafka/MSK, Redis/Valkey).
  • Experience with Azure or GCP environments is welcome.
  • Strong understanding of microservice architectures and Agile/ Scrum processes.
  • Experience with Golang, TypeScript, and React is a plus.
  • Comfort with AI-assisted development tools and workflows.

Responsibilities

  • Lead a team of 3-4 Site Reliability Engineers and contribute hands-on to coding, architecture, and incident response.
  • Own technical debt and modernization across the team's scope while balancing reliability and delivery.
  • Maintain deep technical understanding of observability, DR, infrastructure, and platform initiatives.
  • Drive reliability, platform adoption, and developer self-service improvements.
  • Guide system design, architectural decisions, and perform code reviews.
  • Manage career development, performance, hiring, and team health.
  • Evolve Agile/Scrum/Kanban processes with engineering leadership.
  • Balance delivery urgency with quality and long-term maintainability.
  • Model AI-assisted development practices and broaden adoption across the team.
  • Participate in onboarding, planning, reviews, on-call, and incident response.

Skills

Hands-on coding
Site Reliability Engineering
People leadership
Incident response
Observability
Disaster recovery
Infrastructure modernization
DevOps
AWS
Kubernetes/EKS
HashiCorp Vault
Grafana
GitHub Actions
PostgreSQL
Kafka/MSK
Redis/Valkey
Golang
TypeScript
React
AI-assisted development tools

Tools

AWS
Kubernetes/EKS
kops
HashiCorp Vault
Grafana
GitHub Actions
PostgreSQL
Kafka/MSK
Redis/Valkey
Azure
GCP

Job description

This role combines hands-on software engineering, Site Reliability Engineering, and people leadership within a high-ownership technical team. You will lead a small team of 3--4 Site Reliability Engineers while remaining deeply involved in coding, architecture, infrastructure, and incident response. The position is designed for a technical leader who can own both engineering direction and team development without relying on a separate Tech Lead. You will drive reliability, platform adoption, infrastructure modernization, and improvements to developer self-service and delivery speed. Success will be measured through stronger system performance, faster incident detection and recovery, reduced operational toil, and higher-quality engineering practices. This is an opportunity to shape both the technology and the way a growing engineering team operates.

Accountabilities:
  • Serve as a hands-on contributor to Site Reliability and platform engineering work, including incident response, infrastructure automation, tooling, and technical improvements.
  • Own technical debt and infrastructure modernization across the team's scope, prioritizing and resolving issues directly while balancing reliability with delivery needs.
  • Maintain a deep technical understanding of the team's projects, including observability, disaster recovery, infrastructure, and platform initiatives, and actively challenge and improve technical approaches.
  • Own team delivery outcomes and engineering quality, with accountability for reliability and performance metrics, incident response effectiveness, and adoption of platform tools.
  • Guide technical direction through system design, architectural consultation, technical decision-making, and hands-on code reviews.
  • Lead and develop a team of 3--4 Site Reliability Engineers, including career development, performance management, hiring, coaching, and overall team health.
  • Partner with engineering leadership to evolve Agile, Scrum, or Kanban processes based on what works effectively for the team.
  • Balance delivery urgency with appropriate engineering quality, reliability, and long-term maintainability.
  • Model effective AI-assisted software development through hands-on use of AI tools and establish strong practices for the wider team.
  • Participate in onboarding, codebase exploration, infrastructure work, planning, reviews, on-call activities, and incident response.
  • Contribute directly to technical work within the first months while developing strong relationships with direct reports and gaining a detailed understanding of the team's systems and workflows.
  • Join the on-call rotation and identify opportunities to improve response times, reduce repeated alerts, and strengthen incident management.
  • Lead significant technical or architectural decisions, including improvements to observability, disaster recovery, and infrastructure.
  • Identify and resolve technical or platform risks before they become delivery problems, including through disaster recovery and incident response exercises.
  • Establish a sustainable balance between hands-on technical contribution and people leadership while becoming a trusted technical authority for Site Reliability.
Requirements:
  • 4+ years of professional software engineering experience with strong, current hands-on coding experience.
  • 3+ years of professional platform engineering experience incorporating DevOps and Site Reliability Engineering principles.
  • Previous experience owning technical direction, people management, or both, with clear readiness to combine technical leadership and people leadership responsibilities.
  • Ability to serve as the final technical decision-maker for a team without relying on a separate Tech Lead.
  • Experience with distributed systems and modern infrastructure technologies such as AWS, Kubernetes/EKS, kops, HashiCorp Vault, Grafana, GitHub Actions, PostgreSQL, Kafka/MSK, Redis/Valkey, or comparable platforms.
  • Experience with Azure or GCP environments is also welcome.
  • Strong understanding of modern microservice architectures and engineering methodologies.
  • Experience with Golang, TypeScript, and React is a plus.
  • Experience with AI-assisted development tools and workflows, or a strong willingness and aptitude to adopt them.
  • Demonstrated ability to make pragmatic technical trade-offs while working under real delivery and operational pressure.
  • Experience with Agile or Scrum-based development processes, or willingness to help evolve team processes based on practical experience.
  • Strong written and verbal communication skills, with the ability to collaborate effectively across technical and non-technical stakeholders.
  • Strong coaching, mentoring, and people leadership capabilities.
  • Ability to balance strategic technical thinking with direct execution and hands-on engineering contribution.
Benefits:
  • National target base salary of $146,000--$183,000, depending on experience and skills.
  • Participation in an annual bonus and stock option program.
  • Fully flexible work model, with the option to work remotely, from the Pittsburgh office, or through a combination that suits your needs.
  • 100% employer-paid employee health plan, including vision, dental, and supplemental coverage.
  • Flexible Paid Time Off policy.
  • Work-from-home stipend to help create a personalized and effective home office.
  • 12 weeks of fully paid parental leave for all employees, plus short-term disability for birthing parents.
  • 401(k) plan with employer matching.
  • Opportunity to make an immediate impact on products used by millions of people.
  • High-ownership environment where technical ideas and contributions can directly influence engineering practices and company outcomes.
  • Remote interview process designed to provide a flexible and comfortable candidate experience.
  • Employment is available to candidates legally authorized to work in the United States without current or future sponsorship.
  • Hiring is currently limited to states where the organization already has employees; relocation assistance is not provided.

We appreciate your interest and wish you the best!

Data Privacy Notice:

By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Tech Lead & Product Engineer
Tech Lead & Product Engineer

Remote Raven • United States

Hybrid
USD 14,000 - 19,000
100% remote work
Full-time role
Long-term employment
Infrastructure Engineer
Infrastructure Engineer

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-Tier Compensation
Ownership & Autonomy
Opportunity to work at a high-growth startup
Engineering Lead | Onsite
Engineering Lead | Onsite

Worky • Dallas (TX)

On-site
USD 44,000 - 154,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SupportFinity™ • San Francisco (CA)

On-site
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

United States Digital Space LLC • Michigan

Hybrid
USD 150,000 - 190,000
Senior Manager, Forward Deployed Engineering
Senior Manager, Forward Deployed Engineering

United States Digital Space LLC • United States

Hybrid
USD 234,000 - 310,000
Equity
401(k) Retirement Savings Plan
Employee Stock Participation Plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LeanData Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BetterUp • Austin (TX)

Hybrid
USD 147,000 - 185,000
Access to BetterUp coaching
Medical, dental, and vision insurance
Flexible paid time off
+2