Lead Site Reliability Engineer

Avalara, Inc.

Town of Poland (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonuses
Paid time off
Medical insurance

Job summary

Avalara, Inc. is seeking a senior reliability engineer to lead how reliability is engineered across Avalara's global SaaS platform as we move toward an AI-first operating model.

You will build a modern, automation-first reliability ecosystem that improves stability, reduces risk, and speeds safe product delivery across multi-cloud environments. You will mentor engineers, shape standards, and drive measurable improvements in reliability and performance, including AI-driven monitoring and

Qualifications

  • 10+ years of experience in SaaS, distributed systems, or site reliability engineering.
  • Programming skills in Go, Java, or Python.
  • Deep experience with observability tools such as Prometheus, Grafana, and OpenTelemetry.
  • Hands-on experience with Kubernetes, containerisation, and multi-cloud platforms (AWS, GCP, Azure, or OCI).
  • Strong understanding of Linux systems, networking, and cloud-native architectures.
  • Proven ability to design automation, improve system reliability, and apply AI or machine learning to operational workflows.

Responsibilities

  • Own the reliability strategy for distributed SaaS systems across multi-cloud platforms.
  • Design and implement AI-driven operations, including predictive monitoring and automated root-cause analysis.
  • Build and scale observability using Prometheus, Grafana, and OpenTelemetry.
  • Create self-healing systems and automation to reduce manual work.
  • Improve deployments with feature flags, progressive delivery, and safe rollouts.
  • Ensure reliability of CI/CD pipelines and IaC environments.
  • Strengthen availability, scalability, and fault tolerance on Kubernetes platforms.
  • Lead incident response and drive post-incident improvements.
  • Integrate AI-driven workflows into incident detection and resolution.
  • Mentor engineers and champion automation-first reliability.

Skills

Go
Java
Python

Tools

Prometheus
Grafana
OpenTelemetry

Job description

What You’ll Do

You will lead how reliability is engineered across Avalara's global SaaS platform as we scale and move toward an AI-first operating model. You will focus on building a modern, automation-first reliability ecosystem that improves system stability, reduces operational risk, and enables faster, safer product delivery. You will work across multi-cloud environments to design self-healing systems, advance observability, and modernise deployment practices. As a senior individual contributor, you will also raise the technical bar by shaping standards, mentoring engineers, and driving measurable improvements in reliability and performance.

#LI-REMOTE

What Your Responsibilities Will Be
  • Own and evolve the reliability strategy for distributed SaaS systems across multi-cloud platforms
  • Design and implement AI-driven operations, including predictive monitoring, anomaly detection, and automated root cause analysis
  • Build and scale observability solutions using tools such as Prometheus, Grafana, and OpenTelemetry
  • Create self-healing systems and automation frameworks that reduce manual operational work
  • Improve deployment practices using feature flags, progressive delivery, and safe rollout strategies
  • Ensure reliability and performance of CI/CD pipelines and infrastructure as code environments
  • Strengthen system availability, scalability, and fault tolerance across Kubernetes-based platforms
  • Lead incident response, improve recovery times, and implement lasting fixes through post-incident reviews
  • Integrate AI-driven workflows into incident detection, triage, and resolution to improve operational efficiency
  • Mentor engineers and drive adoption of automation-first and AI-first reliability practices
What You'll Need to be Successful
  • 10+ years of experience in SaaS, distributed systems, or site reliability engineering
  • Programming skills in Go, Java, or Python
  • Deep experience with observability tools such as Prometheus, Grafana, and OpenTelemetry
  • Hands-on experience with Kubernetes, containerisation, and multi-cloud platforms (AWS, GCP, Azure, or OCI)
  • Strong understanding of Linux systems, networking, and cloud-native architectures
  • Proven ability to design automation, improve system reliability, and apply AI or machine learning to operational workflows
Avalara is an AI-first Company

AI is embedded in our workflows, decision-making, and products. Success here requires embracing AI as an essential capability.

  • You’ll bring experience using AI and AI-related technologies, ready to thrive here.
  • You’ll apply AI every day to business challenges - improving efficiency, contributing solutions, and driving results for your team, our company, and our customers.
  • You’ll grow with AI by staying curious about new trends and best practices, and by sharing what you learn so others can benefit too.
How We’ll Take Care of You
Total Rewards

In addition to a great compensation package, paid time off, and paid parental leave, many Avalara employees are eligible for bonuses.

Health & Wellness

Benefits vary by location but generally include private medical, life, and disability insurance.

Inclusive culture and diversity

Avalara strongly supports diversity, equity, and inclusion, and is committed to integrating them into our business practices and our organizational culture. We also have a total of 8 employee-run resource groups, each with senior leadership and exec sponsorship.

What You Need To Know About Avalara

We’re defining the relationship between tax and tech.

We’ve already built an industry-leading cloud compliance platform, processing over 54 billion customer API calls and over 6.6 million tax returns a year. Our growth is real - we're a billion dollar business - and we’re not slowing down until we’ve achieved our mission - to be part of every transaction in the world.

We’re bright, innovative, and disruptive, like the orange we love to wear. It captures our quirky spirit and optimistic mindset. It shows off the culture we’ve designed, that empowers our people to win. We’ve been different from day one. Join us, and your career will be too.

We’re An Equal Opportunity Employer

Supporting diversity and inclusion is a cornerstone of our company — we don’t want people to fit into our culture, but to enrich it. All qualified candidates will receive consideration for employment without regard to race, color, creed, religion, age, gender, national orientation, disability, sexual orientation, US Veteran status, or any other factor protected by law. If you require any reasonable adjustments during the recruitment process, please let us know.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Manager, Software Engineering Remote, United States - Engineering
Sr. Manager, Software Engineering Remote, United States - Engineering

Socotra, Inc. • California (MO)

Hybrid
USD 120,000 - 160,000
Bonuses
Health Insurance
Paid Time Off
+1
Senior Director, Engineering Remote, United States - Engineering - Software Engineering
Senior Director, Engineering Remote, United States - Engineering - Software Engineering

Socotra, Inc. • United States

Remote
USD 150,000 - 200,000
Paid time off
Paid parental leave
Health, life, and disability insurance
+1
Senior Full Stack Engineer, AI-Assisted Product Engineering
Senior Full Stack Engineer, AI-Assisted Product Engineering

avalara • United States

Hybrid
USD 120,000 - 150,000
Health and wellness benefits
Paid parental leave
Inclusive culture support
Sr. Director, Global Support Operations
Sr. Director, Global Support Operations

Avalara, Inc. • United States

On-site
USD 180,000 - 250,000
Bonuses
Private medical, life, and disability 
Employee resource groups (ERGs)
Sr. Software Development Engineer Remote, United States - Engineering
Sr. Software Development Engineer Remote, United States - Engineering

Socotra, Inc. • United States

Remote
USD 100,000 - 130,000
Paid time off
Health and wellness benefits
Potential bonuses
Staff Technical Account Manager Remote, United States - Engineering
Staff Technical Account Manager Remote, United States - Engineering

Socotra, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Sr. Director, Global Support Operations Remote, United States - Customer Service and Support - Global Support
Sr. Director, Global Support Operations Remote, United States - Customer Service and Support - Global Support

Socotra, Inc. • Northern (KY)

Hybrid
USD 150,000 - 200,000
Bonuses
Private medical insurance
Sr. Director, Global Support Operations
Sr. Director, Global Support Operations

Avalara, Inc. • Northern (KY)

Hybrid
USD 220,000 - 320,000
Staff Technical Account Manager
Staff Technical Account Manager

Avalara, Inc. • Northern (KY)

On-site
USD 90,000 - 130,000
Bonus eligibility
Private medical insurance
Diversity and inclusion
Staff Technical Account Manager
Staff Technical Account Manager

Avalara, Inc. • United States

On-site
USD 140,000 - 190,000
Bonus eligibility
Private medical insurance
Retirement benefits