Staff Site Reliability Engineer

Filevine

United States

Remote

USD 235,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
Maternity/Paternity Leave
Disability Insurance
Leadership mentorship
Company swag

Job summary

Filevine is seeking a Staff Site Reliability Engineer to shape reliability culture, set technical standards for production, and align business goals with scalable engineering execution. You will own Observability & Alerting and Platform Infrastructure, guiding SRE strategy and tooling for durable reliability across the org.

You will mentor engineers, partner with leadership on major decisions, and drive AI/ML-driven reliability practices while ensuring uptime, incident response, and change

Qualifications

  • 12+ years in software engineering/sre/infra with 6+ years in SRE leadership.
  • Expert in observability, incident response, and automation.
  • Experience with Kubernetes and cloud platforms.
  • Proven ability to mentor engineers and communicate risk to stakeholders.
  • Experience in regulated environments preferred.

Responsibilities

  • Define and execute strategies for Observability & Alerting and Platform Infrastructure.
  • Lead evolution of scalable cloud platforms and distributed systems.
  • Champion SLIs/SLOs, error budgets, and capacity planning.
  • Lead major production incidents and drive permanent improvements.
  • Build self-service platforms to reduce toil and boost safety.
  • Mentor engineers and set reliability direction.

Skills

SRE leadership
Observability expertise
Cloud infrastructure
Python/Go/Bash
Mentoring engineers
Communication with execs
Incident response
Capacity planning

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
New Relic
Datadog
Production tooling

Job description

Role Summary

As a Staff Site Reliability Engineer at Filevine, you are the senior technical authority on the SRE team and a strategic partner to engineering leadership. You don’t just maintain systems - you shape engineering culture, define the technical standard for how Filevine runs in production, and bridge the gap between high-level business goals and robust, internet-scale technical execution. You bring a forward-looking perspective - actively shaping how AI and machine learning drive the future of reliability practice. You own the roadmap across two critical SRE domains - Observability & Alerting and Platform Infrastructure - and are accountable for ensuring the team solves reliability problems permanently rather than absorbing them as toil. You operate as the senior IC counterpart to the Engineering Manager: technical correctness lives with you. You partner with the Reliability Architect and engineering leadership on significant technical decisions, mentor engineers across experience levels, and influence reliability strategy across the broader organization. Reliability at Filevine protects revenue. You are the senior technical voice responsible for ensuring that uptime, incident response, and every production change meet the operational standard the business demands. This role does not participate in on-call rotation, but you are deeply invested in the engineers who do - shaping the on-call strategy, tooling, and culture that make production support sustainable and effective.

Who You Are
  • Master of the Craft: You bring deep expertise in distributed systems, cloud infrastructure, observability, and reliability engineering. You raise the technical standard for every engineer around you and thrive where the challenges are complex and the stakes are real.
  • Technical Leader and Mentor: You are passionate about mentoring engineers and investing in their growth. You influence technical direction and communicate production risk clearly across engineering, product, and executive audiences.
  • Forward-Thinking & AI/ML Fluent: You bring deep knowledge of AIOps and drive the use of AI and machine learning in observability, anomaly detection, incident response, automated remediation, and resource optimization.
  • Production-Scale Problem Solver: You turn ambiguous, complex reliability challenges into durable solutions for systems where availability, performance, and production changes carry meaningful business impact.
  • Software-Minded Builder: You use software, automation, Infrastructure as Code, and platform capabilities to eliminate toil and make systems safer, more scalable, and easier to operate.
What you will do
  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.
  • Lead the organization through complex production incidents and turn post-incident learning into permanent engineering improvements.
  • Build self-service platform capabilities that reduce toil, improve engineering safety and velocity, and make every team more capable of owning their own reliability.
  • Mentor engineers and serve as a trusted technical authority for long-term reliability and platform direction.
Qualifications
  • 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE, including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives for distributed production systems.
  • Expert-level depth in observability and platform infrastructure, with broad expertise in incident response, capacity planning, automation, and reliability engineering.
  • Advanced experience with a major container-orchestration platform, preferably Kubernetes, and an observability platform such as New Relic, Datadog, or equivalent.
  • Strong software-engineering ability in Python, Go, Bash, or another general-purpose language, with experience building production tooling, automation, or platform capabilities.
  • Proven ability to mentor engineers and communicate technical risk clearly to engineering, product, and executive audiences.
  • Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is strongly preferred.

$235,000 - $275,000 a year

Cool Company Benefits:
  • A dynamic, rapidly growing company, focused on helping organizations thrive
  • Medical, Dental, & Vision Insurance (for full-time employees)
  • Competitive & Fair Pay
  • Maternity & paternity leave (for full-time employees)
  • Short & long-term disability
  • Opportunity to learn from a dedicated leadership team
  • Top-of-the-line company swag
Privacy Policy Notice

Filevine will handle your personal information according to what’s outlined in our Privacy Policy.

Communication about this opportunity, or any open role at Filevine, will only come from representatives with email addresses using "filevine.com". Other addresses reaching out are not affiliated with Filevine and should not be responded to.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VP of Engineering, Reliability
VP of Engineering, Reliability

Filevine • United States

On-site
USD 260,000 - 360,000
Medical, Dental, & Vision Insurance
Competitive & Fair Pay
Maternity & Paternity Leave
+2
Sr. SRE II ›
Sr. SRE II ›

Filevine • United States

On-site
USD 120,000 - 160,000
Medical, Dental, & Vision Insurance
Maternity & paternity leave
Short & long-term disability
+2
Site Reliability Engineer (AI Forms Platform)
Site Reliability Engineer (AI Forms Platform)

Filevine • United States

On-site
USD 100,000 - 140,000
Medical, Dental, & Vision Insurance
Competitive & Fair Pay
Maternity & paternity leave
+2
Software Architect I
Software Architect I

Filevine • United States

On-site
USD 250,000 - 300,000
Medical, Dental, & Vision Insurance (F
Competitive & Fair Pay
Maternity & paternity leave
+3
Senior Database Reliability Engineer ›
Senior Database Reliability Engineer ›

Filevine • United States

On-site
USD 150,000 - 210,000
Medical, Dental, & Vision Insurance
Maternity & paternity leave
Short & long-term disability
+2
Software Engineering Manager ›
Software Engineering Manager ›

Filevine • United States

On-site
USD 185,000 - 212,000
Medical, Dental, & Vision Insurance (F
Parental leave
Disability insurance
+1
Senior Database Reliability Engineer
Senior Database Reliability Engineer

Filevine • United States

On-site
USD 145,000 - 180,000
Medical, Dental, & Vision Insurance
Paid time off
Maternity & paternity leave
+2
Senior Software Development Engineer ›
Senior Software Development Engineer ›

Filevine • United States

On-site
USD 140,000 - 180,000
Medical, Dental, & Vision Insurance (f
Maternity & paternity leave
Short & long-term disability
+2
Senior SRE Architect: AI-Driven Reliability & Platform
Senior SRE Architect: AI-Driven Reliability & Platform

Filevine • United States

Remote
USD 235,000 - 275,000
Medical Insurance
Dental Insurance
Vision Insurance
+4
Architect I ›
Architect I ›

Filevine • United States

On-site
USD 250,000 - 300,000
Medical, Dental, & Vision Insurance
Maternity & paternity leave
Short & long-term disability
+2