Lead Site Reliability Engineer New

Mattermost, Inc.

Northern (KY)

Hybrid

USD 145,000 - 200,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mattermost is seeking a Lead Site Reliability Engineer to guide reliability and operational excellence for our secure collaboration platform. You will shape the SRE strategy, drive scalable, compliant cloud deployments, and mentor a globally distributed team in a remote-first environment.

Ideal candidates bring 5+ years in SRE/DevOps, strong Kubernetes and Terraform experience, and a track record of automation, observability, and cross-functional collaboration with security and product teams.

Qualifications

  • BS in CS, cybersecurity, software engineering, or equivalent experience.
  • 5+ years in site reliability engineering, DevOps, or cloud infra.
  • Proven Kubernetes expertise with cloud tooling.
  • Experience with Terraform and AWS in large-scale environments.
  • Strong incident management and on-call discipline.
  • Proven remote-first leadership across distributed teams.

Responsibilities

  • Define strategy, architecture, and roadmap for Mattermost's SRE function.
  • Lead design, deployment, and optimization of containerized workloads and IaC.
  • Establish observability, monitoring, and alerting at scale.
  • Drive incident management, root cause analysis, and reliability improvements.
  • Partner with security/compliance to meet FedRAMP/DoD requirements.
  • Champion automation and operational excellence.
  • Oversee cloud cost management and capacity planning.
  • Build a developer platform for fast, secure delivery.
  • Mentor SREs and foster a culture of learning.

Skills

Kubernetes
Terraform
AWS
Monitoring
Incident response
DevOps
Scripting
Leadership
Remote team leadership

Education

BS in CS/related field

Tools

Grafana
Prometheus
CI/CD (Jenkins)

Job description

Mattermost is the leading collaborative workflow platform for defense, intelligence, security, and critical infrastructure. Trusted by the U.S. Department of War and Fortune 500s, our platform runs on-premises and in private clouds, delivering secure messaging, file sharing, workflow automation, audio/screenshare, and project management—all with full data and operational control. Mattermost powers high-stakes workflows across mission planning, real-time, real-world operations, DevSecOps, incident response, and cyber defense—enabling secure collaboration from tactical edge and DDIL environments to enterprise HQ. Teams operate across web, desktop, and mobile, with embedded interoperability for Microsoft Teams, Outlook, and Microsoft 365.

Mattermostis seeking an experienced and visionary Lead Site Reliability Engineer (SRE)to guide the architecture, reliability, and operational excellence of the infrastructure powering our secure, mission-critical collaboration platform.

In this role, you will provide technical leadership across our SRE function, driving strategic initiatives for scalability, observability, performance, and automation across cloud and hybrid environments. You will mentor engineers, establish best practices, and collaborate closely with development, security, and operations teams to ensure our customers in defense, government, and critical infrastructure sectors experience exceptional reliability and performance.

Responsibilities Include:

  • Define the strategy, architecture, and roadmap for Mattermost’s site reliability engineering function, aligning infrastructure initiatives with product and business goals.
  • Lead the design, deployment, and optimization of production-grade containerized workloads, infrastructure-as-code, and compliant cloud environments for regulated domains (e.g., FedRAMP, DoD).
  • Establish and evolve observability, monitoring, and alerting frameworks to ensure performance, reliability, and capacity planning at scale.
  • Drive incident management processes, including on-call rotations, root cause analysis, and systemic reliability improvements.
  • Partner with security and compliance teams to meet data sovereignty, security, and regulatory requirements.
  • Champion automation and operational excellence to improve efficiency, reduce risk, and scale operations.
  • Oversee cloud cost management and capacity planning to optimize infrastructure spending while meeting performance targets.
  • Build and maintain a developer platform that enables fast, secure software delivery and improves application stability in production.
  • Mentor and coach SRE team members, fostering a culture of learning, collaboration, and technical excellence.

Requirements:

  • BS in Computer Science, Cybersecurity, Software Engineering, or a related technical field, or equivalent experience, with 5+ years of relevant experience in site reliability engineering, DevOps, or cloud infrastructure roles.
  • Proven expertise in container orchestration platforms, ideally Kubernetes.
  • Extensive experience with infrastructure-as-code, ideally Terraform.
  • Strong background in cloud platforms, ideally AWS.
  • Demonstrated experience designing and implementing monitoring, alerting, and performance optimization strategies.
  • Exceptional troubleshooting and incident management skills for distributed systems.
  • Proficiency in at least one scripting or programming language for automation.
  • Excellent communication skills with a track record of influencing cross-functional teams.
  • Experience leading globally distributed teams in a remote-first environment.

Preferences:

  • Familiarity with observability stacks such as Grafana and Prometheus.
  • Experience designing high-availability, disaster recovery, and scaling architectures.
  • Exposure to GCP and Azure cloud environments.
  • Leadership experience in highly regulated industries such as defense, finance, or critical infrastructure.
  • Experience with U.S. federal compliance frameworks and authorization processes, including FedRAMP, DoD ATO, NIST 800-53, and related government standards.
  • Experience preparing, delivering, and maintaining software offerings through AWS Marketplace and other cloud provider marketplaces (e.g., Azure Marketplace, Google Cloud Marketplace), including packaging, compliance validation, and ongoing operational support.
  • Open-source contributions in reliability, DevOps, or infrastructure tooling.
  • Certifications in cloud infrastructure, reliability, or DevOps engineering (e.g., CKA, CKAD, AWS Certified Solutions Architect).

Compensation

Salary range:$145,000 – $200,000

Mattermost takes a market-based approach to pay. Compensation is determined based on skills, experience, qualifications, and work location. Ranges may be updated as market conditions evolve.

Mattermost is an EEO Employer, we are a remote-first, open-source company.

We are continually working to expand our hiring in more countries and regions, ensuring compliance with local laws and regulations, which takes time.

Mattermost values your unique perspective—we welcome all applicants. We encourage individuals from all backgrounds to apply and are committed to assessing candidates based on their skills and qualifications. We do not tolerate discrimination against staff or applicants based on race, religion, national origin, age, disability, pregnancy status, veteran status, or other personal characteristics.

If you require accommodations during the interview process, please let us know—we’re happy to assist.

Voluntary Self-Identification of Disability

Form CC-305

Page 1 of 1

OMB Control Number 1250-0005

Expires 04/30/2026

We use Greenhouse’s AI-powered Talent Matching tool to compare your application against our job requirements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Full Stack Engineer New
Senior Full Stack Engineer New

Mattermost, Inc. • Northern (KY)

Hybrid
USD 140,000 - 180,000
Remote-first culture
Open-source at the core
Autonomy and ownership
Remote Lead Site Reliability Engineer — Mission-Critical
Remote Lead Site Reliability Engineer — Mission-Critical

Mattermost, Inc. • Northern (KY)

Hybrid
USD 145,000 - 200,000
Staff Software Engineer, Testing Infrastructure
Staff Software Engineer, Testing Infrastructure

Mattermost • United States

Remote
USD 144,000 - 200,000
Senior Account Executive, Navy Sector
Senior Account Executive, Navy Sector

Mattermost • San Diego (CA)

Remote
USD 195,000 - 260,000
Open-source company culture
Remote work flexibility
Diverse and inclusive work environment
Senior Account Executive, Navy Sector
Senior Account Executive, Navy Sector

Mattermost • San Diego (CA)

Remote
USD 195,000 - 260,000
Competitive salary
Remote-first work environment
Health and wellness benefits
Lead Site Reliability Engineer, Platforms
Lead Site Reliability Engineer, Platforms

Zoom • San Jose (CA), Northern (KY)

Hybrid
USD 124,000 - 271,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Lead DevOps (Site Reliability Engineer)
Lead DevOps (Site Reliability Engineer)

Anza Mortgage Insurance Corporation • Wilmington (DE)

On-site
USD 140,000 - 200,000
Competitive pay
Comprehensive benefits
401(k) plan
+3
Senior Engineering Manager, Site Reliability
Senior Engineering Manager, Site Reliability

Horizon3.ai • United States

On-site
USD 260,000 - 280,000
Hybrid & Remote Work
Competitive Compensation
Equity package
Site Reliability Engineer, Lead
Site Reliability Engineer, Lead

Booz Allen Hamilton • Chantilly (VA)

On-site
USD 99,000 - 225,000