Director, Global DevOps, Site Reliability, & Infrastructure (Vancouver, B.C. or Austin, TX)

BitKernel

Austin (TX)

On-site

USD 160,000 - 200,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BitKernel seeks a Director of DevOps, Site Reliability & Infrastructure to lead reliability, performance, security, and operational maturity of our SaaS platform. This is a high-visibility leadership role with ownership across DevOps, SRE, Infrastructure Operations, and Front-Line Technical Support.

You will advise executives, lead critical incidents, improve observability, modernize infrastructure, and drive cost-conscious, scalable solutions aligned with growth.

Qualifications

  • 10+ years in Infrastructure Operations, DevOps, SRE, or related fields.
  • 5+ years in a leadership role managing tech and support teams.
  • Proven experience with mission-critical SaaS platforms and large-scale cloud/hybrid environments.
  • Strong knowledge of Linux, networking, security, automation, and monitoring.

Responsibilities

  • Lead DevOps, SRE, Infrastructure, and Front-Line Support teams with clear ownership.
  • Own platform availability, performance, reliability, and service delivery.
  • Lead incident management, root cause analysis, and post-incident reviews.
  • Define SLIs/SLOs, KPIs, and executive reporting for operational health.
  • Develop and execute a future-state infrastructure roadmap aligned with business needs.
  • Oversee budgeting, forecasting, and cost controls for infra spend.
  • Mentor teams and build a culture of ownership and continuous improvement.
  • Collaborate with Engineering, Product, Security, and Finance leadership.

Skills

Leadership
Incident management
Strategic planning
Cross-functional collaboration
Communication

Education

Bachelor's degree in Computer Science or related

Tools

Terraform
Kubernetes
CloudFormation
Pulumi

Job description

Role Overview

We’re looking for an experienced Director of DevOps, Site Reliability & Infrastructure to lead the teams responsible for the reliability, performance, security, and operational maturity of our SaaS platform.

This is a highly visible leadership role with ownership across DevOps, Site Reliability Engineering (SRE), Infrastructure Operations, and Front-Line Technical Support. You’ll be responsible for keeping our current environment stable and performing at a high level while building the roadmap, operating practices, and infrastructure needed to scale.

We’re looking for a leader who can move comfortably between strategy and execution—someone who can advise executives, lead teams through critical incidents, improve observability and operational processes, modernize infrastructure, and make smart decisions about where and how we invest.

Success means creating an operation that is more reliable, scalable, measurable, cost-efficient, and predictable.

What You’ll Own
Reliability & Service Operations
  • Lead DevOps, SRE, Infrastructure, and Front-Line Support teams and establish clear ownership, accountability, and operating standards.
  • Own platform availability, performance, reliability, operational readiness, and service delivery.
  • Lead incident management, root cause analysis, corrective actions, and post-incident reviews.
  • Establish and continuously improve SLIs, SLOs, KPIs, service-level expectations, operational scorecards, and executive reporting.
  • Create clear escalation paths across Support, SRE, DevOps, Engineering, Product, and other technical teams.
  • Improve Tier 1/Tier 2 support effectiveness, including resolution rates, escalation quality, MTTR, and customer-impacting incidents.
  • Build and maintain effective SOPs, runbooks, troubleshooting guides, and operational documentation.
Infrastructure & Platform Evolution
  • Own the operation, maintenance, and evolution of our cloud, hybrid, and bare-metal infrastructure.
  • Develop and execute a pragmatic future-state infrastructure roadmap aligned with business growth and technical requirements.
  • Lead modernization initiatives that improve scalability, resilience, maintainability, and operational efficiency.
  • Oversee high availability, disaster recovery, backup, business continuity, and capacity planning.
  • Establish effective infrastructure governance while balancing performance, security, reliability, cost, and complexity.
  • Define and execute our monitoring and observability strategy across infrastructure, applications, platforms, and customer experience.
  • Identify visibility gaps and improve alerting, issue detection, diagnosis, escalation, and resolution.
  • Expand Infrastructure as Code, deployment automation, and operational automation to reduce manual work and improve consistency.
  • Use metrics, retrospectives, and operational data to continuously improve reliability and team effectiveness.
  • Evaluate emerging technologies, including AI-driven operational tools and AIOps, where they can meaningfully improve performance or efficiency.
  • Own infrastructure and operational budgeting, forecasting, cost controls, and financial planning.
  • Improve visibility into cloud, hosting, licensing, monitoring, support tooling, and other technology expenditures.
  • Apply FinOps and cloud-governance principles to improve utilization and financial accountability.
  • Identify and eliminate waste, overprovisioning, unused resources, and inefficient technology spending.
  • Partner with Finance and Executive Leadership on future infrastructure investments and operating expenses.
Leadership & Operational Development
  • Build, mentor, and develop high-performing DevOps, SRE, Infrastructure, and Support teams.
  • Establish clear roles, expectations, performance metrics, and development plans.
  • Recruit and retain strong technical talent while building a culture of ownership, accountability, collaboration, and continuous improvement.
  • Create clarity and momentum in a fast-changing environment.
  • Build strong partnerships across Engineering, Product, QA, Security, Customer Success, Finance, and Executive Leadership.
  • Communicate operational health, risks, priorities, investments, and progress clearly to technical and executive audiences.
What We’re Looking For...

Required

  • 10+ years of experience across Infrastructure Operations, DevOps, SRE, Cloud Operations, or related disciplines.
  • 5+ years of leadership experience managing engineering, DevOps, infrastructure, SRE, support, or service-delivery teams.
  • Proven experience operating and supporting mission-critical SaaS platforms.
  • Strong experience operating production cloud, hybrid, and/or bare-metal environments at scale.
  • Deep understanding of Linux, networking, security, virtualization, storage, databases, infrastructure automation, and systems operations.
  • Demonstrated experience improving monitoring, observability, incident response, operational maturity, and service reliability.
  • Experience establishing effective SOPs, runbooks, operational standards, and support processes.
  • Experience managing technology budgets, infrastructure costs, vendors, capacity, and operational expenditures.
  • Strong planning, prioritization, decision-making, and organizational skills.
  • Excellent communication skills with the ability to work effectively with engineers, business leaders, executives, partners, and customers.
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
Bonus Points for Experience in several of the following:
  • DevOps, SRE, or infrastructure transformation
  • Infrastructure as Code such as Terraform, Pulumi, Bicep, or CloudFormation
  • CI/CD and deployment automation
  • Enterprise monitoring and observability platforms
  • Kubernetes, containerized workloads, and platform engineering
  • FinOps, cloud governance, and cost allocation
  • Global SaaS platforms and 24x7 production environments
  • AI-driven operational tooling, automation platforms, or AIOps
  • Scaling technical service organizations through periods of business growth or transformation
What Success Looks Like...

During your first 12 months, you’ll be expected to:

  • Improve the stability and effectiveness of DevOps, SRE, Infrastructure, and Front-Line Support operations.
  • Establish clear ownership, operating standards, KPIs, and executive visibility.
  • Deliver and begin executing a measurable monitoring and observability roadmap .
  • Strengthen incident response, service reliability, uptime, and operational readiness.
  • Improve SOP and runbook coverage for critical systems and services.
  • Define and execute a future-state infrastructure roadmap aligned with business priorities.
  • Establish stronger infrastructure cost controls, forecasting, governance, and optimization.
  • Strengthen collaboration across technical and business teams.
  • Build confidence in the organization through predictable planning, clear communication, disciplined execution, and strong operational leadership .
The Opportunity

This role is ideal for a technology leader who enjoys solving complex operational problems and building organizations that scale. You’ll have the opportunity to influence not only how our infrastructure operates today, but how we technology organization evolves for the future —strengthening reliability, modernizing our platform, developing our teams, and creating the operational foundation for continued growth

#LI-ATX

About Bitkernel

Streaming is having its moment — and BitKernel is building the infrastructure powering it. We're a global tech company on a mission to make world-class streaming accessible to anyone with a great video idea, from scrappy startups to established media brands.

We've spent years in the trenches of IPTV, OTT, cloud TV, billing, and video distribution — and we've taken all of that hard-won expertise and rebuilt it into something faster and smarter: a next-generation streaming SaaS platform.

In November 2025, we launched Goose , our flagship product that lets enterprises and SMEs spin up white-label streaming platforms . And we didn't stop there — in early 2026, we shipped interactive streaming tools built for the next generation of video business models, because the way people watch (and engage) is changing fast, and we intend to lead the way.

We're growing quickly and building a team that moves fast, thinks big, and isn't afraid to challenge how streaming "has always been done." If that sounds like you, let's talk.

The pay range for this role is:

160,000 - 200,000 USD per year (HQ - Austin)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, Global DevOps, Site Reliability, & Infrastructure (Vancouver, B.C. or Austin, TX)
Director, Global DevOps, Site Reliability, & Infrastructure (Vancouver, B.C. or Austin, TX)

BitKernel Technology Inc. • Austin (TX)

On-site
USD 160,000 - 200,000
Director, Global DevOps, Site Reliability, & Infrastructure (Vancouver, B.C. or Austin, TX)
Director, Global DevOps, Site Reliability, & Infrastructure (Vancouver, B.C. or Austin, TX)

Bitkernel Technology Inc • Austin (TX)

On-site
USD 160,000 - 200,000
Global DevOps & SRE Director: Scale & Reliability
Global DevOps & SRE Director: Scale & Reliability

Bitkernel Technology Inc • Austin (TX)

On-site
USD 160,000 - 200,000
Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Relocation assistance
Learning and growth opportunities
+1
Forward Deployed Engineer (Staff)
Forward Deployed Engineer (Staff)

Socket.dev • California (MO)

Hybrid
USD 111,000 - 177,000
Discretionary annual bonus
Equity grants
401(k) with company match
+2
Staff Engineer - Core Infrastructure
Staff Engineer - Core Infrastructure

United States Digital Space LLC • San Francisco (CA)

On-site
USD 232,000 - 258,000
401(k) plan
Bonus program
Equity award
Senior Infrastructure Engineer (Core Infra, US)
Senior Infrastructure Engineer (Core Infra, US)

Workato • Charlotte (NC)

On-site
USD 180,000 - 240,000
Health insurance
Paid time off
401(k) plan
Sr. Platform Engineer
Sr. Platform Engineer

The Bronx • Austin (TX)

Hybrid
USD 114,000 - 190,000
Hybrid schedule
Relocation assistance
Benefits package
+1
Head of Infrastructure - Base up to $290k
Head of Infrastructure - Base up to $290k

CyberCoders • New York (NY)

On-site
USD 225,000 - 290,000
Unlimited PTO
Wellhub Platinum membership
Comprehensive health, dental, and vis​
+1
Senior Software Engineer, Platform Infrastructure
Senior Software Engineer, Platform Infrastructure

Roku • San Jose (CA)

Hybrid
USD 280,000 - 380,000
Health insurance
Equity awards
Disability benefits
+2