Lead Operations Engineer

CacheFly

United States

Remote

CAD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote-first culture
Annual bonus 15%
Unlimited PTO
Health benefits
401(k) with company match
Work-from-home stipend

Job summary

CacheFly is seeking a hands-on, technically elite Operations leader to guide our global CDN platform. You will run the day-to-day operations, provision servers and PoPs, perform upgrades, and manage configurations with a focus on reliability and speed.

You will lead a team from day one as its technical anchor, with potential for people management as the role matures. Strong Linux, networking, and scripting skills are essential for success in a remote-first, 24/7 environment.

Qualifications

  • In-depth Linux system administration experience at scale.
  • Deep understanding of HTTP and caching.
  • Strong networking knowledge — routing and TCP/IP.
  • Proven ability to troubleshoot and resolve complex, customer-facing issues.
  • Reverse‑engineer system architectures and document them for the team.
  • Clear, professional communication with customers under pressure.
  • Hands-on experience with Nginx, HAProxy, Varnish, BGP, anycast, and CentOS/RHEL/Rocky Linux.
  • Scripting in Bash, Python, or PHP, or similar.
  • Experience operating in high-availability, latency-sensitive environments.

Responsibilities

  • Serve as the primary technical resource and senior escalation point for the Operations team
  • Work hands-on across provisioning, upgrades, configuration management, performance work, and customer support
  • Lead the response on large-scale incidents and oversee the handling of smaller ones
  • Own communication with large and strategic customers during incidents and escalations — clearly, calmly, and professionally under pressure
  • Author incident post-mortems and drive actionable follow-ups and process improvements
  • Continuously improve CDN reliability, performance, and operational excellence across the platform
  • On-call & Escalations: 24/7 coverage; top technical link in escalation chain; defined path

Skills

Linux administration
HTTP & caching
Networking
Troubleshooting
Documentation
Scripting

Tools

Nginx
HAProxy
Varnish
BGP
Anycast
CentOS/RHEL/Rocky
Bash
Python
PHP

Job description

About CacheFly

CacheFly powers fast, secure content delivery across the globe. We're a performance-driven, remote-first technology company that values ownership, accountability, and speed — and we hold our infrastructure to the same standard. Our global CDN serves customers who depend on us around the clock, which means the team behind it needs to be elite. We're small enough that every person has real impact, and serious enough about reliability that we never stop raising the bar.

The Role

CacheFly is hiring the technical leader of our Operations team, the people who keep our global CDN platform running, fast, and reliable around the clock. This person will be a hands‑on, technically elite operator who leads from the front and sets the standard by doing.

You are the most capable operator in the room, deep in Linux systems, HTTP and caching internals, and the network itself, troubleshooting, tuning, automating, and resolving the hardest problems alongside the team rather than from a distance.

The Operations team is the backbone of CacheFly's CDN. The team runs the day‑to‑day platform, provision servers and PoPs, upgrade components and fleets, and manage service configuration across the network. We obsess over configuration consistency and continuously raise the performance baseline — finding bottlenecks and degradation before they become incidents. We also own customer support tickets and partner with pre‑sales on proof‑of‑concept engagements.

You will lead all of it. The full Operations team reports to you, and you report directly to the VP of Engineering.

You'll lead a team from day one as its technical anchor. Formal people management is not the focus of this role right now, but it will become part of the job as the role matures. We expect you to develop that muscle over time, with support from technical leadership and HR as you ramp. If you've never managed before but want to, this role is designed for that path.

Key Responsibilities
  • Serve as the primary technical resource and senior escalation point for the Operations team
  • Work hands‑on across all core responsibilities — provisioning, upgrades, configuration management, performance work, and customer support
  • Lead the response on large‑scale incidents and oversee the handling of smaller ones
  • Own communication with large and strategic customers during incidents and escalations — clearly, calmly, and professionally under pressure
  • Author incident post‑mortems and drive the actionable follow‑ups and process improvements that come out of them
  • Continuously improve CDN reliability, performance, and operational excellence across the platform
  • On‑call & Escalations: we provide 24/7 coverage for incidents and customer support, and you are the top technical link in the escalation chain. You will take escalations outside of normal working hours. Major incidents will reach you when they happen, not when it's convenient. This is a defining part of the role, and we want candidates who go in clear‑eyed about that. In exchange, you have a real escalation path of your own — you're never carrying a major event alone. Escalation path is as follows: Operations team > Lead Operations Engineer > VP of Engineering > CTO
Who You Are
  • In‑depth Linux system administration experience at scale
  • Deep understanding of HTTP and caching
  • Strong networking knowledge — routing and TCP/IP
  • Proven ability to troubleshoot and resolve complex, customer‑facing issues
  • Ability to reverse‑engineer system architectures with limited documentation and document them for the team
  • Clear, professional communication with customers under pressure
  • Hands‑on experience with Nginx, HAProxy, Varnish, BGP, anycast, and CentOS / RHEL / Rocky Linux
  • Proficiency in a scripting language (Bash, Python, PHP, or similar)
  • Experience operating in high‑availability, latency‑sensitive environments
Benefits
  • People‑First Culture; While we're obsessed with performance, we're also deeply committed to taking care of our people through fair compensation, competitive global benefits, and a culture of mutual respect and authenticity
  • Competitive salary and annual bonus (15% target)
  • Cutting‑edge technology and resources to help you succeed in your role
  • Remote‑first culture
  • Unlimited PTO and all National holidays off
  • $3,000 work‑from‑home stipend, monthly internet/phone or coworking allowance
  • Full health benefits (medical, dental, vision, LTD/STD, life) - most at no cost to you*
  • 401(k) with 4% Safe Harbor company match, fully vested from day one*
  • Benefits marked with an asterisk (*) are available to U.S.-based employees. For team members outside the United States, equivalent value is incorporated into total compensation and country‑specific benefits where applicable
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Linux Systems Admin
Senior Linux Systems Admin

CacheFly • United States

Remote
CAD 80,000 - 120,000
Unlimited PTO
Remote-first culture
$3,000 annual work-from-home stipend
Senior Lead Operations Engineer - CDN Reliability (Remote)
Senior Lead Operations Engineer - CDN Reliability (Remote)

CacheFly • Washington

On-site
USD 180,000 - 240,000
Remote-first culture
Unlimited PTO and holidays
$3,000 work-from-home stipend
+2
Marketing Manager
Marketing Manager

CacheFly • United States

On-site
USD 83,000 - 103,000
Remote-first culture
Annual work-from-home stipend
Full health benefits
+1
Senior Linux Systems Admin — CDN (Remote, Unlimited PTO)
Senior Linux Systems Admin — CDN (Remote, Unlimited PTO)

CacheFly • United States

Remote
CAD 80,000 - 120,000
Unlimited PTO
Remote-first culture
$3,000 annual work-from-home stipend
Software Engineer, CDN Content
Software Engineer, CDN Content

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 172,000 - 258,000
Equity
Healthcare
Mentorship & events
+2
Senior Cloud Engineer
Senior Cloud Engineer

Fastly • San Francisco (CA)

Hybrid
USD 122,280 - 172,632
Hybrid work model
Equity program
Discretionary bonus
+3
DevOps I – Linux, AWS, Release Management
DevOps I – Linux, AWS, Release Management

Partner's Consulting, Inc. • Englewood (CO)

On-site
USD 90,000 - 120,000
Engineering Manager - Workers Runtime, Edge Team
Engineering Manager - Workers Runtime, Edge Team

Cloudflare • Austin (TX)

Hybrid
USD 210,000 - 320,000
Equity
Health insurance
401(k) matching
+1
Engineering Manager
Engineering Manager

Webhosting • Austin (TX)

Hybrid
USD 210,000 - 320,000
Equity plan
Data Center Engineer
Data Center Engineer

United States Digital Space LLC • United States

Hybrid
USD 85,000 - 125,000