Principal SRE, Infrastructure & Platform

F5 Networks, Inc

United States

On-site

USD 140,000 - 210,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

F5 Networks, Inc. is seeking a Principal SRE to design, deploy, and operate a large-scale infrastructure spanning 30+ PoPs across the Americas, EMEA, and APAC. You will own physical provisioning, Proxmox clusters, container workloads, and automation in a hybrid cloud/on-prem environment.

You will participate in 24x7 on-call, drive incident resolution, and contribute to security hardening and audit readiness in a PCI-DSS compliant setting.

Qualifications

  • 7+ years in a Site Reliability, DevOps, or infra engineering role in production.
  • Strong Linux systems administration (RHEL/CentOS).
  • Proficiency with Ansible for large-scale CM/CI integration.
  • Hands-on hypervisor experience (Proxmox VE preferred).
  • Production on-premise Kubernetes experience.
  • AWS/Azure experience including IAM, VPC, DNS.
  • Networking fundamentals incl. VLANs, routing, BGP, DNS.

Responsibilities

  • Design, deploy, and operate a large multi-datacenter platform.
  • Own provisioning from bare metal to app layer, incl. Proxmox clusters.
  • Develop CI/CD pipelines and automation across regions.
  • Manage secrets with Vault and PKI integration.
  • Maintain CMDB/IPAM accuracy in NetBox.

Skills

7+ years SRE/DevOps infra
Linux admin
Ansible
Proxmox VE
Kubernetes
AWS
Azure
CI/CD (GitLab)
Vault/Secret mgmt
Networking fundamentals

Tools

Proxmox VE
HashiCorp Vault
Kubernetes
Docker
GitLab CI
AWS
Azure
Ansible
BIND9
Unbound

Job description

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.

Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.

F5 is bringing a better digital world to life by helping organizations create, secure, and run applications that power our lives. Within the Platform Engineering team, this role helps ensure our platform is operated safely, reliably, and with operational excellence.

We’re looking for a Principal SRE who leads with kindness, operates well in a global, follow-the-sun environment, and brings strong execution, documentation, and cross-functional coordination skills. You will be responsible for designing, deploying, and operating the foundational infrastructure that underpins a large-scale, multi-datacenter platform spanning 30+ Points of Presence across the Americas, EMEA, and APAC.

This is a hands-on engineering role. You will own systems from bare metal to application layer – provisioning physical servers via out-of-band management, building and maintaining Proxmox-based hypervisor clusters, managing containerized workloads, and driving automation across a heterogeneous on-premises and cloud environment. You will operate within a PCI-DSS compliant environment and be expected to contribute to hardening, audit readiness, and security tooling.

You will join an on-call rotation and be expected to respond to and lead incident resolution for production systems across a 24x7 global environment.

What You'll Do

Infrastructure Automation & Configuration Management

  • Author, maintain, and refactor Ansible playbooks and roles across a large-scale multi-datacenter inventory, covering the full lifecycle from bare-metal provisioning to application deployment

Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions

  • Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions

Manage secrets lifecycle using HashiCorp Vault, including AppRole authentication, secret rotation, and PKI integration

  • Manage secrets lifecycle using HashiCorp Vault, including AppRole authentication, secret rotation, and PKI integration

Maintain CMDB/IPAM accuracy in NetBox as a source of truth for all infrastructure assets

  • Maintain CMDB/IPAM accuracy in NetBox as a source of truth for all infrastructure assets

Compute & Virtualization

  • Deploy and manage Proxmox VE hypervisor clusters on bare-metal HPE hardware, including cluster formation, OVS networking, ZFS storage, and VM replication

  • Provision and lifecycle-manage virtual machines using cloud-init, QCOW2 images, and Proxmox API automation

  • Manage physical server provisioning end-to-end via HPE iLO (firmware updates, SPP deployment, OS installation via virtual media)

Container & Kubernetes Platforms

  • Manage self-hosted Kubernetes clusters on-premises, including control plane operations, node provisioning, workload deployment, and upgrade management

  • Operate Docker-based workloads on infrastructure VMs using compose-driven deployments and container health monitoring

  • Maintain container image pipelines and registry infrastructure (Azure Container Registry or AWS ECR)

Cloud Platforms

  • Engineer and maintain infrastructure on AWS and Azure, integrating cloud resources with on-premises systems (DNS, monitoring, identity, networking)

  • Apply cloud cost awareness, security best practices, and IaC principles (IAM, security groups, networking, storage) across AWS and Azure environments

Networking & Core Services

  • Operate and troubleshoot core distributed services including authoritative DNS (BIND9), recursive DNS (Unbound), load balancing (HAProxy), and high-availability VIPs (Keepalived/VRRP)

  • Maintain directory services (OpenLDAP master-replica topology) and AAA infrastructure (FreeRADIUS) used for SSH, VPN, and network device authentication

  • Manage OVS-based network configurations, VLAN topologies, and bonded NIC arrangements across hypervisor fleets

Observability & Security

  • Maintain and extend monitoring infrastructure (Prometheus, Observium) across a global fleet including SNMP polling, metrics collection, and alerting

  • Manage centralised log aggregation pipelines (Fluentbit) and ensure log delivery integrity across DCs

  • Operate runtime security tooling (Falco) and file integrity monitoring (AIDE) in production environments

  • Support PCI-DSS compliance activities including CIS hardening, audit logging (auditd), and participation in control reviews

Reliability & Incident Response

  • Participate in a 24x7 on-call rotation, responding to and leading production incident resolution

  • Conduct blameless post-mortems and drive remediation of root causes through automation and system improvements

  • Define and track SLOs/SLIs for critical infrastructure services

  • Identify and address single points of failure; design and implement HA improvements

What We're Looking For
Required
  • 7+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role in a production environment

  • Strong Linux systems administration skills (RHEL/CentOS preferred) including systemd, networking, storage, kernel tuning, and package management

  • Proficiency with Ansible (or similar tool) for large-scale configuration management, including role design, inventory management, and CI/CD integration

  • Hands-on experience with at least one hypervisor platform, preferably Proxmox VE, Harvester (Kubevirt), or similar (VMware vSphere, KVM)

  • Production experience operating on-premise Kubernetes clusters (rke2, k3s, et etc)

  • Practical AWS or Azure experience including compute, networking (VPC/VNet, security groups, DNS), IAM, and managed services

  • Solid understanding of networking fundamentals: VLANs, bonding/LAG, routing, BGP concepts, DNS, load balancing, and firewall rule management

  • Experience with secrets management platforms (HashiCorp Vault or equivalent)

  • Familiarity with PCI-DSS requirements as they apply to infrastructure – hardening standards (CIS benchmarks), audit logging, access control

  • Experience writing and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent)

  • Demonstrable on-call experience and comfort leading incident response in a global production environment

Preferred
  • Experience managing bare-metal server fleets including out-of-band management tools (HPE iLO, IPMI, or equivalent)

  • Experience with CMDB/IPAM tooling

  • Working knowledge of LDAP directory services and RADIUS authentication

  • Exposure to network monitoring tooling (Observium, LibreNMS, or similar SNMP-based platforms)

  • Scripting proficiency in Python or Bash for tooling and automation tasks

  • Experience operating in colocation / carrier-neutral data center environments (Equinix, Interxion, or similar)

What You'll Need to Succeed
  • Comfort operating autonomously in a globally distributed, remote-first team across multiple time zones

  • A bias toward automation: if you've done something manually twice, you should be automating it

  • Strong written communication skills – you will document systems, write runbooks, and produce post-mortems

  • The ability to context-switch between strategic work (architecture improvements, tooling development) and urgent operational issues (incident response)

  • Willingness to participate in a 24x7 on-call rotation with a fair and well-supported rota

Nice to Have
  • Experience with OSTree/ atomic update workflows for OS lifecycle management

  • Familiarity with DDoS mitigation platforms (Corero or similar)

  • Exposure to BGP route reflector concepts and network-function virtualization

  • Experience with Pulp or other on-premises RPM/package repository management systems

  • Contributions to open-source infrastructure tooling

The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.

Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or @myworkday.com).

Equal Employment Opportunity

It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE, Infrastructure & Platform
Senior SRE, Infrastructure & Platform

F5 Networks, Inc • United States

On-site
USD 150,000 - 190,000
Network Engineer
Network Engineer

F5 • Virginia (MN)

On-site
USD 111,000 - 167,000
Site Reliability Engineer III
Site Reliability Engineer III

Relha LLC • San Jose (CA)

Hybrid
USD 150,000 - 225,000
Annual bonus
Stock options
Benefits
Sr. Software Development Engineer
Sr. Software Development Engineer

F5 Networks, Inc.  • United States

Hybrid
USD 180,000 - 269,000
Sr. Software Development Engineer
Sr. Software Development Engineer

F5 Networks, Inc.  • San Jose (CA)

Hybrid
USD 180,000 - 269,000
Security Engineer III
Security Engineer III

F5 • Reston (VA)

On-site
USD 152,000 - 228,000
Senior Software Engineer - Platform Security
Senior Software Engineer - Platform Security

ffive • Reston (VA)

On-site
USD 180,000 - 269,000
Senior Software Engineer – Platform Security
Senior Software Engineer – Platform Security

F5 Networks, Inc.  • Reston (VA)

Hybrid
USD 180,000 - 269,000
Security Engineer II
Security Engineer II

F5 Networks, Inc • Reston (VA)

On-site
USD 105,000 - 157,000
SecOps Platform Engineer III (DLP/EDR)
SecOps Platform Engineer III (DLP/EDR)

F5 Networks, Inc • Warsaw (IN)

On-site
USD 110,000 - 150,000