Principal Site Reliability Engineer, Infrastructure & Platform

F5 NETWORKS SINGAPORE PTE LTD

Singapore

On-site

SGD 120,000 - 180,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

F5 NETWORKS SINGAPORE PTE LTD is seeking a seasoned DevOps/SRE professional to design, build, and maintain large-scale infrastructure across multi-datacenter environments. You will author and refactor Ansible playbooks, enhance GitLab CI pipelines, and manage secrets with Vault while ensuring CMDB accuracy in NetBox.

Responsibilities include deploying Proxmox hypervisor clusters, provisioning VMs with cloud-init, managing on-prem Kubernetes, and operating core services like DNS, NTP, and logging

Qualifications

  • Strong Linux systems administration skills (RHEL/CentOS preferred) including systemd, networking, storage, kernel tuning, and package management
  • Proficiency with Ansible (or similar tool) for large-scale configuration management, including role design, inventory management, and CI/CD integration
  • Hands-on experience with at least one hypervisor platform, preferably ProxmoxVE, Harvester (Kubevirt) or similar (VMware vSphere, KVM)
  • Production experience operating on-premise Kubernetes clusters (rke2, k3s, etc)
  • Practical AWS or Azure experience including compute, networking (VPC/VNet, security groups, DNS), IAM, and managed services
  • Solid understanding of networking fundamentals: VLANs, bonding/LAG, routing, BGP concepts, DNS, load balancing, and firewall rule management
  • Experience with secrets management platforms (HashiCorp Vault or equivalent)
  • Familiarity with PCI-DSS requirements as they apply to infrastructure — hardening standards (CIS benchmarks), audit logging, access control
  • Experience writing and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent)
  • Demonstrable on-call experience and comfort leading incident response in a global production environment
  • 7+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role in a production environment

Responsibilities

  • Author, maintain, and refactor Ansible playbooks and roles across a large-scale multi-datacenter inventory, covering bare-metal provisioning to application deployment
  • Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions
  • Manage secrets lifecycle using HashiCorp Vault, including AppRole authentication, secret rotation, and PKI integration
  • Maintain CMDB/IPAM accuracy in NetBox as a source of truth for all infrastructure assets
  • Deploy and manage ProxmoxVE hypervisor clusters on bare-metal hardware, including cluster formation, OVS networking, ZFS storage, and VM replication
  • Provision and lifecycle-manage virtual machines using cloud-init, QCOW2 images, and Proxmox API automation
  • Manage physical server provisioning end-to-end via HPE iLO (firmware updates, SPP deployment, OS installation via virtual media)
  • Manage self-hosted Kubernetes clusters on-premises, including control plane operations, node provisioning, workload deployment, and upgrade management
  • Operate Docker-based workloads on infrastructure VMs using compose-driven deployments and container health monitoring
  • Maintain container image pipelines and registry infrastructure (Azure Container Registry or AWS ECR)
  • Engineer and maintain infrastructure on AWS and Azure, integrating cloud resources with on-premises systems (DNS, monitoring, identity, networking)
  • Apply cloud cost awareness, security best practices, and IaC principles (IAM, security groups, networking, storage) across AWS and Azure environments
  • Operate and troubleshoot core distributed services including authoritative DNS (BIND9), recursive DNS (Unbound), load balancing (HAProxy), and high-availability VIPs (Keepalived/VRRP)
  • Maintain directory services (OpenLDAP master-replica topology) and AAA infrastructure (FreeRADIUS) used for SSH, VPN, and network device authentication
  • Manage OVS-based network configurations, VLAN topologies, and bonded NIC arrangements across hypervisor fleets
  • Maintain and extend monitoring infrastructure (Prometheus, Observium) across a global fleet including SNMP polling, metrics collection, and alerting
  • Manage centralized log aggregation pipelines (Fluentbit) and ensure log delivery integrity across DCs
  • Operate runtime security tooling (Falco) and file integrity monitoring (AIDE) in production environments
  • Support PCI-DSS compliance activities including CIS hardening, audit logging (auditd), and participation in control reviews
  • Identify and address single points of failure; design and implement HA improvements
  • Participate in a 24x7 on-call rotation, responding to and leading production incident resolution
  • Conduct blameless post-mortems and drive remediation of root causes through automation and system improvements
  • Define and track SLOs/SLIs for critical infrastructure services

Skills

Linux administration
CI/CD experience
Ansible
On-call experience
Networking fundamentals

Tools

ProxmoxVE
Harvester/Kubevirt
VMware vSphere
GitLab CI
HashiCorp Vault
NetBox
Prometheus
Observium
HAProxy
BIND9
Unbound
Keepalived
OpenLDAP
FreeRADIUS

Job description

Responsibilities


  • Author,maintain, and refactor Ansible playbooks and roles across a large-scale multi-datacenter inventory, covering the full lifecycle from bare-metal provisioning to application deployment

  • Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions

  • Manage secrets lifecycle usingHashiCorpVault, includingAppRoleauthentication, secret rotation, and PKI integration

  • Maintain CMDB/IPAM accuracy inNetBoxas a source of truth for all infrastructure assets


Compute & Virtualization


  • Deploy and manageProxmoxVE hypervisor clusters on bare-metal HPE hardware, including cluster formation, OVS networking, ZFS storage, and VM replication

  • Provision and lifecycle-manage virtual machines using cloud-init, QCOW2 images, andProxmoxAPI automation

  • Manage physical server provisioning end-to-end via HPEiLO(firmware updates, SPP deployment, OS installation via virtual media)


Container & Kubernetes Platforms


  • Manage self-hosted Kubernetes clusters on-premises, including control plane operations, node provisioning, workload deployment, and upgrade management

  • OperateDocker-based workloads on infrastructure VMs using compose-driven deployments and container health monitoring

  • Maintaincontainer image pipelines and registry infrastructure (Azure Container Registryor AWS ECR)


Cloud Platforms


  • Engineer andmaintaininfrastructure on AWS and Azure, integrating cloud resources with on-premises systems (DNS, monitoring, identity, networking)

  • Apply cloud cost awareness, security best practices, andIaCprinciples (IAM, security groups, networking, storage) across AWS and Azure environments


Networking & Core Services


  • Operateand troubleshoot core distributed services including authoritative DNS (BIND9), recursive DNS (Unbound), load balancing (HAProxy), and high-availability VIPs (Keepalived/VRRP)

  • Maintaindirectory services (OpenLDAPmaster-replica topology) and AAA infrastructure (FreeRADIUS) used for SSH, VPN, and network device authentication

  • Manage OVS-based network configurations, VLAN topologies, and bonded NIC arrangements across hypervisor fleets


Observability & Security


  • Maintainand extend monitoring infrastructure (Prometheus,Observium) across a global fleet including SNMP polling, metrics collection, and alerting

  • Managecentralisedlog aggregation pipelines (Fluentbit) and ensure log delivery integrity across DCs

  • Operateruntime security tooling (Falco) and file integrity monitoring (AIDE) in production environments

  • Support PCI-DSS compliance activities including CIS hardening, audit logging (auditd), and participation in control reviews


Reliability & Incident Response


  • Participatein a 24x7 on-call rotation, responding to and leading production incident resolution

  • Conduct blameless post-mortems and drive remediation of root causes through automation and system improvements

  • Define and track SLOs/SLIs for critical infrastructure services

  • Identifyand address single points of failure; design and implement HA improvements


Requirements


  • Strong Linux systems administration skills (RHEL/CentOS preferred) includingsystemd, networking, storage, kernel tuning, and package management




  • Proficiencywith Ansible(or similar tool)for large-scale configuration management, including role design, inventory management, and CI/CD integration




  • Hands-on experience with at least one hypervisor platform, preferablyProxmoxVE, Harvester (Kubevirt)or similar (VMware vSphere, KVM)




  • Production experienceoperatingon-premiseKubernetes clusters (rke2, k3s,etc)




  • Practical AWSorAzure experience includingcompute, networking (VPC/VNet, security groups, DNS), IAM, and managed services




  • Solid understanding of networking fundamentals: VLANs, bonding/LAG, routing, BGP concepts, DNS, load balancing, andfirewallrule management




  • Experience with secrets management platforms (HashiCorpVault or equivalent)




  • Familiarity with PCI-DSS requirements as they apply to infrastructure -- hardening standards (CIS benchmarks), audit logging, access control




  • Experience writing andmaintainingCI/CD pipelines (GitLab CI, GitHub Actions, or equivalent)




  • Demonstrable on-call experience and comfort leading incident response in a global production environment




  • 7+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role in a production environment

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Principal Site Reliability Engineer, Infrastructure & Platform
Senior Principal Site Reliability Engineer, Infrastructure & Platform

f5 networks singapore pte ltd • Singapore

On-site
SGD 180,000 - 240,000
IT Infrastructure Engineer
IT Infrastructure Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 90,000 - 140,000
Senior Cloud Engineer
Senior Cloud Engineer

BOND FINANCIAL GROUP PTE. LTD. • Shenton Way

On-site
SGD 120,000 - 180,000
Cloud Operations Engineer – Infrastructure
Cloud Operations Engineer – Infrastructure

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 110,000 - 170,000
Director, Systems Engineering – Compute Operations, Linux
Director, Systems Engineering – Compute Operations, Linux

Jobtailor • Singapore

On-site
SGD 240,000 - 360,000
Senior Devops Engineer
Senior Devops Engineer

THALES SOLUTIONS ASIA PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
CLOUD ENGINEER
CLOUD ENGINEER

RFNET TECHNOLOGIES PTE LTD • Singapore

On-site
SGD 120,000 - 170,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

PhillipCapital • Singapore

On-site
SGD 180,000 - 240,000
DevSecOps Engineer
DevSecOps Engineer

YM GLOBAL TECHNOLOGIES SDN. BHD. • Singapore

On-site
SGD 90,000 - 130,000
Platform Engineer (RHEL & OpenShift)
Platform Engineer (RHEL & OpenShift)

ZENITH INFOTECH (S) PTE LTD. • Singapore

On-site
SGD 120,000 - 180,000