Senior Principal Site Reliability Engineer, Infrastructure & Platform

f5 networks singapore pte ltd

Singapore

On-site

SGD 180,000 - 240,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

f5 networks singapore pte ltd is seeking a Senior Site Reliability Engineer to lead automation across a global, multi-datacenter infra. You will build internal tools, Ansible playbooks, and CI/CD pipelines to reduce toil and enable self-service at scale.

The role emphasizes Python/Go-based tooling, Kubernetes operations, and PCI-DSS compliant environments, with on-call rotation and deep problem-solving responsibilities.

Qualifications

  • 5+ years in an SRE/DevOps role with strong automation focus.
  • Proficiency in Python; building CLI tools, APIs, and automation frameworks.
  • Expert-level Ansible skills: large-scale inventory, modules, CI/CD integration.
  • Solid Linux systems knowledge (RHEL/CentOS).
  • Experience building and maintaining GitLab CI pipelines for infrastructure automation.
  • Production experience with self-hosted Kubernetes: cluster ops and automation.
  • AWS and Azure experience with an IaC mindset.
  • API-driven infra management experience (REST, Vault, iLO, APIs).
  • HashiCorp Vault or equivalent secrets management.
  • Understanding PCI-DSS in automated infrastructure.

Responsibilities

  • Design and develop internal tools, CLIs, and APIs to enable infrastructure self-service and automate complex workflows.
  • Build integrations between infrastructure systems (CMDB/IPAM, Vault, Proxmox) into cohesive automations.
  • Develop API clients and libraries for infrastructure services (Proxmox API, Vault API, NetBox API).
  • Write well-tested, documented code with proper versioning and release processes.

Skills

Python
Ansible
Linux (RHEL/CentOS)
CI/CD pipelines
Kubernetes
Cloud platforms (AWS/Azure)
API automation
HashiCorp Vault
PCI-DSS compliance
Software engineering fundamentals

Tools

GitLab CI
Proxmox
NetBox

Job description

We are looking for a Senior Site Reliability Engineerthat leads withkindness, andpossessesa strong software development background to join our Infrastructure Engineering team. Your primary focus will be building automation, tooling, and internal platforms that enable our team tooperatea global, multi-datacenter infrastructure spanninga growing number ofPoints of Presenceacross the globe. Deep familiarity with production infrastructure -- bare-metal hypervisors, containerized workloads, Kubernetes clusters, and cloud platforms -- is essential, but your primary lens shouldbe onautomation and code.You will develop internal tools and APIs in Python and Go,designandmaintainAnsible automation across hundreds of hosts, build CI/CD pipelines, and create self-service interfaces that reduce toil andeliminatemanual operations. You will work within a PCI-DSS compliant environment andparticipatein a 24x7 on-call rotation.

Responsibilities:

Internal Tooling & Application Development
  • Design and develop internal tools, CLIs, and APIs (primarily in Go and Python) that enable infrastructure self-service, automate complex workflows, and improve operational efficiency

  • Build integrations between infrastructure systems -- connecting CMDB/IPAM (NetBox), secrets management (HashiCorpVault), hypervisor APIs (Proxmox), monitoring platforms, and CI/CD pipelines into cohesive automated workflows

  • Develop andmaintainAPI clients and libraries for interacting with infrastructure services (ProxmoxAPI, Vault API,NetBoxAPI,iLORedfish, container registries)

  • Write well-tested, documented, and maintainable code with proper versioning, release processes, and code review practices

Infrastructure as Code & Ansible Development
  • Architect, develop, and refactor Ansible roles and playbooks across a large-scale inventory spanning 30+ datacenters, 80+ group variable files, and 40+ roles

  • Design reusable, composable Ansible role patterns that scale cleanly as the DC footprint grows -- new DCs should be deployable with minimal variable additions

  • Improve idempotency, error handling, and test coverage across the existing Ansible codebase

  • Develop custom Ansible modules, plugins, and lookup plugins where upstream modulesmay be insufficient (e.g., custom Vault integration,ProxmoxAPI interactions,iLOautomation)

  • Automate bare-metal server lifecycle end-to-end: fromiLObootstrap through OS installation, hypervisor configuration, VM provisioning, and service deployment

CI/CD Pipeline Engineering
  • Design, write, and maintain GitLab CI pipelines for infrastructure automation, including multi-stage deployment workflows with linting, validation, canary testing, and regional rollout

  • Build pipeline patterns for safe infrastructure changes: staged rollouts, automated rollback, drift detection, and change validation

  • Create reusable pipeline templates and shared CI components thatstandardisehow infrastructure changes are tested and deployed

  • Implement automated testing for Ansible roles and infrastructure changes(molecule, ansible-lint, integration testing in epic environments)

Kubernetes & Container Platform Automation
  • Develop automation for self-hosted Kubernetes cluster lifecycle management: provisioning, upgrades, scaling, and disaster recovery

  • Build andmaintaincontainer image build pipelines, registry management, and image promotion workflows

  • Create Kubernetes operators or controllers (in Go) where custom automation of cluster-level concerns is needed

  • Automate workload deployment patterns, including Helm chart development andGitOpsworkflows

Cloud Infrastructure Automation
  • DevelopIaCand automation for AWS and Azure resources, integrating cloud infrastructure with on-premises systems

  • Build automation that spans hybrid environments -- coordinating deployments across bare-metal,virtualized, and cloud targets from a unified pipeline

Observability & Reliability Engineering
  • Instrument internal tools and automation with proper logging, metrics, and tracing

  • Build automated remediation workflows that respond to monitoring alerts and reduce mean time to recovery

  • Develop reporting and dashboards that provide visibility into infrastructure state, automation success rates, and toil metrics

  • Identifyand automate away recurring operational toil; track and quantify toil reduction over time

Security & Compliance Automation
  • Automate PCI-DSS compliance workflows including CIS benchmark hardening, audit evidence collection, and configuration drift detection

  • Build automated secret rotation pipelines usingHashiCorpVault

  • Develop security scanning integration into CI/CD pipelines (container image scanning, infrastructure configuration validation)

What We're Looking For
  • 5+ years of experience in an SRE, DevOps, or Infrastructure Engineering role with a strong emphasis on writing code and building automation

  • Proficiencyin Python, with experience building CLI tools, APIs (Flask/FastAPIor equivalent), and automation frameworks

  • Expert-level Ansible skills: custom role development, module/plugin authorship, complex Jinja2 templating, inventory management at scale, and CI/CD integration

  • Solid Linux systems knowledge (RHEL/CentOS) -- you need to understand the systemsyou'reautomating at a depth that lets you debug failures and design robust automation

  • Experience building andmaintainingCI/CD pipelines (GitLab CI preferred) for infrastructure automation, not just application builds

  • Production experience with self-hosted Kubernetes: cluster operations, controller/operator development, and workload automation

  • Practical AWS and Azure experience with anIaCmindset -- provisioning and managing cloud resources through automation, not console clicks

  • Experience with API-driven infrastructure management (RESTful APIs, Redfish/iLO, hypervisor APIs)

  • Familiarity withHashiCorpVault or equivalent secrets management platforms, including programmatic integration

  • Understanding of PCI-DSS requirements as they apply to automated infrastructure management -- audit trails, change control, hardening automation

  • Strong software engineering fundamentals: version control workflows, code review, testing practices, documentation, and release management

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer, Infrastructure & Platform
Principal Site Reliability Engineer, Infrastructure & Platform

F5 NETWORKS SINGAPORE PTE LTD • Singapore

On-site
SGD 120,000 - 180,000
Platform Infrastructure Engineer
Platform Infrastructure Engineer

Cognizant • Singapore

On-site
SGD 120,000 - 180,000
Remote work option
Senior DevOps Engineer
Senior DevOps Engineer

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 80,000 - 120,000
Technical leadership opportunities
Access to mentorship and certifications
Modern DevOps environment
Automation and DevSecOps Engineer
Automation and DevSecOps Engineer

NETS • Singapore

On-site
SGD 120,000 - 180,000
Senior Devops Engineer
Senior Devops Engineer

THALES SOLUTIONS ASIA PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

PhillipCapital • Singapore

On-site
SGD 180,000 - 240,000
Technical Lead – DevOps / SRE
Technical Lead – DevOps / SRE

CORNERSTONE GLOBAL PARTNERS PTE. LTD. • Singapore

On-site
SGD 180,000 - 260,000
DevSecOps Engineer
DevSecOps Engineer

YM GLOBAL TECHNOLOGIES SDN. BHD. • Singapore

On-site
SGD 90,000 - 130,000
DevOps Lead (Platform & Infrastructure Engineering)
DevOps Lead (Platform & Infrastructure Engineering)

CDG ZIG PTE. LTD. • Singapore

On-site
SGD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000