The Mission
What does it take to modernize critical infrastructure in the telecommunications industry? Join a platform engineering engagement focused on replatforming on-premises infrastructure, building highly available and resilient platforms, and improving how engineering teams develop and deploy applications.
You'll work across Kubernetes, bare-metal infrastructure, virtualization, GitOps, and Infrastructure-as-Code, helping transform complex environments into secure, scalable, and automated platforms.
The Role
- Kubernetes & On-Premises Infrastructure: Install, deploy, maintain, and upgrade Kubernetes clusters across on-premises and bare-metal environments, supporting containerized and virtualized workloads.
- Infrastructure as Code & GitOps: Build and manage declarative infrastructure and deployment workflows using Terraform, Ansible, Helm, Argo CD, or Flux, improving consistency, scalability, and operational efficiency.
- Platform Engineering & Automation: Design and evolve reusable platform capabilities, automate infrastructure provisioning and CI/CD pipelines, and provide tools and environments that improve developer experience.
- Reliability, Security & Operations: Maintain highly available and resilient platforms through monitoring, alerting, incident response, post-incident reviews, and security-first engineering practices, including participation in an on-call rotation.
- Networking & Observability: Support Linux systems, container networking, load balancing, and distributed infrastructure, with exposure to technologies such as BGP, Cilium/eBPF, Istio, Linkerd, Prometheus, Grafana, and OpenTelemetry.
- Technical Leadership & Collaboration: Work closely with development and infrastructure teams to improve platform architecture, troubleshoot complex production issues, mentor junior engineers, and drive operational excellence.
The "Must-Haves"
- Strong hands-on experience with Kubernetes installation, deployment, cluster lifecycle management, and upgrades in on-premises or bare-metal environments.
- Proven expertise in GitOps (Argo CD or Flux) alongside Infrastructure-as-Code tools such as Terraform, Ansible, or Helm.
- Solid Linux systems administration experience supporting production, containerized, and virtualized workloads.
- Experience with platform automation, CI/CD, production troubleshooting, and maintaining secure, highly available infrastructure.
- Based within 2-3 hours of AEST, such as the Philippines, Vietnam, or Malaysia.
The "Bonus" Skills
- Experience with bare-metal infrastructure, Rancher RKE2, KubeVirt, or Kubernetes Gateway APIs.
- Knowledge of distributed systems, high availability, fault tolerance, and advanced networking, including BGP, Cilium/eBPF, and service mesh.
- Experience in telecommunications, ISP, or service-provider environments.
- Programming skills in Go or Python for custom automation and infrastructure tooling.
- Hands-on experience with Prometheus, Grafana, Loki, OpenTelemetry, or Jaeger for monitoring, logging, and tracing.
- Platform Engineer at Heart: Passionate about building scalable, resilient platforms that enable engineering teams to deliver efficiently.
- Hands-On Problem Solver: Enjoys working across Kubernetes, Linux, networking, and infrastructure to solve complex technical challenges.
- Automation-Minded: Focused on reducing manual processes, improving workflows, and embracing declarative infrastructure.
- Reliability-Focused: Prioritizes security, availability, observability, and operational excellence in production environments.
- Strong Communicator: Comfortable collaborating with technical and non-technical stakeholders, providing guidance, and mentoring others.
- Independent & Collaborative: Takes ownership, works autonomously, and thrives in a fast-paced environment while supporting the wider team.