We are building a product that runs across on-premises environments, GCP, and AWS. We need an infrastructure engineer who can own the deployment, reliability, and scaling of this product across all three environments. This is a hands‑on role, not a planning‑and‑PowerPoint role. You will build and maintain the infrastructure that keeps the product running in production.
Responsibilities
- Design, deploy, and maintain infrastructure across on-prem, GCP, and AWS.
- Manage networking: VPNs, VPCs, peering, DNS, load balancers, and firewalls.
- Set up and manage Kubernetes clusters (GKE on GCP, EKS on AWS, self‑managed on‑prem).
- Set up monitoring, alerting, and logging (Prometheus, Grafana, ELK).
- Build and maintain CI/CD pipelines for multi‑environment deployments.
- Handle secrets management, IAM policies, and security hardening.
- Write and maintain Infrastructure as Code (Terraform, Pulumi, CloudFormation).
- Troubleshoot production incidents end‑to‑end, from the app layer to bare metal.
- Automate repetitive operational tasks; reduce toil systematically.
- Work with dev teams to make applications deployment‑ready and observable.
Requirements
- Bias toward over‑discussing. Ships working infrastructure, iterates, and improves.
- Debugs methodically. Traces problems from user symptoms through app, network, and infra.
- Thinks about failure modes proactively. Designs for recoverability, not just uptime.
- Communicates clearly with engineers who are not infra specialists.
- Owns infrastructure decisions independently, without waiting to be told what to do.
- Has been on‑call and knows what it means to keep systems running at 2 AM.
Must Have
- 3+ years of hands‑on infrastructure / DevOps experience.
- Kubernetes: Deploying, operating, and debugging clusters; GKE or EKS experience.
- GCP: GKE, Compute Engine, VPC, IAM, Cloud SQL, Cloud Storage, Cloud Monitoring.
- Linux: Administration, networking, systemd, storage, OS‑level troubleshooting.
- IaC: Proficiency with Terraform or equivalent.
- Networking: TCP/IP, DNS, HTTP, TLS, subnets, routing, firewall rules.
- CI/CD: Cloud Build, GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
- Scripting: Bash + at least one of Python or Go.
Good to Have
- AWS: EC2, VPC, IAM, S3, RDS, EKS, Route 53, ALB/NLB, CloudWatch.
- On‑prem: Bare metal, VMware/Proxmox, NFS/Ceph, on‑prem networking.
- Hybrid networking: Site‑to‑site VPN, Cloud Interconnect, Direct Connect.
- GitOps: Helm, Kustomize, ArgoCD.
- Service mesh: Istio, Linkerd, or Consul Connect.
- Container security: Falco, OPA/Gatekeeper, Trivy, Snyk.
- DB ops: Backup, replication, failover (PostgreSQL, MySQL).
- Cost optimization: Right‑sizing, preemptible/spot instances.
- DR/HA: Multi‑region or hybrid disaster recovery.
- Compliance: CIS benchmarks, SOC 2, or equivalent.
Bonus Points
- You have deployed the same product across both cloud and on‑prem and dealt with the real pain of environment parity.
- You have operated Kubernetes in production (not just set it up once in a sandbox).