Must have hands-on operational experience with:
- AWS (strong): EC2, ECS/Fargate
- AWS (working knowledge): RDS (Postgres/MySQL)
- GCP (strong): GCE (VMs)
- GCP (working knowledge): Cloud SQL, Alloy DB
Expected capabilities:
- Troubleshoot performance, availability, and connectivity issues
- Support upgrades, patching coordination, and lifecycle management
- Work across dev, QA, and production environments with strong awareness of production impact
2. Networking (Practical Knowledge)
Strong working understanding of:
- VPCs, subnets, routing tables
- Security groups and firewall rules
- DNS (internal and external)
- Basic connectivity patterns (peering, private endpoints / PSC)
Ability to:
- Troubleshoot cross-service connectivity issues
- Collaborate effectively with networking teams (clear understanding of ownership boundaries)
3. IAM, Security & Access Governance (Critical)
Required experience with:
- AWS IAM roles and policies
- GCP IAM (projects, folders, service accounts)
- Least privilege access model
- PAM / temporary elevation concepts
- Service account usage and restrictions
4. Infrastructure as Code (IaC)
Required:
- Terraform (hands-on)
- AWS CloudFormation
Experience with:
- AWS CloudWatch
- GCP Monitoring and Logging
- Splunk or similar tools
Ability to:
- Troubleshoot incidents using logs and metrics
- Tune alerts and reduce noise
6. DevOps Tooling Support
Must have experience supporting:
Familiarity with:
- GitHub / repository management
7. Incident Management & Operations
Required experience with:
- On-call support model
- Incident triage, escalation, and ownership
- Own issues end-to-end (triage ? resolution ? handoff if needed)
- Communicate clearly during incidents
- Work across Cloud Engineering, application teams, and vendors
8. Documentation & Runbooks
Strong expectation for:
- Writing and maintaining runbooks
- Keeping documentation accurate and up to date
Ability to:
Experience with:
- Scripting (Python or Bash)
- Automating operational tasks and workflows
Mindset:
- Focus on reducing manual effort and improving operational efficiency
10. Cloud Lifecycle Management
Experience with:
- Provisioning, upgrades, and decommissioning
Understanding of:
- Patching responsibilities (shared vs application-owned)