Seniority Level:Senior Level (5+ Years Experience)
MandaapX is a forward-thinking IT consulting company committed to transforming how organizations operate and scale. We blend cloud-native technologies, robust engineering practices, and intelligent systems to deliver real-world business outcomes. We partner with enterprises and fast-growing startups to engineer future-ready technology, focusing deeply on Cloud Computing, Full-Stack Engineering, Salesforce Solutions, and AI/Data Engineering. With a highly specialized global team, we design resilient, secure, and scalable digital ecosystems that help businesses adapt and lead in their respective industries.
Role Overview:
We are seeking aSenior DevOps Engineerwith5+ years of experienceto lead and elevate our cloud infrastructure, continuous delivery pipelines, and reliability practices. In this role, you will play a pivotal role in designing, building, and managing the foundational infrastructure across our production-grade Google Cloud Platform (GCP) ecosystem.
As a technical authority and thought leader, you will collaborate with cross-functional software engineering teams to design cloud-native architectures that meet demanding performance, scalability, and security requirements. You will set engineering standards, establish SRE paradigms, mentor team members, and build frictionless, self-healing platforms while supporting hybrid or multi-cloud environments.
Key Responsibilities:
GCP Infrastructure, Operations & FinOps
- Design, deploy, and manage scalable, highly available GCP infrastructure across Compute Engine, GKE, Cloud Run, Cloud SQL, BigQuery, Pub/Sub, Cloud Armor, and Cloud Storage.
- Implement and maintain complex VPC architectures, firewall rules, Cloud NAT, Shared VPC, and Private Google Access configurations.
- Lead FinOps initiatives and cost optimization strategies using committed use discounts, rightsizing, budget alerts, resource labeling, and capacity planning.
Infrastructure as Code (IaC) & Automation
- Build and maintain all infrastructure using Terraform, enforcing strict modular, reusable, and DRY code conventions.
- Manage Terraform state in GCS backends with proper state locking, workspace management, and automated policy enforcement.
- Automate environment provisioning and CI/CD workflows using GitHub Actions with systematic plan, validate, and apply stages.
- Design and manage Google Kubernetes Engine (GKE) clusters, including node pools, autoscaling, and workload configurations.
- Build, optimize, and maintain Docker images (multi-stage builds) and container registries (Artifact Registry / GCR).
- Implement Kubernetes best practices including namespaces, RBAC, network policies, resource quotas, pod security standards, and GitOps deployments via Helm charts.
SRE, Observability, & Incident Management
- Monitor infrastructure health using Cloud Monitoring, Cloud Logging, and proactive alerting policies, integrating cloud-native monitoring with Datadog for end-to-end observability (APM, metrics, tracing).
- Drive SRE principles by establishing SLOs, SLIs, and Error Budgets across core services.
- Participate in on-call rotations, lead high-priority incident response efforts, conduct Blameless Post-Mortems, and implement self-healing mechanisms.
Security, Identity, & Compliance
- Enforce GCP security best practices including VPC Service Controls, Security Command Center, Binary Authorization, and GCP Organization Policies.
- Manage Identity and Access Management (IAM), service accounts, Workload Identity Federation, and least-privilege policies.
- Manage secrets via Secret Manager, enforce encryption at rest and in transit, and align systems with FERPA, NIST, SOC 2, and CIS benchmarks.
- Partner with software engineers to design cloud-native architectures that meet performance and scalability requirements
- Produce and maintain clear infrastructure documentation, architecture diagrams, and runbooks.
- Participate in on-call rotations and respond to infrastructure incidents
- Mentor team members on GCP services and cloud engineering best practices
Requirements & Qualifications
This role requires a hands-on approach to problem-solving. Candidates must demonstrate clear grasp of the following fundamental concepts, supported by practical, real-world examples of their prior work.
- 5+ years of hands-on experience with Google Cloud Platform in a production environment.
- Deep knowledge of core GCP services: Compute Engine, GKE, Cloud Run, Cloud SQL, Cloud Storage, BigQuery, Pub/Sub, Cloud Armor, and VPC networking.
- Strong proficiency with Terraform for infrastructure as code; experience managing modules and remote state.
- Solid experience with Kubernetes — cluster management, workloads, RBAC,
- networking, and troubleshooting.
- Proficiency with Docker — writing Dockerfiles, multi-stage builds, image optimization, and registry management.
- Experience building and maintaining CI/CD pipelines with GitHub Actions.
- Strong understanding of cloud networking concepts: BGP, DNS, load balancing, CDN, and hybrid connectivity.
- Proven ability to implement GCP security best practices including IAM, VPC Service Controls, and encryption.
- Scripting proficiency in Python, Bash, or Go for automation tasks
- Excellent written and verbal communication skills
Preferred Qualifications
- Google Cloud Professional certifications (Cloud Architect, DevOps Engineer, or Security Engineer)
- Experience with AWS — ability to manage multi-cloud environments or support cloud migrations
- Hands-on experience with Datadog for infrastructure monitoring, APM, and alerting
- Experience with database services including Cloud Spanner, Firestore, or Bigtable
- Knowledge of FinOps practices and GCP cost management tools
- Familiarity with Agile/Scrum methodologies.
Our Ideal Candidate
We are looking for an individual who can showcase real-world experience that confirms their ability to apply these skills in a production environment. During the interview process, be prepared to discuss specific examples of problems you have solved and projects you have