DevOps Manager

GCS Recruitment

Redwood City (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GCS Recruitment is seeking an experienced DevOps Manager to lead a high-performing Technical Operations team responsible for reliable, scalable cloud infrastructure. This leadership role combines technical depth with people management to drive operational excellence across engineering, security, and product teams.

You will establish SLOs/SLIs, oversee Kubernetes platforms, and drive CI/CD adoption while ensuring security and regulatory compliance on a large multi-cloud environment in the United

Qualifications

  • 8+ years in DevOps, Technical Operations, Infrastructure Engineering, or SRE.
  • 4+ years managing technical operations or SRE teams.
  • Experience with production cloud infrastructure in AWS and/or GCP.
  • Hands-on with Kubernetes, Docker, and container orchestration.
  • Strong IaC experience with Terraform.
  • Experience building CI/CD pipelines and deployment automation.
  • Strong Linux systems administration knowledge.
  • Incident management and on-call production support experience.
  • Networking fundamentals and cloud security knowledge.
  • Agile methodologies and Jira or similar tools.
  • Excellent stakeholder management and executive presentation skills.

Responsibilities

  • Lead, mentor, and develop a Technical Operations/DevOps team.
  • Ensure 24/7 operational stability of production infrastructure.
  • Drive continuous improvement and operational excellence.
  • Lead incident response and post-incident reviews (RCA).
  • Establish and manage SLOs, SLIs, and error budgets.
  • Oversee IaC-based infrastructure automation.
  • Manage Kubernetes-based platforms and cloud infrastructure.
  • Improve deployment processes through CI/CD and automation.
  • Collaborate with Product, Engineering, Security, and Infrastructure.
  • Drive disaster recovery, resiliency, and business continuity planning.
  • Develop monitoring, alerting, and observability strategies.
  • Present operational metrics and project updates to senior leadership.
  • Ensure security standards, audits, and regulatory compliance.

Skills

Leadership
People management
AWS
GCP
Kubernetes
Terraform
CI/CD
Linux administration
Incident management
SRE
Agile/Jira
Networking basics
Executive communication

Education

Bachelor's degree in CS/Engineering

Tools

Kubernetes
Docker
Terraform
CI/CD pipelines
Linux

Job description

About the Role

We are seeking an experienced DevOps Manager to lead a high-performing Technical Operations team responsible for ensuring the reliability, scalability, and performance of a large-scale cloud infrastructure. This leadership role combines technical expertise with people management, driving operational excellence, infrastructure automation, and cross-functional collaboration across engineering, security, and product teams.

Key Responsibilities
  • Lead, mentor, and develop a high-performing Technical Operations/DevOps team.
  • Ensure 24/7 operational stability, availability, and reliability of production infrastructure.
  • Drive operational excellence through continuous improvement initiatives.
  • Lead incident response, major incident management, and post-incident reviews (RCA).
  • Establish and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
  • Oversee infrastructure automation using Infrastructure as Code (IaC) principles.
  • Manage Kubernetes-based container platforms and cloud infrastructure.
  • Improve deployment processes through CI/CD and automation.
  • Collaborate with Product, Engineering, Security, and Infrastructure teams on strategic initiatives.
  • Drive disaster recovery, resiliency, and business continuity planning.
  • Develop monitoring, alerting, and observability strategies.
  • Present operational metrics and project updates to senior leadership.
  • Ensure compliance with security standards, audits, and regulatory requirements.
Required Qualifications
  • 8+ years of experience in DevOps, Technical Operations, Infrastructure Engineering, or Site Reliability Engineering.
  • 4+ years of experience managing technical operations, infrastructure, or SRE teams.
  • Strong leadership, mentoring, and people management skills.
  • Experience managing production cloud infrastructure in AWS and/or GCP.
  • Hands-on experience with Kubernetes, Docker, and container orchestration.
  • Strong experience with Terraform and Infrastructure as Code (IaC).
  • Experience implementing CI/CD pipelines and deployment automation.
  • Strong understanding of Linux systems administration.
  • Experience with incident management, on-call operations, and production support.
  • Knowledge of networking fundamentals, cloud security, and infrastructure best practices.
  • Experience with Agile methodologies, Kanban, Jira, or similar project management tools.
  • Excellent communication, stakeholder management, and executive presentation skills.
Preferred Qualifications
  • Experience implementing Site Reliability Engineering (SRE) practices.
  • Experience with monitoring and observability platforms such as Prometheus, Grafana, APM tools, and centralized logging solutions.
  • Experience with disaster recovery planning and multi-region infrastructure.
  • Knowledge of compliance frameworks including SOC 2 and other security standards.
  • Experience supporting enterprise-scale cloud platforms.
  • Cloud certifications such as AWS Solutions Architect or Google Cloud Professional certifications.
  • Experience with enterprise networking, telecommunications, wireless networking, or IoT environments.
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
Technical Environment
  • Google Cloud Platform (GCP)
  • Amazon Web Services (AWS)
  • Kubernetes
  • Docker
  • Terraform
  • Linux
  • CI/CD Pipelines
  • Git
  • Prometheus
  • Grafana
  • Application Performance Monitoring (APM)
  • Infrastructure as Code (IaC)
  • Jira
  • Agile / Kanban
  • Networking & Cloud Security
What You'll Gain
  • Opportunity to lead and mentor a high-performing DevOps and Technical Operations team.
  • Work on large-scale, cloud-native infrastructure supporting enterprise applications and mission-critical services.
  • Drive strategic initiatives in infrastructure automation, cloud modernization, and Site Reliability Engineering (SRE).
  • Gain hands-on experience with cutting-edge technologies including Kubernetes, Terraform, AWS/GCP, CI/CD, and Infrastructure as Code (IaC).
  • Influence technical direction and collaborate with cross-functional engineering, security, and product teams.
Get your free, confidential resume review.
or drag and drop your file here.