Principal Cloud Operations Engineer

Jobtailor

California (MO)

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JobTailor seeks an seasoned cloud and infrastructure engineer to design, implement, and maintain scalable cloud/hybrid infrastructure for production workloads. You will lead IaC adoption with Terraform/CloudFormation and architect multi-cloud solutions spanning OCI, AWS, Azure, and on-prem data centers.

You will drive observability, SRE practices, and automated deployments while mentoring engineers and partnering with security/compliance teams to ensure production readiness and cost efficiency.

Qualifications

  • 10+ years in cloud and infrastructure engineering with 3+ years in senior roles
  • Experience with OCI (preferred), AWS and/or Azure
  • Proven ability to manage production-scale environments for mission-critical apps and services
  • Strong proficiency with IaC, CI/CD toolchains, and container orchestration
  • Knowledge of security, compliance, and networking in hybrid environments
  • Experience leading cross-functional initiatives from concept to execution

Responsibilities

  • Design, implement, and maintain cloud/hybrid infrastructure for production workloads.
  • Lead IaC adoption using Terraform, CloudFormation, or similar tools for repeatable deployments.
  • Architect scalable, fault-tolerant solutions across OCI, AWS, Azure, and on-prem data centers.
  • Evaluate emerging cloud services for business needs and scalability.
  • Serve as technical lead for production operations ensuring uptime and reliability.
  • Develop observability frameworks with metrics, logs, and traces.
  • Partner with engineering to implement SRE practices including SLOs and post-incident reviews.
  • Drive root cause analysis and performance tuning of production services
  • Collaborate with DevOps to build automated deployment pipelines for frequent releases
  • Integrate security/compliance checks into CI/CD workflows
  • Design self-healing infrastructure with automated rollback mechanisms
  • Ensure secure configuration management and environment orchestration with Ansible, Chef or Puppet
  • Establish operational best practices for monitoring, patching, and change management
  • Lead production readiness reviews for new releases and major changes
  • Collaborate with Security/Compliance to meet policy and hardening standards
  • Participate in on-call rotations for critical systems
  • Mentor cloud and infrastructure engineers
  • Lead architectural reviews and capacity planning discussions
  • Advise on cloud modernization, resilience engineering, and cost optimization

Skills

Analytical skills
Problem-solving
Mentoring
Collaboration
Leadership
Incident management
Capacity planning
Cross-functional leadership

Education

Bachelor's degree in Computer Science or Information Systems
Master's degree preferred

Tools

Terraform
CloudFormation
Jenkins
GitLab
ArgoCD
Kubernetes
Docker
Prometheus
Grafana
Datadog
ELK
Python
Bash
PowerShell

Job description

Responsibilities
  • Design, implement, and maintain cloud and hybrid infrastructure supporting production workloads, enterprise systems, and CI/CD pipelines
  • Lead the adoption of infrastructure-as-code (IaC) using Terraform, CloudFormation, or similar tools to enable repeatable, auditable, and secure deployments
  • Architect scalable and fault-tolerant solutions across OCI, AWS, Azure, and on-prem data centers, ensuring high availability and cost efficiency
  • Evaluate emerging cloud services and technologies for applicability to business needs and long-term scalability goals
  • Serve as the technical lead for production operations, ensuring uptime, performance, and reliability of customer-facing and internal systems
  • Develop and maintain observability frameworks leveraging metrics, logs, and traces to ensure proactive detection and rapid response
  • Partner with engineering teams to implement SRE-inspired practices, including service level objectives (SLOs), error budgets, and post-incident reviews
  • Drive root cause analysis, performance tuning, and continuous improvement of production services
  • Collaborate with DevOps and application engineering teams to build and optimize automated deployment pipelines supporting frequent, low-risk releases
  • Integrate security and compliance checks into CI/CD workflows to ensure production readiness and alignment with internal standards
  • Design self-healing infrastructure and automated rollback mechanisms to reduce operational risk
  • Ensure secure and reliable configuration management and environment orchestration using tools such as Ansible, Chef, or Puppet
  • Establish and enforce operational best practices for monitoring, patching, and change management across production systems
  • Lead production readiness reviews for new releases and large-scale changes
  • Collaborate with the Security and Compliance teams to ensure systems adhere to policy, hardening standards, and regulatory requirements
  • Participate in and occasionally lead on-call rotations for critical production systems, ensuring rapid triage and resolution
  • Act as a technical mentor to cloud and infrastructure engineers, fostering a culture of knowledge sharing and engineering excellence
  • Lead architectural reviews, design sessions, and capacity planning discussions
  • Serve as a trusted advisor to management on cloud modernization, resilience engineering, and cost optimization strategies.
Requirements
  • Bachelor’s degree in Computer Science, Information Systems, or related field; Master’s preferred
  • 10+ years of experience in cloud and infrastructure engineering, including 3+ years in a senior or principal role
  • Expertise with OCI (preferred), AWS and/or Azure cloud services, including networking, compute, storage, and identity management
  • Proven experience managing production-scale environments supporting mission-critical applications and services
  • Strong proficiency in:
  • -Infrastructure-as-code (Terraform, CloudFormation)
  • -CI/CD and DevOps toolchains (Jenkins, GitLab, ArgoCD)
  • -Container orchestration (Kubernetes, Docker)
  • -Monitoring and observability platforms (Prometheus, Grafana, Datadog, ELK)
  • -Scripting and automation (Python, Bash, PowerShell)
  • Solid understanding of security, compliance, and networking principles in hybrid environments
  • Exceptional analytical, problem-solving, and incident management skills
  • Demonstrated ability to lead complex, cross-functional initiatives from concept to execution.
ATS Optimization Keywords

Below are skills and terms extracted directly from this job posting to improve Applicant Tracking System (ATS) visibility. This unique feature helps candidates tailor their applications more effectively — a feature exclusive to JobTailor job listings.

Hard Skills
  • Cloud Services (OCI, AWS, Azure)
  • Scripting and Automation (Python, Bash, PowerShell)
  • Security and Compliance Principles
  • Production-Scale Environment Management
  • Performance Tuning
  • Root Cause Analysis
  • Configuration Management (Ansible, Chef, Puppet)
  • Service Level Objectives (SLOs)
  • Error Budgets
  • Incident Management
Soft Skills
  • Analytical Skills
  • Problem-Solving
  • Mentoring
  • Collaboration
  • Leadership
Certifications & Qualifications
  • Bachelor’s Degree in Computer Science or Information Systems
  • Master’s Degree Preferred
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Engineer
Senior Cloud Engineer

Jobtailor • Kentucky

On-site
USD 110,000 - 150,000
Senior Cloud Engineer
Senior Cloud Engineer

Jobtailor • Scottsdale (AZ)

On-site
USD 140,000 - 190,000
VP, Engineering & Operations
VP, Engineering & Operations

Jobtailor • Plano (TX)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • United States

On-site
USD 120,000 - 180,000
Cloud Engineer
Cloud Engineer

Jobtailor • Idaho Falls (ID)

On-site
USD 120,000 - 150,000
Principal Cloud Engineer – AI
Principal Cloud Engineer – AI

Jobtailor • West Chester

On-site
USD 150,000 - 210,000
Cloud Infrastructure DevOps Engineer
Cloud Infrastructure DevOps Engineer

Jobtailor • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Engineer, Applications Systems
Senior Engineer, Applications Systems

Jobtailor • Town of Florida (NY)

On-site
USD 120,000 - 160,000
Principal DevOps Engineer
Principal DevOps Engineer

Jobtailor • Westbrook (ME)

On-site
USD 110,000 - 160,000
Senior Engineer, IT Infrastructure Engineering – Data Center
Senior Engineer, IT Infrastructure Engineering – Data Center

Jobtailor • Atlanta (GA)

Hybrid
USD 140,000 - 180,000