Systems Engineer - Cloud Ops

AutoZone

Memphis (TN)

On-site

USD 110,000 - 150,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

AutoZone is seeking a Systems Engineer for the Cloud Operations team to design, deploy, and optimize cloud infrastructure on Google Cloud Platform (GCP). You will work with Terraform, GKE, GitOps/ArgoCD, CI/CD pipelines, and observability tools to ensure reliable and scalable platform operations.

You will contribute to AI/ML platform initiatives and support infrastructure for LLM-based applications and AI-enabled automation, collaborating with cross-functional teams to deliver high availability

Qualifications

  • Kubernetes in production environments is essential.

Responsibilities

  • Design, build, and maintain cloud infrastructure in GCP.
  • Develop CI/CD pipelines and GitOps workflows.
  • Monitor performance and troubleshoot production issues.
  • Manage Kubernetes deployments and cluster upgrades.
  • Support AI/ML platform workloads and tools.
  • Participate in on-call rotation and incident response.
  • Collaborate with development teams, SREs, and architects.

Skills

Kubernetes expertise
GCP experience
GitOps
CI/CD pipelines
Observability
Automation scripting
On-call readiness
Cross-functional collaboration

Education

Bachelor's degree in CS/IT or related field
Equivalent experience

Tools

Terraform
GKE
ArgoCD
GitLab CI
Helm
Kustomize
Istio
Dynatrace
Prometheus
Grafana

Job description

As a Systems Engineer on the Cloud Operations team, you will be responsible for deploying, managing, and optimizing our cloud-based infrastructure on Google Cloud Platform (GCP). You will work with technologies such as Terraform, Kubernetes (GKE), GitOps/ArgoCD, CI/CD pipelines, and observability tools to ensure reliable, secure, and scalable platform operations. You will also contribute to our AI/ML platform initiatives, supporting infrastructure for LLM-based applications and AI-powered automation tools that enhance developer productivity and operational efficiency. You will collaborate with development teams, SREs, and platform architects to ensure seamless deployment and delivery of applications while maintaining the highest standards of reliability, security, and performance.

Cloud Infrastructure, Automation & Operations
  • Design, build, and maintain cloud infrastructure using Terraform to automate provisioning, scaling, and lifecycle management of resources on GCP
  • Develop and maintain CI/CD pipelines using GitLab CI to automate build, test, and deployment workflows. Implement and maintain GitOps practices using ArgoCD for declarative, version-controlled application deployment
  • Monitor system performance using observability tools (Dynatrace, Cloud Monitoring, Prometheus/Grafana) and troubleshoot production issues
  • Participate in on-call rotation to provide 24/7 support for critical infrastructure incidents
  • Perform root cause analysis on incidents and implement preventive measures. Document runbooks, architecture decisions, and operational procedures
Kubernetes Platform Management
  • Deploy, configure, and manage containerized applications on Google Kubernetes Engine (GKE), including GKE Autopilot and Standard clusters Manage cluster lifecycle including upgrades, node pool configurations, and capacity planning
  • Troubleshoot pod failures, CrashLoopBackOff, OOMKilled events, and container resource issues
  • Configure and optimize resource requests/limits, Horizontal Pod Autoscaler (HPA), and Vertical Pod Autoscaler (VPA)
  • Manage Kubernetes networking including Services, Ingress controllers, Network Policies, and DNS configurations. Implement and manage service mesh (Istio) for traffic management, observability, and security
  • Manage secrets and configurations using Kubernetes Secrets, ConfigMaps, and external secret management tools. Implement pod security standards, RBAC policies, and workload identity configurations
AI/ML Platform & Automation
  • Support infrastructure for AI/ML workloads including LLM-based applications and model serving platforms
  • Deploy and manage AI-powered developer tools such as coding assistants (Claude Code, GitHub Copilot) and agentic AI systems. Explore and implement AI-assisted incident response and automated remediation workflows
  • Build and maintain infrastructure for Retrieval-Augmented Generation (RAG) pipelines and vector databases
  • Configure GPU-enabled node pools and optimize resource allocation for AI/ML workloads
  • Implement MCP (Model Context Protocol) servers and AI agent integrations for operational automation
  • Stay current with emerging AI technologies and evaluate their applicability for infrastructure automation
Kubernetes Expertise (Essential)
  • 3+ years hands‑on experience with Kubernetes in production environments
  • Deep understanding of Kubernetes architecture: API server, etcd, scheduler, controller manager, kubelet
  • Experience with GKE (Standard and Autopilot modes), including cluster creation, upgrades, and maintenance
  • Proficiency in troubleshooting workloads: analyzing pod logs, events, describe outputs, and container states
  • Strong understanding of resource management: requests, limits, QoS classes, and resource quotas
  • Experience with Kubernetes networking: Services (ClusterIP, NodePort, LoadBalancer), Ingress, Network Policies
  • Knowledge of Kubernetes storage: PersistentVolumes, PersistentVolumeClaims, StorageClasses, dynamic provisioning
  • Experience with Helm charts for application packaging and deployment
  • Familiarity with Kubernetes security: RBAC, Pod Security Standards, Secrets management, Workload Identity
  • Understanding of Kubernetes observability: metrics-server, kubectl top, container resource monitoring
  • Experience debugging common issues: ImagePullBackOff, CrashLoopBackOff, OOMKilled, Evicted pods, pending pods
Cloud & Infrastructure
  • 3+ years of experience with Google Cloud Platform (GCP) services including GKE, Cloud Run, Cloud SQL, Memorystore, Pub/Sub, and Cloud Logging
  • Strong experience with Terraform for infrastructure as code (IaC)
  • Understanding of cloud networking: VPCs, subnets, firewall rules, Cloud NAT, Private Service Connect
CI/CD & GitOps
  • Proficiency with GitLab CI/CD pipelines
  • Experience with ArgoCD or similar GitOps tools
  • Understanding of Helm charts and Kustomize for Kubernetes manifest management
Observability & Troubleshooting
  • Experience with monitoring and APM tools (Dynatrace, Datadog, Prometheus, Grafana)
  • Ability to analyze logs, metrics, and traces to diagnose production issues
  • Familiarity with JVM troubleshooting (heap dumps, thread analysis, GC tuning, connection pool issues)
AI/ML Knowledge
  • Basic understanding of LLM concepts, prompt engineering, and AI model deployment
  • Familiarity with AI coding assistants and their integration into development workflows
  • Interest in agentic AI systems and autonomous automation tools
  • Exposure to vector databases (Pinecone, Weaviate, pgvector) and RAG architectures is a plus
Systems & Networking
  • Strong Linux administration skills
  • Understanding of networking concepts (DNS, load balancing, firewalls, TCP/IP)
  • Experience with service mesh (Istio) is a plus
General
  • Excellent problem-solving and analytical skills
  • Strong written and verbal communication
  • Ability to work effectively in a collaborative, cross-functional environment
  • Experience working in an Agile/DevOps culture
  • Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent experience)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Jobtailor • Town of Florida (NY)

On-site
USD 180,000 - 240,000
Senior Staff Engineer – DevOps
Senior Staff Engineer – DevOps

Jobtailor • California (MO)

On-site
USD 140,000 - 180,000
Platform Engineer
Platform Engineer

Jobtailor • Denver (CO)

On-site
USD 140,000 - 190,000
GCP Cloud Platform Automation Engineer
GCP Cloud Platform Automation Engineer

TechDigital Group • New York (NY)

On-site
USD 120,000 - 150,000
Principal Cloud Engineer – AI/ML
Principal Cloud Engineer – AI/ML

Jobtailor • Town of Florida (NY)

On-site
USD 150,000 - 190,000
Senior Platform Engineer
Senior Platform Engineer

Jobtailor • Denver (CO)

On-site
USD 120,000 - 150,000
Senior Cloud Engineer Multi-Cloud (GCP. AWS, and Azure)
Senior Cloud Engineer Multi-Cloud (GCP. AWS, and Azure)

Creative Solutions Services, LLC • Irving (TX)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

NexTier Completion Solutions Inc. • Houston (TX)

On-site
USD 110,000 - 150,000
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE

ContractStaffingRecruiters.com • Stamford (CT)

Hybrid
USD 120,000 - 150,000
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE

ContractStaffingRecruiters.com • Branford (CT)

Remote
USD 120,000 - 160,000