Overview We are seeking a highly skilled Staff Software Engineer with deep expertise in DevOps, Site Reliability Engineering (SRE), Cloud Platforms, Kubernetes, and GitOps practices. This role will architect, scale, and operate enterprise-grade Kubernetes platforms while driving reliability, scalability, automation, and operational excellence across critical enterprise applications and infrastructure. The ideal candidate will provide technical leadership on Kubernetes platform strategy, influence engineering best practices, and partner with cross-functional teams to build resilient, secure, and highly available cloud-native platforms at scale. This position requires deep hands-on experience with Kubernetes architecture, ArgoCD, GitOps methodologies, Infrastructure Automation, and Production Engineering.
What You'll Do
- DevOps & Platform Engineering Design, build, and optimize scalable CI/CD pipelines and Kubernetes-native platform solutions. Drive Infrastructure as Code (IaC), automation, and platform standardization initiatives. Improve developer experience through self-service infrastructure and deployment automation. Lead architecture discussions and establish Kubernetes and DevOps platform standards across engineering teams.
- Site Reliability Engineering (SRE) Define and champion reliability standards, SLAs, SLOs, and operational excellence frameworks. Lead incident response, root cause analysis, and reliability improvement programs. Drive performance optimization, scalability enhancements, capacity planning, and disaster recovery strategies. Build proactive monitoring, observability, and alerting capabilities to improve system health and availability.
- GitOps & Automation Architect and manage GitOps practices using ArgoCD for multi-cluster Kubernetes application delivery. Automate application deployment, configuration management, and environment provisioning. Establish deployment governance, release management processes, and operational controls. Ensure secure, consistent, and auditable deployments across all environments.
- Kubernetes & Cloud Infrastructure Define and evolve enterprise Kubernetes platform architecture, standards, and multi-cluster/multi-region deployment strategies. Design, deploy, and operate production Kubernetes clusters at scale across cloud environments (EKS, AKS, GKE, or equivalent). Build and optimize containerized platform solutions using Docker and Kubernetes for high availability and performance. Lead adoption of Kubernetes-native tooling including Helm, Kustomize, operators, and service mesh technologies (Istio, Linkerd, or equivalent). Drive cluster lifecycle management including upgrades, autoscaling, capacity planning, disaster recovery, and cost optimization. Architect Kubernetes networking, ingress, storage (CSI), CNI, and workload isolation patterns for complex enterprise workloads. Establish Kubernetes security frameworks including RBAC, network policies, pod security standards, secrets management, and DevSecOps integration. Drive Infrastructure as Code adoption using Terraform or equivalent technologies for cloud and Kubernetes infrastructure. Partner with Security and Engineering teams to implement platform governance, compliance, and operational controls.
- Technical Leadership & Collaboration Provide technical leadership and mentorship to engineers across DevOps and platform teams. Collaborate with Engineering, Product, Infrastructure, and Security stakeholders to drive strategic initiatives. Participate in architecture reviews and technology roadmap discussions. Work closely with global teams across EMEA and US regions. Demonstrate flexibility to work across overlapping business hours and shifts when required to support global stakeholders and critical production environments.
What We Are Looking For
- 12+ years of experience in Software Engineering, DevOps, Platform Engineering, or Site Reliability Engineering (SRE).
- Strong hands-on expertise in DevOps, SRE, CI/CD, Infrastructure Automation, and Cloud Engineering.
- Proven experience implementing GitOps practices using ArgoCD.
- Deep expertise in Kubernetes platform architecture, cluster operations, and troubleshooting at enterprise scale, including Docker, Microservices, and Cloud Platforms (AWS/Azure/GCP).
- Proven experience architecting, deploying, and operating large-scale production Kubernetes platforms in enterprise environments.
- Expert-level hands-on experience with Kubernetes tooling such as Helm, Kustomize, kubectl, cluster APIs, and platform engineering frameworks.
- Experience leading Kubernetes migration, modernization, and cloud-native transformation initiatives across multiple teams.
- Strong understanding of Kubernetes internals, container orchestration patterns, and cloud-native architecture at scale.
- Experience with Infrastructure as Code (Terraform or equivalent).
- Strong understanding of observability, monitoring, alerting, logging, and production operations.
- Experience with monitoring tools such as Prometheus, Grafana, Datadog, Splunk, ELK, New Relic, or similar platforms.
- Strong troubleshooting, incident management, root cause analysis, and production support experience.
- Experience developing automation and operational tooling using Python, Shell, Go, or similar technologies.
- Strong understanding of security, compliance, and reliability engineering principles.
- Excellent stakeholder management, communication, and technical leadership skills.
- Experience working with globally distributed teams across EMEA and US regions.
- Ability to drive technical decisions, define Kubernetes platform roadmaps, influence engineering practices, and lead large-scale platform initiatives.
- Flexibility to collaborate with global stakeholders across multiple time zones and support business-critical operations when needed.
Our Values
If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.
Who are we? We are a proven, passionate bunch of disruptors. Our work is all about tapping into your potential so we can deliver the best solutions and customer experiences on the planet. Collaboration, respect, and a great work-life balance earned us the title of "Best Place to Work- Employees' Choice" by Glassdoor. Our people are smart, creative, rock stars with over 400 patents and 10,000 people years of domain expertise.
What do we do? Blue Yonder is the world leader in digital supply chain and omni-channel commerce fulfillment. Our intelligent, end-to-end platform enables retailers, manufacturers and logistics providers to seamlessly predict, pivot and fulfill customer demand. With Blue Yonder, you can make more automated, profitable business decisions that deliver greater growth and re-imagined customer experiences. Blue Yonder - Fulfill your Potential. blueyonder.com
Blue Yonder is a trademark or registered trademark of Blue Yonder, Inc. Any trade, product or service name referenced in this document using the name “Blue Yonder” is a trademark and/or property of Blue Yonder, Inc. Blue Yonder, Inc. 15059 N Scottsdale Rd, Ste 400 Scottsdale, AZ 85254