CMK Resources is seeking a Senior Container Platform Engineer to support one of our strategic partners and their enterprise customer. This individual will be responsible for the engineering, reliability, security, and lifecycle management of enterprise container platforms across production environments.
This is a hands-on engineering role for someone with deep Kubernetes and/or Red Hat OpenShift experience who can go beyond day-to-day administration to design, automate, troubleshoot, and continuously improve enterprise container platforms. You will work closely with application, security, network, storage, and CI/CD teams to deliver a stable, secure, automated, and scalable platform for containerized workloads.
What You Will Be Doing
- Design, deploy, administer, and upgrade production Kubernetes and/or Red Hat OpenShift clusters across on-premises and cloud environments.
- Engineer core platform capabilities including namespaces, operators, ingress, service discovery, load balancing, storage classes, secrets management, certificates, and image registries.
- Build reusable automation using technologies such as Ansible, Terraform, Helm, GitOps, Python, and shell scripting to reduce manual effort and configuration drift.
- Integrate container platforms with CI/CD pipelines, artifact repositories, vulnerability scanners, identity providers, logging, monitoring, and IT service management tools.
- Define and implement platform standards for high availability, capacity management, backup and recovery, disaster recovery, patching, upgrades, and lifecycle management.
- Implement security controls including RBAC, network policies, pod security standards, image governance, secrets protection, and compliance requirements.
- Troubleshoot complex Kubernetes/OpenShift, Linux, networking, storage, container runtime, and application deployment issues.
- Lead root cause analysis efforts and implement permanent corrective actions for recurring or high-impact platform issues.
- Create reference architectures, technical standards, runbooks, operational dashboards, and knowledge-transfer materials.
- Mentor engineers and L2 support teams while providing technical leadership during major incidents and platform changes.
What We Are Looking For
- 6+ years of experience in infrastructure, platform engineering, DevOps, SRE, Linux systems engineering, or a related field.
- 4+ years of hands-on experience supporting production Kubernetes, Red Hat OpenShift, or comparable enterprise container orchestration environments.
- Strong Linux administration and troubleshooting experience.
- Strong understanding of TCP/IP networking, DNS, TLS/PKI, load balancing, and enterprise storage fundamentals.
- Hands-on experience with container runtimes and tooling such as CRI-O, containerd, Docker/Podman, Helm, kubectl, and/or oc.
- Experience building automation and CI/CD integrations using technologies such as Git, Ansible, Terraform, Jenkins, GitLab CI, GitHub Actions, or equivalent tools.
- Experience designing, deploying, upgrading, and supporting highly available, business-critical container platforms.
- Strong troubleshooting skills with the ability to diagnose complex platform, infrastructure, and application deployment issues.
- Experience participating in incident, change, and problem-management processes.
- Strong documentation, communication, and cross-functional collaboration skills.
Preferred Experience
- CKA, CKAD, CKS, Red Hat Certified OpenShift Administrator, or similar certification.
- Public cloud Kubernetes experience with platforms such as EKS, AKS, or GKE.
- GitOps experience with tools such as Argo CD.
- Observability experience with Prometheus, Grafana, Splunk, Elastic, or similar platforms.
- Knowledge of service mesh, policy-as-code, software supply-chain security, and container cost/capacity optimization.
What Success Looks Like
- Maintain platform availability, performance, reliability, and capacity across business-critical environments.
- Deliver cluster upgrades, patching, and lifecycle activities with minimal disruption and controlled risk.
- Reduce deployment lead time and manual support effort through automation and self-service capabilities.
- Reduce recurring incidents, security findings, and configuration drift through proactive engineering and root cause resolution.
- Establish strong platform standards, documentation, and knowledge-sharing practices across engineering and support teams.