Senior Cloud Engineer

APEX Analytix, LLC

India

On-site

INR 2,500,000 - 4,500,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

APEX Analytix, LLC is seeking a production Kubernetes specialist to design, build, and operate clusters across bare-metal, virtualization, and cloud environments. The role emphasizes GitOps with Flux, Kustomize, and Helm, plus strong identity work with Active Directory and Entra ID.

You will manage RBAC, backups, and recovery, while integrating VMware and Windows infrastructure. Ideal candidates will have CKAs or equivalent depth, hands-on AD/M365, and a proven track record with cluster

Qualifications

  • Production Kubernetes experience including cluster lifecycle.
  • Certified Kubernetes Administrator (CKA) or equivalent depth.
  • Production Active Directory and Microsoft 365 experience.

Responsibilities

  • Design, build, upgrade and operate Kubernetes across bare-metal, virtualised and cloud environments.
  • Own the control plane: etcd backup/restore, certificate rotation, upgrades, RBAC and admission.
  • GitOps as operating model with Flux, Kustomize, Helm; infra as code with OpenTofu/ Terraform/ Ansible.

Skills

Kubernetes admin
GitOps
Active Directory
Microsoft 365
CI/CD pipelines

Tools

Kubernetes
Flux
Kustomize
Helm
Harbor
Keycloak
Entra ID
Entra Connect
Exchange Online
vSphere/ESXi
Windows Server
OpenTofu
Ansible
PowerShell
Prometheus
Grafana
AlmaLinux
Ubuntu
Go
Python
Bash
AKS
Cluster API
KubeVirt

Job description

At apexanalytix, we’re lifelong innovators! Since the date of our founding nearly four decades ago we’ve been consistently growing, profitable, and delivering the best procure-to-pay solutions to the world. We’re the perfect balance of established company and start-up. You will find a unique home here.


And you’ll recognize the names of our clients. Most of them are on The Global 2000. They trust us to give them the latest in controls, audit and analytics software every day. Industry analysts consistently rank us as a top supplier management solution, and you’ll be helping build that reputation.


Most cloud jobs in India are someone else's public cloud. This one is not. The prime focus is Kubernetes on hardware we own: clusters on metal, and tenant clusters those hosts provision. Git is the only path to production. You would own that platform.


You will not get an interview on Active Directory and Microsoft 365 alone, and you will not get one on a Kubernetes certificate alone. The person this is written for has already operated both, and wants Kubernetes to be the majority of the work going forward. On-prem Exchange is useful. It is not the job.


The stack you will work with. Kubernetes (kubeadm, Cluster API, KubeVirt, hosted control planes, AKS), Flux, Kustomize, Helm, Harbor, Keycloak, Active Directory, Entra ID, Entra Connect, Microsoft 365, Exchange Online, VMware vSphere/ESXi, Windows Server, Cisco UCS, OpenTofu, Ansible, PowerShell, Prometheus, Grafana, AlmaLinux, Ubuntu, Go, Python, Bash


What you would own


  • Cluster lifecycle, both layers. Design, build, upgrade and operate Kubernetes across bare-metal (kubeadm), virtualised (KubeVirt with Cluster API and hosted control planes) and managed cloud (AKS). Own the control plane: etcd backup and restore, certificate rotation, upgrades, RBAC and admission. Operate Cluster API so create, upgrade and retire is a repeatable path, not an artisan one. Workload topology (disruption budgets, anti-affinity, requests and limits) and multi-tenancy that actually holds. Diagnose the hard failures and turn them into runbooks.

  • GitOps as the operating model. Flux, Kustomize, Helm. Desired state lives in Git, not in someone's shell history. Infrastructure as code with OpenTofu or Terraform and Ansible. Automate toil in PowerShell, Bash, and Python or Go. Written blast-radius and a rollback path before production changes.

  • Cluster networking and storage, as a platform consumer. Own service exposure on the clusters: Gateway API and HTTPRoute, ingress, kube-vip and CoreDNS. Use NetworkPolicy, StorageClasses, PVCs, snapshots and the CSI drivers the platform is given. Diagnose whether a fault is the application, the CNI, the CSI or the layer beneath, then work it with the network or storage engineer rather than owning those layers. Deep Cilium and Ceph operations sit with those roles.

  • Active Directory and hybrid identity. Architect and maintain Active Directory: forest and domain design, Group Policy at scale, replication health, trusts, schema changes, and identity hygiene across multi-domain or multi-site environments. Manage hybrid identity through Entra ID and Entra Connect. Own enterprise DNS and DHCP at scale, including AD-integrated DNS, conditional forwarders, scope design, failover and reservations. Join Linux and Windows hosts into the directory where the platform requires it, and integrate Kubernetes access with the identity you already run. Windows Server lifecycle, provisioning, patching, hardening, performance tuning and decommissioning, is part of this, including Windows worker nodes on the clusters.

  • Microsoft 365 and application mail. Exchange Online, hybrid mail flow, and the Microsoft 365 admin surface. Application email enablement: SMTP relay, OAuth 2.0 for mail, SPF, DKIM and DMARC. On-prem Exchange Server is useful if you have it. It is not required.

  • VMware and KubeVirt virtualization. Operate the VMware estate (vSphere/ESXi, vCenter, HA, VM lifecycle) and the boundary with Kubernetes. On the Kubernetes side, KubeVirt is how VM workloads run on the clusters, on KVM. You decide which workload belongs where rather than lifting everything by default.

  • Platform services. Container registry (Harbor or equivalent) and cluster identity (Keycloak or equivalent, federated with Active Directory and Entra ID), or the depth to take both on. Secrets, cert-manager, and image policy as part of how the fleet is run.

  • Kubernetes-native backup and restore. Own backup and disaster recovery for stateful workloads with Velero, Kopia and etcd snapshots, and prove it by restoring, not by reading a green job status. A backup nobody has restored is not a backup.

  • Observability and reliability. Prometheus and Grafana the team actually uses. On-call as a shared rotation, incident command when the platform is the fault, and a root cause analysis afterwards. Capacity planning, and DR you have tested rather than documented.

  • The path from a new server to a production node. Repeatable provisioning, not a manual build: out-of-band management (IPMI, iLO, iDRAC), RAID, SAN/NAS attach, Cisco UCS, firmware, then kubeadm or Cluster API with CNI and CSI validated before the node takes workload. Linux and Windows worker nodes on the same clusters.


What the first year looks like


  • Make cluster create, upgrade and retire a GitOps path that someone other than you can run.

  • Establish verified high availability on production: replica counts, host anti-affinity and disruption budgets, applied and checked continuously rather than reconstructed after an outage.

  • Establish provable backup and restore for the Kubernetes platform and its stateful workloads, with independently witnessed evidence. This is a stated company priority, and the platform half of it would be yours.

  • Bring Active Directory, hybrid identity, Microsoft 365 and the VMware estate under the same operating standard as the Kubernetes platform: documented, monitored, and not a single point of knowledge.

  • Remove the single points of knowledge, so no part of the platform depends on one person being reachable.


What we need you to have done

Three hard requirements. Miss any one and this is not the right role.


1. Production Kubernetes you have operated, including cluster lifecycle. Not a single managed cluster you clicked in a console, and not a certificate you collected to know it. Walk us through a cluster you built or upgraded, a control-plane failure you diagnosed, and how desired state got from Git onto the nodes.


2. Certified Kubernetes Administrator (CKA), or equivalent depth you can demonstrate live. The certificate is the floor. We will still ask you to debug a real failure. CKS or CKAD is a plus. Neither substitutes for (1).


3. Production Active Directory and Microsoft 365 you have operated. Multi-site or multi

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Engineer
Cloud Engineer

APEX Analytix, LLC • India

On-site
INR 4,000,000 - 7,000,000
Senior Kubernetes and Open Source Platform Engineer
Senior Kubernetes and Open Source Platform Engineer

Zappsec Inc. • India

On-site
INR 3,500,000 - 7,500,000
Kubernetes Data Platform
Kubernetes Data Platform

Tata Consultancy Services • Bengaluru Urban

On-site
INR 1,800,000 - 3,000,000
Forward Deployed Engineer (FDE)
Forward Deployed Engineer (FDE)

Emergys Corp. • Pune District

On-site
INR 1,800,000 - 3,000,000
Senior DevOps Engineer
Senior DevOps Engineer

Benchmarkit • Pune District

On-site
INR 1,400,000 - 2,400,000
Lead DevOps Engineer
Lead DevOps Engineer

Marktine Technology Solutions Pvt Ltd • Jaipur

On-site
INR 2,500,000 - 4,200,000
Senior DevOps Engineer
Senior DevOps Engineer

Asymbl • Jaipur

On-site
INR 3,000,000 - 6,000,000
Senior Backend / Distributed Systems Engineer
Senior Backend / Distributed Systems Engineer

Keka Technologies Private Limited • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Senior Platform Engineer
Senior Platform Engineer

InfoVision Inc. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior DevOps Engineer
Senior DevOps Engineer

Asymbl • Rajasthan

On-site
INR 2,500,000 - 4,500,000