Senior Kubernetes Platform Engineer

Aivar Innovations

Bengaluru

On-site

INR 2,800,000 - 4,600,000

Full time

41 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We seek an experienced Kubernetes-focused engineer to own control loops, multi-cluster tooling, and workload agents that orchestrate GPU/ML workloads across clusters.

You will design and implement CRDs, controllers, and reconciliation logic, while ensuring secure remote operations and robust observability.

Qualifications

  • 5+ years in software, platform, infrastructure, SRE, or cloud engineering.
  • Strong Go. Design production services, concurrency, interfaces, testing, profiling.
  • Deep Kubernetes internals: controllers, CRDs, reconciliation, API machinery, RBAC.
  • Built Kubernetes software, not only operated clusters; Kubebuilder/Operator SDK experience valued.
  • Strong distributed-systems instincts: idempotency, retries, eventual consistency, leader election.
  • Experience operating Kubernetes across environments (cloud and on-prem).
  • Comfort with networking: TCP/TLS, mTLS, proxies, DNS, load balancers, ingress.
  • Production troubleshooting from controllers to workloads.
  • Strong security: identities, certificates, least privilege, auditability.
  • Fluency with agentic coding tools; AI acceleration while ensuring architecture and reliability.

Responsibilities

  • Own Kubernetes control loops: design CRDs, controllers, operators, and reconciliation.
  • Build multi-cluster control for onboarding, authentication, observation, upgrades, control of clusters.
  • Develop workload-cluster agent: registration, heartbeat, inventory, command execution, rollout.
  • Translate platform intent into Kubernetes state: projects, resources, deployments, notebooks, training jobs.
  • Manage resource isolation and placement: namespaces, quotas, limits, scheduling, GPU resources.
  • Integrate GPU support: NVIDIA Operator, device plugins, MIG, topology-aware scheduling.
  • Enable secure remote operations: browser kubectl/exec/log access, mTLS, auditable commands.
  • Deploy platform capabilities: install only necessary operators/services per project/cluster.
  • Design for unreliable environments: handle disconnections, restarts, partial upgrades, repeated reconciliation.
  • Work deeply with Kubernetes API machinery: informers, server-side apply, RBAC, API conventions.

Skills

Go programming
Kubernetes
Distributed systems
Security fundamentals
Cloud experience

Tools

Kubebuilder
controller-runtime
Operator SDK
Custom schedulers

Job description

Aivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We partner with enterprises to deploy intelligent agents that automate complex business processes—from intelligent customer interactions to enterprise knowledge systems.

Aivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We partner with enterprises to deploy intelligent agents that automate complex business processes—from intelligent customer interactions to enterprise knowledge systems.

Experience: 5–9 years | 4+ years building or operating production Kubernetes platforms, controllers, operators, or cloud-native infrastructure

The Role

You build the Kubernetes substrate that makes Kubogent possible.

Kubogent is a Kubernetes-native AI infrastructure and MLOps platform. A central control plane manages multiple workload clusters, allocates infrastructure to tenants and projects, deploys platform capabilities on demand, runs GPU and ML workloads, exposes remote operational access, and continuously reconciles desired state with what is actually running.

What You'll Do
  • Own Kubernetes control loops. Design and build CRDs, controllers, operators, reconcilers, finalizers, watches, status models, and lifecycle state machines using Go and controller-runtime.
  • Build multi-cluster control. Help design and implement how Kubogent onboards, authenticates, observes, upgrades, and controls workload clusters that may only initiate outbound connections to the control plane.
  • Build the workload-cluster agent. Own registration, heartbeat, inventory, command execution, reconnect behaviour, versioning, rollout, credential rotation, and failure recovery.
  • Translate platform intent into Kubernetes state. Convert concepts such as project placement, capabilities, resource allocation, model deployments, notebooks, training jobs, and shared services into safe, idempotent Kubernetes reconciliation.
  • Own resource isolation and placement. Work with namespaces, quotas, limits, priority, scheduling, taints/tolerations, affinity, topology, GPU resources, gang scheduling, and workload placement.
  • Build GPU and accelerator support. Integrate with NVIDIA GPU Operator, device plugins, MIG where appropriate, node feature discovery, topology-aware scheduling, and accelerator-specific runtime requirements.
  • Build secure remote operations. Design mechanisms for browser-based kubectl/exec/log access, tunneled cluster connectivity, least-privilege credentials, mTLS, authorization, and auditable command execution.
  • Own cluster capability deployment. Build mechanisms that install only the operators and services required by capabilities enabled on projects or clusters, rather than treating every cluster as identical.
  • Design for unreliable environments. Clusters disappear, links break, agents restart, watches expire, APIs throttle, upgrades partially fail, and reconciliation gets repeated. Your systems must remain correct anyway.
  • Work deeply with Kubernetes API machinery. Informers, watches, admission, status conditions, server-side apply, resource versions, optimistic concurrency, garbage collection, RBAC, and API conventions.
  • Own platform upgrades. Design safe version skew, agent upgrades, CRD evolution, migration, backward
  • compatibility, and rollout/rollback strategies.
  • Drive observability for the platform itself. Instrument operators and agents with logs, metrics, traces, health checks, queue depth, reconciliation latency, and actionable failure signals.
  • Set the bar. Review designs and code, mentor engineers, and establish patterns for building reliable Kubernetes-native systems.
What We're Looking For
  • 5+ years in software, platform, infrastructure, SRE, or cloud engineering, with substantial hands-on Kubernetes experience.
  • Strong Go. You should be comfortable designing production services, concurrency, interfaces, testing, profiling, and failure handling in Go.
  • Deep Kubernetes internals. Controllers, CRDs, reconciliation, API machinery, watches/informers, RBAC, admission, scheduling, storage, networking, and workload lifecycle.
  • You have built Kubernetes software, not only operated clusters. We especially value experience with Kubebuilder, controller-runtime, Operator SDK, custom schedulers, admission webhooks, or Kubernetes-integrated platforms.
  • Strong distributed-systems instincts. Idempotency, retries, at-least-once execution, eventual consistency, leases, leader election, partial failure, backpressure, and state convergence should be familiar ideas.
  • Experience operating Kubernetes across environments. Cloud-managed Kubernetes and on-prem/private-cloud experience are both valuable.
  • Comfort with networking. TCP/TLS, mTLS, proxies, reverse tunnels, WebSockets or streaming RPC, DNS, load balancers, ingress/gateway, and debugging connectivity failures.
  • Production troubleshooting ability. You can move from symptom to root cause across controllers, API servers, networking, scheduling, container runtimes, storage, and workloads.
  • Strong security fundamentals. Service identities, certificates, RBAC, secrets, credential rotation, least privilege, tenant isolation, and auditability.
  • Fluency with agentic coding tools. We expect AI to accelerate implementation and investigation while you remain responsible for architecture, correctness, failure handling, and operational quality.
Strong Pluses
  • Multi-cluster management platforms.
  • Kubernetes API aggregation or extension patterns.
  • Cluster API, Crossplane, Argo CD, Flux, Rancher, Rafay, Open Cluster Management, or similar systems.
  • Envoy, reverse tunnels, relay systems, or secure remote cluster access.
  • GPU scheduling, NVIDIA GPU Operator, MIG, CUDA-aware workloads, or distributed training infrastru
  • Prometheus, OpenTelemetry, Loki, or Kubernetes observability stacks.
  • Kubernetes conformance, upgrade testing, chaos testing, or large-scale fleet management.
This Role May Not Be The Best Match If
  • Your Kubernetes experience is primarily writing manifests, Helm charts, and maintaining CI/CD pipelines. We arebuilding Kubernetes-native control-plane software.
  • You expect infrastructure to be reliable and synchronous. Disconnected clusters, duplicate commands, partial
  • upgrades, stale state, and repeated reconciliation are normal operating conditions here.
  • You prefer solving platform problems by adding manual operational procedures. Kubogent must turn those procedures into productized, automated control loops.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer
Senior Platform Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 1,800,000 - 2,500,000
Lead Engineer - Managed Kubernete
Lead Engineer - Managed Kubernete

Target • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Software Engineer II - Kubernetes
Software Engineer II - Kubernetes

sghcorp.com • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Azure Kubernetes Engineer
Azure Kubernetes Engineer

Neurealm • Chennai District

On-site
INR 1,200,000 - 2,400,000
Platform Engineer
Platform Engineer

Rakuten • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Platform Engineer
Platform Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 800,000 - 1,500,000
Lead Engineer - Managed Kubernetes
Lead Engineer - Managed Kubernetes

Target • Bengaluru

On-site
INR 3,500,000 - 5,200,000
Senior Frontend Engineer
Senior Frontend Engineer

Keka Technologies Private Limited • Coimbatore District

On-site
INR 4,000,000 - 7,500,000
Senior MLOps / AI Platform Engineer
Senior MLOps / AI Platform Engineer

Keka Technologies Private Limited • Coimbatore District

On-site
INR 3,000,000 - 6,000,000
Software Engineer II - Kubernetes (Bangalore, IN, 560048)
Software Engineer II - Kubernetes (Bangalore, IN, 560048)

Penguin Computing • India

Hybrid
INR 2,500,000 - 4,000,000