Forward Deployed Engineer (FDE)

Emergys Corp.

Pune District

On-site

INR 1,800,000 - 3,000,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Emergys Corp. in Pune, India, seeks a Senior Platform Engineer to own and operate the cluster and platform layer for managed inference.

You will build cloud-connected Kubernetes control planes, join AI accelerator nodes as workers, and deploy the vendor platform stack via Helm charts while maintaining security, observability, and automation. This role requires strong Linux, Kubernetes, and cloud experience, with on-call rotation and occasional data-centre/site presence.

Qualifications

  • Strong Linux administration to diagnose service, storage, network and kernel problems.
  • Production Kubernetes lifecycle experience: building, upgrading, and recovering clusters.
  • Helm proficiency beyond chart installation; values management and multi-chart upgrades.
  • Hands-on experience with at least one major public cloud and knowledge of a second.
  • Terraform and Ansible at production scale as reusable code.
  • Networking fundamentals: routing, NAT, DNS, TLS termination; debug hybrid connectivity.
  • OIDC authentication knowledge and integration with Kubernetes.
  • Experience running Prometheus/Grafana and centralised log pipeline.
  • Scripting in Python and Bash; YAML-heavy configuration.
  • Strong ownership, on-call availability, ability to write runbooks.

Responsibilities

  • Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.
  • Register, label and taint accelerator worker nodes so workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.
  • Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.
  • Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.
  • Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.
  • Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.
  • Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.
  • Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.
  • Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.
  • Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.
  • Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.
  • Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.
  • Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.
  • Support model bundle and deployment configuration changes through the platform’s Kubernetes custom resources, in coordination with ML systems engineers.
  • Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each change.

Skills

Linux administration
Kubernetes
Networking
Terraform
Ansible
CI/CD
Python Bash

Tools

RKE2
Helm
Terraform
Ansible
Kustomize
OCI Helm registries
Prometheus

Job description

This role owns the cluster and the platform layer that the managed inference service runs on. You will build Kubernetes control planes in the cloud, join AI accelerator nodes hosted in the data centre as worker nodes, install and configure the vendor AI platform stack on top, and keep the whole thing running, patched and observable. The hardware vendor retains the physical rack, the on-premise network and the local tunnel endpoint. Everything from the cloud network and our side of the tunnel upward is ours, and this role is the technical owner of most of it. On-call rotation and occasional data-centre or customer-site presence are part of the job.

Key Responsibilities
  • Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.
  • Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.
  • Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.
  • Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.
  • Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.
  • Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.
  • Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.
  • Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.
  • Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.
  • Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.
  • Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.
  • Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.
  • Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.
  • Support model bundle and deployment configuration changes through the platform’s Kubernetes custom resources, in coordination with ML systems engineers.
  • Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each change.
Tools and Technologies

Category

Tools and technologies

Kubernetes and platform RKE2, Kubernetes, kubectl, Helm v3.19 and above, OCI Helm registries, Kustomize, custom resource definitions and operators, etcd, node labels, taints and tolerations, device plugins

Operating systems RHEL 8 and 9, systemd, SELinux, dnf and Satellite or equivalent patch management, kernel and sysctl tuning, LVM, Bash

Cloud platforms AWS, Azure or GCP in depth plus a working second: virtual networks and subnets, route tables, security groups, peering and transit gateways, VPN gateways, IAM, KMS, managed DNS, object storage

Networking and connectivity IPSec site-to-site VPN, BGP and static routing, Calico or Cilium, MetalLB and cloud load balancers, ingress-nginx or Traefik, CoreDNS, TLS and PKI, cert-manager

IaC, GitOps and CI/CD Terraform including module design and remote state, Ansible, Helmfile, Argo CD or Flux, GitHub Actions, GitLab CI or Jenkins, Git, Python, Bash

Storage and backup CSI drivers, local path and local PV provisioners, NFS and object storage, Velero, etcd snapshot and restore

Identity and security Keycloak or an equivalent OIDC provider, OAuth 2.0, OIDC and SAML, Kubernetes RBAC, HashiCorp Vault or a cloud secrets manager, CIS benchmarks, Trivy, Falco, audit logging

Observability Prometheus and the Prometheus operator, kube-prometheus-stack, Grafana, Alertmanager, Node Exporter, Fluent Bit, OpenSearch, Loki, OpenTelemetry collectors

AI platform layer Helm-packaged inference platform stacks, accelerator node scheduling, model bundle and deployment custom resources, OpenAI-compatible inference endpoints, operator-managed PostgreSQL, Redis queues

AI-assisted engineering Agentic coding and operations assistants used for infrastructure code generation, review and diagnostics, with human verification before production

  • Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network,and kernel problems without escalation.
  • Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.
  • Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.
  • Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.
  • Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.
  • Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.
  • Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.
  • Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.
  • Scripting in Python and Bash, and comfort with YAML-heavy configuration.
  • Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call rotation.
Preferred Requirements
  • Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.
  • Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.
  • Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.
  • Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments.
  • CKA and CKS, plus a professional or expert level cloud certification. Foundational certifications carry limited weight.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Devops Engineer
Devops Engineer

Airtel • Gurugram District

On-site
INR 800,000 - 1,400,000
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)

Annova Solutions • Indore District

On-site
INR 2,800,000 - 4,200,000
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)

Annova Solutions Corp. • Indore District

On-site
INR 2,400,000 - 4,200,000
Agentic AI developer
Agentic AI developer

Tredence • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Senior AI Applications Engineer
Senior AI Applications Engineer

GE HealthCare • Bengaluru

On-site
INR 1,500,000 - 2,700,000
Senior Kubernetes and Open Source Platform Engineer
Senior Kubernetes and Open Source Platform Engineer

Zappsec Inc. • India

On-site
INR 3,500,000 - 7,500,000
Cloud Engineer
Cloud Engineer

APEX Analytix, LLC • India

On-site
INR 4,000,000 - 7,000,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Innovalus Technologies • Mumbai

On-site
INR 400,000 - 650,000
Kubernetes Data Platform
Kubernetes Data Platform

Tata Consultancy Services • Bengaluru Urban

On-site
INR 1,800,000 - 3,000,000
Senior Staff Engineer, DevOps
Senior Staff Engineer, DevOps

Sierra Wireless • India

On-site
INR 2,500,000 - 4,000,000