Network Consultant II

Gruve

Maharashtra

On-site

INR 1,800,000 - 3,000,000

Full time

28 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Gruve in Maharashtra is seeking a Senior engineer and shift anchor for the NOC pod and L2 for PulseAI infrastructure operations. Owns P1/P2 first response and bridges to L3, remediate GPU-server and OpenShift-network faults, and coordinates vendor escalations to restoration objectives.

You will lead complex changes, manage hardware incidents across OpenShift/GKE, monitor fabric health, and mentor juniors while aligning with security and data center operations.

Qualifications

  • 4–6 years network operations in data-center environments.
  • Strong hands-on data-center switching/routing and next-generation firewall experience; incident bridge experience.
  • Working understanding of Kubernetes networking — CNI (Cilium), Services/Ingress and east-west flows — to troubleshoot fabric-to-GKE connectivity end to end.
  • Hands-on Red Hat OpenShift / Kubernetes operations in production — node lifecycle and MachineConfig, operators, cluster networking (OVN-Kubernetes/CNI), storage, oc/kubectl troubleshooting — plus Linux (RHEL) administration.
  • Working understanding of GPU-server operations: NVIDIA GPU Operator and driver stack, DCGM-class telemetry, firmware/driver update procedures, RDMA/RoCEv2 NIC health and common GPU failure modes.
  • Multi-vendor switch monitoring and troubleshooting across data-center platforms; out-of-band management proficiency; disciplined change execution.
  • Mentoring aptitude for the junior bench.

Responsibilities

  • Act as shift anchor: own P1/P2 first response, open and run technical bridges until L3 engagement.
  • Execute complex changes (fabric node additions, firewall policy pushes) with senior oversight.
  • Remediate PulseAI infrastructure incidents within Gruve's scope: GPU-server firmware and driver faults, OpenShift node and cluster-networking issues, front-end and RoCEv2 back-end fabric connectivity, switch configuration within the PulseAI fabric; diagnose storage capacity/health and node-hardware faults and escalated to the vendor with complete diagnostics.
  • Execute firmware and GPU driver updates, switch configuration changes and node lifecycle operations within agreed maintenance windows with the required customer notice; verify post-change health.
  • Manage hardware escalations: drive OEM, neocloud-provider and storage-vendor cases to closure, deliver daily status on Premium, and manage restoration-clock pauses correctly.
  • Drive preventive maintenance: firmware/BIOS/driver currency, EOL/EOS tracking, spare posture — for network devices and PulseAI GPU and control-plane nodes.
  • Own infrastructure observability for the NOC pod: GPU-cluster and fabric health dashboards, log queries and alert thresholds in Grafana; contribute to onboarding telemetry checks.
  • Quality-gate tickets, handovers and runbook adherence across the pod; mentor juniors; coordinate with the SOC anchor on security-relevant network events; coordinate facility alerts with data-center operations.

Skills

Network operations
OpenShift
Kubernetes
GPU/Driver management
Bridge incident management
Vendor escalation

Tools

NVIDIA GPU Operator
OpenShift
GKE networking

Job description

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position Summary

Senior engineer and shift anchor for the NOC pod, and L2 for PulseAI infrastructure operations. Owns P1/P2 first response and bridge coordination, remediates GPU-server, node, fabric and OpenShift-network faults within Gruve's scope, executes firmware/driver and switch-configuration changes within maintenance windows, and manages hardware vendor escalations to closure within the restoration objectives.

Key Responsibilities
  • Act as shift anchor: own P1/P2 first response, open and run technical bridges until L3 engagement.
  • Execute complex changes (fabric node additions, firewall policy pushes) with senior oversight.
  • Remediate PulseAI infrastructure incidents within Gruve's scope: GPU-server firmware and driver faults (within the change process), OpenShift node and cluster-networking issues, front-end and RoCEv2 back-end fabric connectivity, switch configuration within the PulseAI fabric; diagnose storage capacity/health and node-hardware faults and escalated to the vendor with complete diagnostics.
  • Execute firmware and GPU driver updates, switch configuration changes and node lifecycle operations (MachineConfig, node drain/reboot) within agreed maintenance windows with the required customer notice; verify post-change health.
  • Manage hardware escalations: drive OEM, neocloud-provider and storage-vendor cases to closure, deliver daily status on Premium, and manage restoration-clock pauses correctly.
  • Drive preventive maintenance: firmware/BIOS/driver currency, EOL/EOS tracking, spare posture — for network devices and PulseAI GPU and control-plane nodes.
  • Own infrastructure observability for the NOC pod: GPU-cluster and fabric health dashboards, log queries and alert thresholds in Grafana; contribute to the environment validation checklist at onboarding (telemetry reachability per switch, storage throughput, access path).
  • Quality-gate tickets, handovers and runbook adherence across the pod; mentor juniors; coordinate with the SOC anchor on security-relevant network events; coordinate facility alerts with data-center operations.
Mandatory Qualifications
  • 4–6 years network operations in data-center environments.
  • Strong hands-on data-center switching/routing and next-generation firewall experience; incident bridge experience.
  • Working understanding of Kubernetes networking — CNI (Cilium), Services/Ingress and east-west flows — to troubleshoot fabric-to-GKE connectivity end to end.
  • Hands-on Red Hat OpenShift / Kubernetes operations in production — node lifecycle and MachineConfig, operators, cluster networking (OVN-Kubernetes/CNI), storage, oc/kubectl troubleshooting — plus Linux (RHEL) administration.
  • Working understanding of GPU-server operations: NVIDIA GPU Operator and driver stack, DCGM-class telemetry, firmware/driver update procedures, RDMA/RoCEv2 NIC health and common GPU failure modes.
  • Multi-vendor switch monitoring and troubleshooting (SNMP, syslog, streaming telemetry) across data-center platforms; out-of-band management proficiency; disciplined change execution.
  • Mentoring aptitude for the junior bench.
Preferred Qualifications
  • Professional-level data-center networking certification; automation exposure (Python/Ansible).
  • GKE cluster networking exposure (VPC-native, LoadBalancer/Gateway) and NetworkPolicy troubleshooting.
  • Red Hat OpenShift Administration certification (EX280) or RHCSA/RHCE; exposure to AI/ML workload scheduling on OpenShift.
  • GPU-cluster performance troubleshooting — RoCEv2/PFC/ECN tuning, ECMP polarisation, NCCL-visible latency/jitter — and hands-on with high-performance fabric telemetry, NVLink/NVSwitch topologies and GPU-node network profiling.
  • AI/HPC, neocloud or hyperscale data-center fabric exposure.
Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.
Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Consultant I
Network Consultant I

Gruve • Pune District

On-site
INR 600,000 - 1,000,000
Network Consultant II
Network Consultant II

gruve • Pune District

On-site
INR 600,000 - 1,200,000
Solution Architect – (AI Infrastructure & Hybrid Cloud)
Solution Architect – (AI Infrastructure & Hybrid Cloud)

Gruve • Pune District

On-site
INR 2,800,000 - 5,200,000
Network Consultant - L2
Network Consultant - L2

Gruve • Pune District

On-site
INR 1,200,000 - 1,800,000
Network Consultant - L2
Network Consultant - L2

Gruve • Pune District

On-site
INR 800,000 - 1,200,000
Solution Architect – (AI Infrastructure & Hybrid Cloud)
Solution Architect – (AI Infrastructure & Hybrid Cloud)

Gruve • Maharashtra

On-site
INR 1,800,000 - 4,000,000
Senior Security Consultant (Red Hat)
Senior Security Consultant (Red Hat)

Gruve • Maharashtra

Hybrid
INR 4,000,000 - 6,400,000
Junior Network Engineer
Junior Network Engineer

Gruve • India

On-site
INR 300,000 - 600,000
Culture of innovation
Diverse and inclusive workplace
Technical Support Engineer
Technical Support Engineer

Gruve • India

On-site
INR 400,000 - 800,000
Technical Account Manager
Technical Account Manager

Gruve • Mumbai

On-site
INR 1,800,000 - 3,200,000