About Gruve
Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.
About Gruve
Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.
Position Summary
Senior engineer and shift anchor for the NOC pod, and L2 for PulseAI infrastructure operations. Owns P1/P2 first response and bridge coordination, remediates GPU-server, node, fabric and OpenShift-network faults within Gruve's scope, executes firmware/driver and switch-configuration changes within maintenance windows, and manages hardware vendor escalations to closure within the restoration objectives.
Key Responsibilities
- Act as shift anchor: own P1/P2 first response, open and run technical bridges until L3 engagement.
- Execute complex changes (fabric node additions, firewall policy pushes) with senior oversight.
- Remediate PulseAI infrastructure incidents within Gruve's scope: GPU-server firmware and driver faults (within the change process), OpenShift node and cluster-networking issues, front-end and RoCEv2 back-end fabric connectivity, switch configuration within the PulseAI fabric; diagnose storage capacity/health and node-hardware faults and escalated to the vendor with complete diagnostics.
- Execute firmware and GPU driver updates, switch configuration changes and node lifecycle operations (MachineConfig, node drain/reboot) within agreed maintenance windows with the required customer notice; verify post-change health.
- Manage hardware escalations: drive OEM, neocloud-provider and storage-vendor cases to closure, deliver daily status on Premium, and manage restoration-clock pauses correctly.
- Drive preventive maintenance: firmware/BIOS/driver currency, EOL/EOS tracking, spare posture — for network devices and PulseAI GPU and control-plane nodes.
- Own infrastructure observability for the NOC pod: GPU-cluster and fabric health dashboards, log queries and alert thresholds in Grafana; contribute to the environment validation checklist at onboarding (telemetry reachability per switch, storage throughput, access path).
- Quality-gate tickets, handovers and runbook adherence across the pod; mentor juniors; coordinate with the SOC anchor on security-relevant network events; coordinate facility alerts with data-center operations.
Mandatory Qualifications
- 4–6 years network operations in data-center environments.
- Strong hands-on data-center switching/routing and next-generation firewall experience; incident bridge experience.
- Working understanding of Kubernetes networking — CNI (Cilium), Services/Ingress and east-west flows — to troubleshoot fabric-to-GKE connectivity end to end.
- Hands-on Red Hat OpenShift / Kubernetes operations in production — node lifecycle and MachineConfig, operators, cluster networking (OVN-Kubernetes/CNI), storage, oc/kubectl troubleshooting — plus Linux (RHEL) administration.
- Working understanding of GPU-server operations: NVIDIA GPU Operator and driver stack, DCGM-class telemetry, firmware/driver update procedures, RDMA/RoCEv2 NIC health and common GPU failure modes.
- Multi-vendor switch monitoring and troubleshooting (SNMP, syslog, streaming telemetry) across data-center platforms; out-of-band management proficiency; disciplined change execution.
- Mentoring aptitude for the junior bench.
Preferred Qualifications
- Professional-level data-center networking certification; automation exposure (Python/Ansible).
- GKE cluster networking exposure (VPC-native, LoadBalancer/Gateway) and NetworkPolicy troubleshooting.
- Red Hat OpenShift Administration certification (EX280) or RHCSA/RHCE; exposure to AI/ML workload scheduling on OpenShift.
- GPU-cluster performance troubleshooting — RoCEv2/PFC/ECN tuning, ECMP polarisation, NCCL-visible latency/jitter — and hands-on with high-performance fabric telemetry, NVLink/NVSwitch topologies and GPU-node network profiling.
- AI/HPC, neocloud or hyperscale data-center fabric exposure.
Why Gruve
At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.
Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.