Security Operations Consultant

Gruve

Maharashtra

On-site

INR 1,200,000 - 2,000,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Gruve is seeking a senior technical lead to oversee SOC, NOC, and PulseAI Managed Services. You will drive major incident response, own detection content, and lead OpenShift lifecycle operations across client environments.

You will coordinate with Red Hat, NVIDIA, and OEMs, mentor L1/L2 teams, and ensure alignment with SLAs. The role requires a strong automation mindset and deep security operations experience in production Kubernetes/OpenShift environments.

Qualifications

  • 8–11 years with senior/lead experience in SOC, NOC or 24×7 platform operations.
  • Incident command capability; deep SIEM content and query skills.
  • Strong Kubernetes/GKE security operations depth — onboarding kube-audit and workload telemetry.
  • Deep Red Hat OpenShift / Kubernetes operations experience in production.
  • GPU-cluster operations experience: NVIDIA GPU Operator/driver lifecycle and AI Enterprise components.
  • Experience operating against contractual SLAs and managing vendor escalations.

Responsibilities

  • Lead technical response on Severity 1/2 incidents across security, network and PulseAI platform.
  • Own root cause analysis for PulseAI and OpenShift incidents; reproduce and characterize defects and route them to engineering.
  • Own the PulseAI operations practice: standards, diagnostics, alerts, and monitoring matrix.
  • Lead OpenShift lifecycle operations for PulseAI customers — upgrades, backups, and remediation.
  • Act as vendor engineering liaison: Red Hat, NVIDIA, OEMs for platform defects and escalations.
  • Lead customer onboarding: deploy OpenShift and PulseAI within 14 days and validate tenancy and connectivity.
  • Own the detection-content backlog and log-pipeline health; manage escalation paths.
  • Approve high-risk changes and drive post-incident reviews.

Skills

Incident command
SIEM content
Kubernetes security
OpenShift operations
Grafana observability
Vendor escalations

Tools

OpenShift
Kubernetes
GPU Operator
NVIDIA AI Enterprise

Job description

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position Summary

Senior technical lead across SOC, NOC and PulseAI Managed Services. Runs major-incident response, owns the detection-content and tuning program, and is L3 for PulseAI operations: owns root cause analysis, works platform defects with the Gruve PulseAI engineering team, leads OpenShift lifecycle operations and emergency remediation, acts as vendor engineering liaison (Red Hat, NVIDIA, OEMs), and deputises for the Security Operation Manager.

Key Responsibilities
  • Lead technical response on Severity 1/2 incidents across security, network and PulseAI platform until management command engages; own the SLA status cadence (hourly / every 30 minutes on Severity 1) to customer authorised contacts.
  • Own root cause analysis for PulseAI and OpenShift incidents; reproduce and characterise platform defects and route them to Gruve PulseAI product engineering; own the corrective-action backlog and problem management to eliminate recurring incidents.
  • Own the PulseAI operations practice: severity classification standards, diagnostic runbooks for platform/OpenShift/GPU/switch/storage layers, Grafana observability standards, alert thresholds for the SLA appendix, and the monitor / remediate /  escalate matrix as applied per customer.
  • Lead OpenShift lifecycle operations for PulseAI customers — y-stream upgrades on customer approval within the agreed window of the Red Hat release, operator and node‑pool changes, GPU driver/firmware currency, emergency vulnerability remediation — enforcing the pre-change backup gate.
  • Act as vendor engineering liaison: Red Hat for OpenShift product defects, NVIDIA for GPU Operator/driver/NVIDIA AI Enterprise issues, OEM/neocloud/storage vendors for hardware — including the third‑party platform variant (Rafay, vCluster, vNode, NVIDIA Run:ai, Red Hat OpenShift AI) where restoration is best-effort with committed vendor escalation.
  • Lead customer onboarding technically: deploy and validate OpenShift and PulseAI within the 14‑day Ready‑for‑Install window, configure identity provider federation and initial tenancy structure, onboard the environment to monitoring via the agreed connectivity pattern (outbound collector, site-to‑site VPN or jump host with just‑in‑time elevation), and complete the countersigned environment validation checklist.
  • Own the detection-content backlog and tuning program; own log‑pipeline integration health with escalation into engineering.
  • Approve and execute high‑risk changes; cross‑train the L1/L2 bench across towers; drive shift‑quality audits and post‑incident reviews.
Mandatory Qualifications
  • 8–11 years with prior senior/lead experience in SOC, NOC or 24×7 platform operations and genuine cross‑domain fluency.
  • Incident command capability; deep SIEM content and query skills.
  • Strong Kubernetes/GKE security operations depth — onboarding kube‑audit and workload telemetry, building detection content for container attack paths (MITRE ATT&CK for Containers), and tuning Cilium/Hubble‑based use cases.
  • Deep Red Hat OpenShift / Kubernetes operations experience in production — multi‑node cluster administration, operators, y/z-stream upgrades, RBAC, storage and networking, observability stack (Grafana, metrics, logs, alerting), backup/restore of cluster and platform state — with a track record of platform diagnostics and RCA.
  • GPU‑cluster operations experience: NVIDIA GPU Operator/driver lifecycle, NVIDIA AI Enterprise components (NIM), DCGM‑class telemetry, node health and capacity management for AI inference workloads on RTX PRO 6000 / HGX B300‑class servers or equivalent.
  • Experience operating against contractual SLAs (acknowledgement, restoration, availability, service credits) and running vendor engineering escalations through to fix.
  • Advanced network troubleshooting; automation mindset.
Preferred Qualifications
  • Advanced incident‑handling / intrusion‑analysis certification, or expert‑level networking certification.
  • CKS or equivalent; experience running K8s posture/runtime security tooling in production.
  • Red Hat certifications (EX280 / EX380 / RHCE); exposure to Rafay, vCluster, NVIDIA Run:ai or Red Hat OpenShift AI; MLOps or AI‑platform operations (model‑serving endpoints, GPU scheduling, quota governance).
  • Exposure to HashiCorp Vault, SAML 2.0 SSO federation, and CSI/NFS storage for model artefacts.
  • MSSP/managed‑services background; AI‑SOC tooling exposure.
  • Correlating GPU‑platform performance anomalies (utilization, thermal, fabric saturation) with security events to separate abuse, crypto‑mining or misconfiguration from genuine workload load.
Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Security Operations Manager
Security Operations Manager

Gruve • Maharashtra

On-site
INR 3,500,000 - 7,000,000
Security Analyst I
Security Analyst I

Gruve • Maharashtra

On-site
INR 900,000 - 1,300,000
Network Consultant II
Network Consultant II

Gruve • Maharashtra

On-site
INR 1,800,000 - 3,000,000
Senior Security Consultant (Red Hat)
Senior Security Consultant (Red Hat)

Gruve • Maharashtra

Hybrid
INR 4,000,000 - 6,400,000
Solution Architect – (AI Infrastructure & Hybrid Cloud)
Solution Architect – (AI Infrastructure & Hybrid Cloud)

Gruve • Pune District

On-site
INR 2,800,000 - 5,200,000
Network Consultant I
Network Consultant I

Gruve • Maharashtra

On-site
INR 1,200,000 - 1,800,000
Security Analyst I
Security Analyst I

gruve • Pune District

On-site
INR 600,000 - 900,000
Senior Software Engineer (Java Fullstack)
Senior Software Engineer (Java Fullstack)

Gruve • Maharashtra

Hybrid
INR 1,800,000 - 2,400,000
Senior Software Engineer (Java Fullstack) Pune, Maharashtra, India
Senior Software Engineer (Java Fullstack) Pune, Maharashtra, India

Gruve Inc. • Pune District

On-site
INR 3,000,000 - 5,000,000
Senior Pre-Sales Consultant - CyberSecurity
Senior Pre-Sales Consultant - CyberSecurity

Gruve • Mumbai

On-site
INR 1,500,000 - 2,500,000
Dynamic work environment
Inclusive workplace culture
Continuous learning opportunities