Gruve
Gruve

Security Analyst II

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position summary:

Senior analyst and shift anchor for the SOC pod, and L2 for PulseAI Managed Services. Owns triage quality across the rotation, handles high-severity security incidents through to handoff, and remediates PulseAI and OpenShift incidents to the Standard/Premium restoration objectives (Severity 1 in 4/2 hours), executes changes and patching within maintenance windows, and manages escalations to L3 platform engineering and to hardware vendors.

Key Roles & Responsibilities:

  • Act as shift senior: final quality gate on investigations and escalations; handle P1/P2 security alerts end-to-end to L3 handoff including scoping and evidence preservation.
  • Drive shift-level metrics: SLA adherence, false-positive rate, reopen rate.
  • Remediate PulseAI platform incidents within the authority matrix — control-plane and platform-service recovery, authentication/SSO and RBAC faults, tenancy and quota-enforcement failures, endpoint deployment and model-serving failures, observability outages — and restore PulseAI configuration state (organisation/department/project hierarchy, quota allocations, users and roles) from Gruve backups when required.
  • Remediate OpenShift cluster incidents — node NotReady, scheduling and capacity, operator degradation, cluster networking, storage/PVC faults, image registry and ingress — using oc/kubectl and cluster diagnostics; hand structured RCAs to L3.
  • Execute PulseAI patch releases and OpenShift z-stream patches in the monthly maintenance window, firmware and GPU driver updates as required, and emergency security remediation within the tier window (72 hours Premium / 5 business days Standard); enforce the pre-change gate that no cluster upgrade proceeds without the customer's written confirmation of a verified backup.
  • Manage hardware escalations opened by L1: drive OEM/neocloud/storage vendor cases to closure, keep the customer informed at the SLA cadence, and pause/resume restoration clocks correctly when waiting on the customer, a vendor or a change approval.
  • Own platform observability hygiene: Grafana dashboards, alert thresholds recorded in the SLA appendix (performance degradation, capacity), and runbooks for recurring platform faults.
  • Coach and quality-review the L1 analysts; own runbook accuracy for the SOC pod; coordinate cross-tower with the NOC anchor on ambiguous events.

Mandatory Qualifications:

  • BE/BTech (CS/IT/E&TC) or equivalent.
  • 4–6 years SOC or 24×7 platform-operations experience with demonstrable incident handling and remediation ownership.
  • Strong multi-source triage across cloud, network and identity telemetry.
  • Hands-on triage of Kubernetes/container security alerts — kube-audit events, workload anomalies and pod-level network flows (Cilium/Hubble) — ideally on GKE.
  • Hands-on Red Hat OpenShift / Kubernetes administration in production — troubleshooting pods, nodes, operators, storage and cluster networking with oc/kubectl; applying z-stream patches and operator updates within change control; reading control-plane and workload logs in a metrics/logging observability stack (Grafana).
  • Working understanding of GPU-node operations — NVIDIA GPU Operator and driver stack, DCGM-class telemetry, GPU firmware/driver update process, common GPU health failure modes — and of the monitor / remediate / escalate boundaries between platform, hardware and customer workload.
  • Audit-grade documentation; ability to run a shift independently and calmly under Severity 1 pressure and to communicate status to customer authorised contacts.
  • Mentoring aptitude — this role carries the shift's junior bench.

Preferred Qualifications:

  • Incident-handling / forensics certification-level knowledge (e.g., GCIH, CHFI or equivalent); scripting for enrichment/automation.
  • CKA/CKS or KCSA; exposure to admission controls and image/runtime security tooling.
  • Red Hat OpenShift Administration certification (EX280) or RHCSA; exposure to NVIDIA AI Enterprise (NIM microservices), model-serving endpoints and GPU workload scheduling on OpenShift.
  • Exposure to HashiCorp Vault, SAML 2.0 SSO / RBAC troubleshooting, and n8n or similar platform applications.
  • Prior GPU-cloud, neocloud, hyperscale or data-center customer exposure.

Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Gruve is a software services startup that empowers enterprises to become AI-driven organizations. We provide specialized expertise in cybersecurity, customer experience, and cloud infrastructure, utilizing advanced technologies such as Large Language Models to enable smarter, data-informed decision-making for our clients.

Founded
Founded 2024
Employees
201-500 employees
Industry
information technology and services
Funding stage
Series A
Total raised
$88M raised
Last funding
Raised February 2026
View company profile