Storage Engineer, Deployment & Support

Groq, Inc.

New York (NY)

On-site

USD 270,000 - 402,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Groq, Inc. is seeking a Storage Engineer, Deployment & Support to deploy, validate, and operate storage infrastructure for Groq's global AI infrastructure. You will own bring-up, validation, performance acceptance, troubleshooting, and production handoff across large-scale environments.

The role requires strong storage and networking fundamentals, Linux expertise, and the ability to coordinate deployments with multiple teams and vendors. Travel to global datacenters is expected as needed.

Qualifications

  • 4+ years of storage engineering, systems/datacenter roles in large-scale environments.

Responsibilities

  • Deploy and validate storage infrastructure across Groq datacenters.
  • Configure and bring up storage systems against approved designs and standards.
  • Drive deployment execution end-to-end with on-site technicians and vendors.
  • Develop and maintain deployment procedures, validation plans, and runbooks.
  • Perform storage acceptance and performance validation (throughput, IOPS, latency).
  • Monitor storage health, capacity, availability, and performance, preempting issues.

Skills

Storage engineering
Networking
Datacenter operations
Linux administration
Python
Bash
Ansible

Tools

VAST
WEKA
DDN
Lustre
Ceph

Job description

Mission: Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. Our Network Engineering & Datacenter teams design and operate the systems that provide the compute, storage, and networking capabilities behind Groq's rapidly growing AI infrastructure.

We are looking for a Storage Engineer, Deployment & Support to deploy, validate, and operate the storage infrastructure behind Groq's global AI infrastructure. As a Storage Engineer, Deployment & Support, you will take approved storage designs from implementation planning through production deployment and operational handoff. You will own storage system bring-up, configuration, validation, performance acceptance, troubleshooting, expansion, upgrades, and ongoing operational health across large-scale AI/HPC environments.

This role owns the underlying storage infrastructure lifecycle and partners closely with Platform Engineering to ensure storage capabilities are reliable, production-ready, and consumable. The role requires strong storage and networking fundamentals, structured troubleshooting, performance analysis, and the ability to drive deployments and operational work across multiple teams and vendors.

Responsibilities & opportunities in this role:
  • Deploy and validate storage infrastructure across new and existing Groq datacenters, including high-performance storage systems supporting AI/HPC workloads.
  • Configure and bring up storage systems against approved designs and implementation standards, including cluster initialization, storage nodes, interfaces, protocols, and supporting infrastructure.
  • Drive storage deployment execution end-to-end, directing and coordinating on-site datacenter technicians and vendors through physical rack/stack, cabling, and installation while owning configuration, validation, troubleshooting, performance acceptance, and production handoff.
  • Develop and maintain implementation procedures, validation plans, deployment checklists, system inventories, as-built documentation, and operational runbooks.
  • Perform storage acceptance and performance validation, including throughput, IOPS, latency, capacity, data distribution, cluster health, and recovery behavior.
  • Monitor and maintain storage health, capacity, availability, and performance, identifying and resolving issues before they impact production workloads.
  • Troubleshoot storage and hardware failures across storage nodes, drives, network paths, protocols, and supporting infrastructure, and perform root cause analysis for deployment and production issues.
  • Execute and coordinate storage expansions, upgrades, hardware refreshes, and other lifecycle activities while minimizing production impact.
  • Manage hardware readiness for deployments and expansions, including rack/stack coordination, inventory and asset record updates, RMA coordination, spares, and vendor shipments.
  • Coordinate deployment audits and execute established QA/QC processes to ensure storage infrastructure meets Groq standards.
  • Partner closely with storage design, network, compute and platform teams to identify and resolve blockers throughout deployment and operations.
  • Provide operational support during and after deployments, including maintenance activities, incidents, remediation, break-fix events, and vendor escalation when required.
  • Develop tools and scripts that improve storage deployment, configuration, validation, testing, health checks, and repeatable operational workflows.
  • Capture lessons-learned from deployments and incidents and contribute to global implementation standards and deployment playbooks.
Ideal candidates have/are:
  • 4+ years of experience in storage engineering, systems engineering, network engineering, datacenter operations, or related infrastructure roles, with hands-on experience deploying, operating, or troubleshooting storage systems.
  • Strong hands-on experience operating and troubleshooting large-scale distributed or high-performance storage environments.
  • Experience with modern storage platforms and technologies such as VAST, WEKA, DDN, Lustre, Ceph, or comparable distributed storage solutions.
  • Strong Linux systems administration and troubleshooting skills in production environments.
  • Strong understanding of storage protocols and data paths, including NFS, NFS over RDMA, NVMe-oF, and high-throughput Ethernet or InfiniBand connectivity.
  • Strong understanding of storage performance concepts, including throughput, IOPS, latency, capacity, and utilization.
  • Understanding of RDMA, RoCE, InfiniBand, and other high-performance networking concepts used in AI/HPC storage environments.
  • Ability to troubleshoot systematically across storage, Linux, network, and hardware layers and drive issues through root cause and resolution.
  • Experience developing deployment or operational tooling using Python, Bash, Ansible, APIs, or similar technologies.
  • Strong ownership, communication, and time-management skills, with the ability to manage multiple deployments and operational priorities under demanding timelines.
  • Comfortable operating in fast-moving environments where deployment plans and requirements can evolve quickly.
  • Ability to travel to global datacenter locations for deployments, maintenance activities, and other site-specific needs.
Why Join Us:
  • Purposeful Hiring: You’re not here by accident, and neither is anyone else. Every teammate is handpicked with intention because who we build with matters.
  • Builders Wanted: You’re not just riding the rocket ship, you’re building it. Your work directly shapes the trajectory of our company.
  • Mission-Driven Work: We’re here to make a real impact. Our mission fuels everything we do.
  • Tackling Hard Problems: If easy isn’t your thing, you’re in the right place. We solve some of the most complex and exciting challenges in our space.
  • Excellence Is The Standard: High performance isn’t just encouraged, it’s the baseline. And it’s contagious.
Compensation

Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $270,400–$401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Storage Engineer, Deployment & Support
Storage Engineer, Deployment & Support

Groq, Inc. • Town of Texas (WI)

On-site
USD 270,000 - 402,000
Storage Engineer, Deployment & Support
Storage Engineer, Deployment & Support

Groq, Inc. • San Francisco (CA)

On-site
USD 270,000 - 402,000
Long-Term Incentive (LTI) Program
Employee benefits
Storage Engineer, Deployment & Support
Storage Engineer, Deployment & Support

Groq, Inc. • Dallas (TX), Northern (KY)

Hybrid
USD 270,000 - 402,000
Sr./Staff Forward Deployed Engineer
Sr./Staff Forward Deployed Engineer

Groq, Inc. • San Francisco (CA)

Hybrid
USD 270,000 - 402,000
Sr./Staff Forward Deployed Engineer
Sr./Staff Forward Deployed Engineer

Groq, Inc. • Town of Texas (WI)

Hybrid
USD 270,000 - 402,000
Sr./Staff Network Design Engineer, AI/HPC
Sr./Staff Network Design Engineer, AI/HPC

Groq, Inc. • Town of Texas (WI)

Hybrid
USD 270,000 - 402,000
Remote work flexibility
Bonus and LTI by Groq
Total Cash compensation
Sr./Staff Network Design Engineer, AI/HPC
Sr./Staff Network Design Engineer, AI/HPC

Groq, Inc. • San Francisco (CA)

Remote
USD 270,000 - 402,000
Purposeful Hiring
Builders Wanted
Mission-Driven Work
+1
Sr./Staff Network Design Engineer, AI/HPC
Sr./Staff Network Design Engineer, AI/HPC

Groq, Inc. • New York (NY)

Hybrid
USD 270,000 - 402,000
Sr./Staff Forward Deployed Engineer
Sr./Staff Forward Deployed Engineer

Groq, Inc. • New York (NY)

Remote
USD 270,000 - 402,000
Data Center Sourcing Manager
Data Center Sourcing Manager

Groq, Inc. • United States

Remote
USD 154,000 - 209,000