Sr./Staff Forward Deployed Engineer

Groq, Inc.

San Francisco (CA)

Hybrid

USD 270,000 - 402,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Groq, Inc. is seeking a Senior/Staff Forward Deployed Engineer focused on AI infrastructure to bring accelerator and infrastructure technology into production.

You will work across GroqCloud, GroqMetal, GPU and LPX infrastructure, and networking, storage, and observability layers to deploy demanding AI workloads. Location options include Dallas, San Francisco, or New York City with remote flexibility as local offices ramp up.

Qualifications

  • 4+ years building, deploying, operating, or troubleshooting cloud/AI infrastructure.
  • Strong Linux and distributed-systems fundamentals; practical Kubernetes or Slurm experience.

Responsibilities

  • Own technical execution across complex customer engagements from discovery to production readiness and handoff.
  • Translate ambiguous requirements into concrete architectures, runbooks and actions.
  • Troubleshoot across Linux, bare-metal infra, Kubernetes, networking, storage, and Groq platform integrations to bring environments online.
  • Support large-scale GPU/LPX deployments including health checks, benchmarking, and production-readiness evidence.
  • Lead PoCs, architecture reviews, demos, and validations with clear tradeoffs and risks for engineering and business stakeholders.
  • Partner with Networking and Security on private connectivity, routing, access controls, and security diligence.
  • Embed with Platform, Cloud, Infrastructure, or Operations teams for priority deployments and reusable tooling.
  • Build reusable tools, reference architectures, and documentation to improve future deployments.

Skills

Linux
Kubernetes
Slurm
Python
Go
Bash
Automation
Diagnostics

Tools

CUDA
NCCL
NVLink/NVSwitch
DCGM
GPU Operator
InfiniBand
RoCE
Kubernetes
Slurm

Job description

As a Sr./Staff Forward Deployed Engineer focused on AI Infrastructure at Groq, you will work at the frontier of large-scale AI systems, taking complex customer infrastructure programs from requirements to working production environments. You’ll help bring some of the newest accelerator and infrastructure technologies into production, spanning next-generation NVIDIA GPU systems alongside Groq’s purpose-built inference platform.


You will operate across GroqCloud, GroqMetal (Groq’s infrastructure platform), GPU and LPX infrastructure, and the networking, storage, orchestration, observability, and workload layers around them. This is an opportunity to work on infrastructure where the playbooks are still being written: bringing up new systems, solving problems that emerge only at scale, and helping customers deploy demanding AI workloads on platforms at the leading edge of the market. You will work directly with customers while partnering closely with Commercial, Field Engineering, Platform and Cloud Engineering, Networking, Data Center Operations, Security, and Support teams.


This is a deeply hands-on individual-contributor role. You will work directly in systems, write code and automation, troubleshoot across layers of the stack, and turn ambiguous customer requirements into deployed and validated solutions. The work you do in the field will also shape what comes next: turning hard-won lessons into reusable tooling, deployment patterns, reference architectures, and improvements to the Groq platform for the customers that follow.


Location: This role will be based in one of our three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired for this role must be based in one of these three areas. You’ll have the flexibility to work remotely while we establish our local Groq office, with the expectation that this role will transition to onsite once the office opens.


Responsibilities & Opportunities in This Role


  • Own technical execution across complex customer engagements, from discovery and architecture through PoCs, demos, deployment, cluster bring-up, validation, acceptance, production readiness, and operational handoff.

  • Translate incomplete or ambiguous customer requirements into practical architectures, implementation plans, test criteria, runbooks, and concrete engineering actions.

  • Work hands-on across Linux, bare-metal infrastructure, Kubernetes and Slurm, networking, storage, observability, automation, and Groq platform integrations to bring customer environments online and resolve issues.

  • Support large-scale GPU and LPX deployments, including infrastructure bring-up, cluster health and performance validation, workload testing, benchmarking, failure isolation, and production-readiness evidence.

  • Understand customer AI workloads well enough to reason about training and inference behavior, concurrency, throughput, latency, data movement, caching, scheduling, and infrastructure bottlenecks.

  • Lead technical portions of customer discovery, architecture reviews, demonstrations, and proofs of concept, clearly explaining design choices, tradeoffs, performance results, and risks to both engineering and business stakeholders.

  • Partner with Networking and Security teams on customer requirements such as private connectivity and peering, routing, ingress and egress, load balancing, network policy, access controls, security architecture reviews, and enterprise security diligence.

  • Troubleshoot production and pre-production issues that cross organizational or technical boundaries, drive them to resolution, and coordinate the right internal experts without losing end-to-end ownership.

  • Embed with Platform, Cloud, Infrastructure, or Operations teams when priority customer deployments expose gaps that require concentrated engineering execution, automation, or integration work.

  • Build reusable tools, automation, reference architectures, test suites, deployment patterns, documentation, and lessons learned so that customer-specific engineering makes the platform better for the next deployment.

  • Bring structured customer feedback and field evidence back to Product and Engineering, identifying recurring gaps and helping turn one-off solutions into repeatable platform capabilities.


Ideal Candidates Have/Are


  • 4+ years of hands-on experience building, deploying, operating, or troubleshooting cloud infrastructure, AI infrastructure, HPC systems, large-scale platforms, or similarly demanding production environments.

  • Strong Linux and distributed-systems fundamentals, with practical experience in Kubernetes, Slurm, bare-metal environments, or comparable infrastructure platforms.

  • Meaningful technical depth in at least one area such as GPU or accelerator systems, networking, storage, orchestration/platform engineering, or infrastructure reliability, with enough breadth to troubleshoot across adjacent layers.

  • Working knowledge of AI training and inference workloads and how workload characteristics affect compute, networking, storage, scheduling, latency, and throughput.

  • Strong Python, Go, Bash, or equivalent scripting/programming skills for diagnostics, automation, deployment tooling, testing, or integrations.

  • A track record of personally debugging and delivering systems rather than operating only at the architecture, project-management, or escalation level.

  • Ability to break ambiguous problems into concrete technical actions and drive issues to resolution when responsibility spans multiple teams.

  • Strong written and verbal communication skills, including the ability to gather requirements from customer engineers, explain technical tradeoffs clearly, and document work so that others can reproduce it.

  • Comfortable operating in a fast-moving environment where customer requirements, platform capabilities, and implementation details can evolve in parallel.


Preferred Qualifications


  • Experience at a neocloud, hyperscaler, AI infrastructure provider, HPC environment, frontier AI company, or other organization operating large-scale accelerator infrastructure.

  • Hands-on experience with NVIDIA GPU infrastructure and technologies such as CUDA, NCCL, NVLink/NVSwitch, DCGM, GPU Operator, InfiniBand, RoCE, Kubernetes, or Slurm.

  • Experience bringing up, qualifying, or operating multi-node GPU clusters, including health checks, burn-in or stress testing, collective-communication testing, performance benchmarking, and acceptance criteria.

  • Familiarity with high-performance storage systems such as VAST, Weka, Lustre, Ceph, or similar technologies and the data-access patterns of distributed AI workloads.

  • Experience with infrastructure automation and lifecycle tooling such as Terraform, Ansible, CI/CD, BMC/Redfish, PXE/iPXE, or related systems.

  • Prior solutions engineering, sales engineering, solutions architecture, or technical pre-sales experience in cloud, networking, security, AI infrastructure, or data center systems.

  • Customer-facing networking experience including private interconnects and peering, BGP and routing, load balancing, Kubernetes/Cilium network policy, and north-south and east-west traffic design.

  • Customer-facing security experience including access-control architecture, network isolation, enterprise security reviews, and SOC 2 / ISO 27001-style diligence or questionnaires.

  • Experience defining or executing technical PoCs, reference architectures, cluster acceptance tests, performance benchmarks, migration plans, or production-readiness criteria.


Why Join Us:


  • Purposeful Hiring: You’re not here by accident, and neither is anyone else. Every teammate is handpicked with intention because who we build with matters.

  • Builders Wanted: You’re not just riding the rocket ship, you’re building it. Your work directly shapes the trajectory of our company.

  • Mission-Driven Work: We’re here to make a real impact. Our mission fuels everything we do.

  • Tackling Hard Problems: If easy isn’t your thing, you’re in the right place. We solve some of the most complex and exciting challenges in our space.

  • Excellence Is The Standard: High performance isn’t just encouraged, it’s the baseline. And it’s contagious.


Compensation

Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $270,400-$401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr./Staff Forward Deployed Engineer
Sr./Staff Forward Deployed Engineer

Groq, Inc. • Town of Texas (WI)

Hybrid
USD 270,000 - 402,000
Sr./Staff Forward Deployed Engineer
Sr./Staff Forward Deployed Engineer

Groq, Inc. • New York (NY)

Remote
USD 270,000 - 402,000
Technical Product Marketing Manager
Technical Product Marketing Manager

Groq, Inc. • United States

Hybrid
USD 154,000 - 209,000
Sr./Staff Network Design Engineer, AI/HPC
Sr./Staff Network Design Engineer, AI/HPC

Groq, Inc. • San Francisco (CA)

Remote
USD 270,000 - 402,000
Purposeful Hiring
Builders Wanted
Mission-Driven Work
+1
Sr./Staff Network Design Engineer, AI/HPC
Sr./Staff Network Design Engineer, AI/HPC

Groq, Inc. • Town of Texas (WI)

Hybrid
USD 270,000 - 402,000
Remote work flexibility
Bonus and LTI by Groq
Total Cash compensation
Sr./Staff Network Design Engineer, AI/HPC
Sr./Staff Network Design Engineer, AI/HPC

Groq, Inc. • New York (NY)

Hybrid
USD 270,000 - 402,000
Technical Product Marketing Manager
Technical Product Marketing Manager

Groq, Inc. • San Francisco (CA)

On-site
USD 154,000 - 209,000
Long-Term Incentive (LTI)
Storage Engineer, Deployment & Support
Storage Engineer, Deployment & Support

Groq, Inc. • New York (NY)

On-site
USD 270,000 - 402,000
Storage Engineer, Deployment & Support
Storage Engineer, Deployment & Support

Groq, Inc. • San Francisco (CA)

On-site
USD 270,000 - 402,000
Long-Term Incentive (LTI) Program
Employee benefits
Storage Engineer, Deployment & Support
Storage Engineer, Deployment & Support

Groq, Inc. • Town of Texas (WI)

On-site
USD 270,000 - 402,000