Senior Software Engineering Manager – KV Cache Platform

DDN

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DDN is seeking a Senior Software Engineering Manager to lead the KV Cache Platform—an advanced distributed memory and storage platform designed to accelerate large-scale LLM inference across GPU clusters.

You will guide geographically distributed engineering teams, shape the technical vision, and collaborate with Product Management, Sales, Customer Engineering, NVIDIA, and executives to deliver AI infrastructure that meets customer needs and aligns with DDN's long-term strategy.

Qualifications

  • 15+ years of experience building distributed systems, cloud infrastructure, storage platforms, or AI infrastructure software.
  • 7+ years leading high-performing software engineering organizations, including geographically distributed teams.
  • Proven experience delivering large-scale distributed infrastructure products from architecture through production deployment.
  • Strong background in distributed systems, Linux, networking, performance engineering, and cloud-native architectures.
  • Hands‑on programming experience with Go and Python; experience with C/C++ is a plus.
  • Demonstrated ability to lead cross-functional initiatives and influence technical direction across multiple organizations.
  • Experience building AI infrastructure, LLM serving platforms, distributed caching systems, or high-performance storage solutions.
  • Experience with technologies such as NVIDIA Dynamo, TensorRT-LLM, Triton, RDMA, GPUDirect Storage, BlueField DPUs, Kubernetes, or related AI infrastructure.
  • Background in HPC, distributed storage, networking, or enterprise infrastructure software.
  • Experience working directly with strategic customers, technology partners, OEMs, or hyperscalers to deliver enterprise AI solutions.

Responsibilities

  • Lead, mentor, and grow a geographically distributed team of software engineers and technical leaders, fostering a culture of technical excellence, innovation, ownership, and collaboration.
  • Define and execute the technical strategy and roadmap for the KV Cache Platform, ensuring scalability, reliability, security, and operational excellence.
  • Drive the architecture, development, and delivery of distributed systems supporting AI inference, GPU memory optimization, distributed caching, RDMA networking, GPUDirect Storage, NVIDIA BlueField DPUs, and emerging AI infrastructure technologies.
  • Partner closely with Product Management, Sales, Customer Engineering, NVIDIA, and strategic technology partners to prioritize customer requirements, drive proof-of-concepts (POCs), influence product direction, and successfully deliver customer deployments.
  • Own day-to-day engineering execution, including feature development, release planning, bug triage, production issues, customer escalations, and cross-functional execution to ensure timely, high-quality software delivery.
  • Establish engineering best practices for software quality, observability, automation, performance, testing, and production readiness.
  • Collaborate across engineering, infrastructure, and hardware teams to deliver scalable, production-ready AI infrastructure while developing future engineering leaders and driving continuous improvement.

Skills

Distributed systems
Leadership
Go
Python
C/C++
Linux
Networking
Performance engineering
Cloud-native architectures
HPC
AI infrastructure
Kubernetes
RDMA
GPUDirect Storage
TensorRT-LLM
Triton
NVIDIA BlueField DPUs

Tools

Go
Python
C/C++
Kubernetes
RDMA
GPUDirect Storage
TensorRT-LLM
Triton

Job description

DDN is seeking a Senior Software Engineering Manager to lead the engineering organization responsible for our KV Cache Platform—a distributed memory and storage platform that accelerates large-scale LLM inference across GPU clusters.

In this role, you will lead geographically distributed engineering teams responsible for building highly scalable, low-latency distributed systems that power AI inference. You will define the technical vision and execution strategy for the platform while partnering closely with Product Management, Sales, Customer Engineering, NVIDIA, and executive leadership to deliver innovative AI infrastructure that meets customer needs and supports DDN's long-term product strategy.

This is a highly visible leadership role with responsibility for engineering execution, customer success, roadmap delivery, and building a world-class engineering organization.

Responsibilities
  • Lead, mentor, and grow a geographically distributed team of software engineers and technical leaders, fostering a culture of technical excellence, innovation, ownership, and collaboration.

  • Define and execute the technical strategy and roadmap for the KV Cache Platform, ensuring scalability, reliability, security, and operational excellence.

  • Drive the architecture, development, and delivery of distributed systems supporting AI inference, GPU memory optimization, distributed caching, RDMA networking, GPUDirect Storage, NVIDIA BlueField DPUs, and emerging AI infrastructure technologies.

  • Partner closely with Product Management, Sales, Customer Engineering, NVIDIA, and strategic technology partners to prioritize customer requirements, drive proof-of-concepts (POCs), influence product direction, and successfully deliver customer deployments.

  • Own day-to-day engineering execution, including feature development, release planning, bug triage, production issues, customer escalations, and cross-functional execution to ensure timely, high-quality software delivery.

  • Establish engineering best practices for software quality, observability, automation, performance, testing, and production readiness.

  • Collaborate across engineering, infrastructure, and hardware teams to deliver scalable, production-ready AI infrastructure while developing future engineering leaders and driving continuous improvement.

Qualifications
Required
  • 15+ years of experience building distributed systems, cloud infrastructure, storage platforms, or AI infrastructure software.

  • 7+ years leading high-performing software engineering organizations, including geographically distributed teams.

  • Proven experience delivering large-scale distributed infrastructure products from architecture through production deployment.

  • Strong background in distributed systems, Linux, networking, performance engineering, and cloud-native architectures.

  • Hands‑on programming experience with Go and Python; experience with C/C++ is a plus.

  • Demonstrated ability to lead cross-functional initiatives and influence technical direction across multiple organizations.

  • Experience building AI infrastructure, LLM serving platforms, distributed caching systems, or high-performance storage solutions.

  • Experience with technologies such as NVIDIA Dynamo, TensorRT-LLM, Triton, RDMA, GPUDirect Storage, BlueField DPUs, Kubernetes, or related AI infrastructure.

  • Background in HPC, distributed storage, networking, or enterprise infrastructure software.

  • Experience working directly with strategic customers, technology partners, OEMs, or hyperscalers to deliver enterprise AI solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineering Manager – KV Cache Platform
Senior Software Engineering Manager – KV Cache Platform

Ddn • Sacramento (CA)

On-site
USD 190,000 - 260,000
Head of AI Inference Infrastructure (KV Cache Platform)
Head of AI Inference Infrastructure (KV Cache Platform)

Ddn • Sacramento (CA)

On-site
USD 190,000 - 260,000
Senior Engineering Leader – AI Inference KV Cache
Senior Engineering Leader – AI Inference KV Cache

DDN • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior Staff Engineer - AI Data Path
Senior Staff Engineer - AI Data Path

Ddn • Sacramento (CA)

On-site
USD 180,000 - 240,000
Senior Staff Engineer - AI Data Path
Senior Staff Engineer - AI Data Path

DDN • California (MO)

Hybrid
USD 180,000 - 250,000
Senior Staff Engineer - AI Data Path
Senior Staff Engineer - AI Data Path

Data Direct Networks • California (MO)

On-site
USD 180,000 - 240,000
Highly Competitive Vacation Plans
Paid Holidays
Bonus Programs
+5
Sr. Technical Product Manager, AI Data Platforms
Sr. Technical Product Manager, AI Data Platforms

DDN • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems

NVIDIA • Massachusetts

On-site
USD 272,000 - 432,000
Sr. Technical Product Manager, AI Data Platforms
Sr. Technical Product Manager, AI Data Platforms

Data Direct Networks • Santa Clara (CA)

Hybrid
USD 175,000 - 225,000
Vacation plans
Paid holidays
Bonus programs
+5
Senior Technical Product Manager
Senior Technical Product Manager

DDN • Santa Clara (CA)

On-site
USD 190,000 - 230,000