Building AI dev tools and infra for LLMs; uses vLLM, llm-d, Ray Serve and other inference frameworks and contributes upstream to open-source inference ecosystems.
About the Role
Lead design and delivery of production-grade, multi-tenant LLM inference systems at DigitalOcean, embedding with customers to debug and optimize GPU-backed inference workloads. The role focuses on cluster-scale distributed inference architecture, performance optimization, and translating customer needs into reusable internal tooling and open-source contributions.
Job Description
Role
Senior Forward Deployed Engineer focused on AI inference. You will lead end-to-end design, development, and delivery of production LLM inference workloads, embed with customer teams to diagnose and optimize latency and GPU utilization, and produce reusable blueprints and upstream fixes for open-source inference ecosystems.
Key Responsibilities
- Act as AI inference lead on the FDE team, driving technical direction and delivery for customer-facing inference deployments.
- Architect and deploy production-grade, multi-tenant LLM inference engines using Kubernetes-native frameworks.
- Embed with external tech leads to debug latency spikes, profile GPU memory utilization, and refactor inference code for high-concurrency workloads.
- Implement cluster-scale optimizations (e.g., prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, expert parallelism for MoE models).
- Guide customers on hardware and compute efficiency: tensor/data parallelism, continuous batching, and quantization strategies (FP8/FP4).
- Build internal tooling and translate customer edge cases into reusable deployment blueprints; contribute performance fixes upstream to open-source inference projects.
Requirements
- 6+ years in AI/ML systems with deep experience in cluster-scale serving and distributed inference challenges.
- Hands-on experience with inference frameworks such as vLLM, llm-d, SGLang, TensorRT-LLM, Modular MAX, or similar.
- Expert-level proficiency in Python or GoLang; familiarity with gRPC and running critical services on Kubernetes at scale.
- Strong distributed systems and architecture skills, including profiling, load balancing, and memory/GPU optimization.
- Customer-facing engineering experience with clear communication skills and the ability to translate business SLAs into technical solutions.
Preferred Qualifications
- Prior Forward Deployed Engineering, AI inference architecture, or technical consulting experience supporting production AI systems.
- Experience collaborating with GPU vendors, infrastructure providers, or model vendors on benchmarking, optimization, or launch readiness.
- Preference for engineers who deliver production-ready code, low-latency container images, and deployment blueprints.
Location & Travel
- This job is located in Bengaluru, India.
- Ability to travel up to 30% for customer engagements, workshops, conferences, and internal collaboration.
- Must consistently overlap with North American business hours, including availability until at least noon Eastern Time.
Skills
Technical Leadership Distributed Systems System Design Performance Optimization Architecture Proficiency Customer-facing Communication Troubleshooting Collaboration Scalability Engineering GPU/Accelerator Optimization
Experience Level
Senior
Employment Type
Full Time, Permanent
- Travel up to 30%
- Opportunity to contribute to open-source projects