Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru

Vibehackers

Bengaluru

On-site

INR 3,500,000 - 6,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Travel up to 30%
Open-source contributions

Job summary

Vibehackers in Bengaluru, India, seeks a Senior Forward Deployed Engineer focused on AI inference. You will lead end-to-end design, development, and delivery of production LLM inference workloads, embedding with customer teams to diagnose latency and optimize GPU usage.

You will build reusable blueprints and upstream fixes for open-source inference ecosystems, using Kubernetes-native frameworks and advanced performance techniques to scale multi-tenant deployments.

Qualifications

  • 6+ years in AI/ML systems with distributed inference experience.
  • Hands-on with inference frameworks such as vLLM, llm-d, SGLang, TensorRT-LLM, Modular MAX, or similar.
  • Proficiency in Python or GoLang; familiarity with gRPC and Kubernetes at scale.
  • Strong distributed systems and architecture skills: profiling, load balancing, memory/GPU optimization.
  • Customer-facing engineering experience with translating SLAs into technical solutions.

Responsibilities

  • Act as AI inference lead on the FDE team, driving technical direction for customer deployments.
  • Architect and deploy multi-tenant LLM inference engines using Kubernetes-native frameworks.
  • Embed with customer teams to diagnose latency, profile GPU memory, and refactor inference code for high concurrency.
  • Implement cluster-scale optimizations such as prefill/decode disaggregation and KV-cache routing.
  • Guide customers on hardware efficiency: tensor/data parallelism, batching, and quantization strategies.

Skills

AI systems
Distributed systems
Kubernetes
Python
GoLang
gRPC
Performance optimization
Customer-facing
System design

Tools

vLLM
llm-d
SGLang
TensorRT-LLM
Modular MAX

Job description

Building AI dev tools and infra for LLMs; uses vLLM, llm-d, Ray Serve and other inference frameworks and contributes upstream to open-source inference ecosystems.

About the Role

Lead design and delivery of production-grade, multi-tenant LLM inference systems at DigitalOcean, embedding with customers to debug and optimize GPU-backed inference workloads. The role focuses on cluster-scale distributed inference architecture, performance optimization, and translating customer needs into reusable internal tooling and open-source contributions.

Job Description
Role

Senior Forward Deployed Engineer focused on AI inference. You will lead end-to-end design, development, and delivery of production LLM inference workloads, embed with customer teams to diagnose and optimize latency and GPU utilization, and produce reusable blueprints and upstream fixes for open-source inference ecosystems.

Key Responsibilities
  • Act as AI inference lead on the FDE team, driving technical direction and delivery for customer-facing inference deployments.
  • Architect and deploy production-grade, multi-tenant LLM inference engines using Kubernetes-native frameworks.
  • Embed with external tech leads to debug latency spikes, profile GPU memory utilization, and refactor inference code for high-concurrency workloads.
  • Implement cluster-scale optimizations (e.g., prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, expert parallelism for MoE models).
  • Guide customers on hardware and compute efficiency: tensor/data parallelism, continuous batching, and quantization strategies (FP8/FP4).
  • Build internal tooling and translate customer edge cases into reusable deployment blueprints; contribute performance fixes upstream to open-source inference projects.
Requirements
  • 6+ years in AI/ML systems with deep experience in cluster-scale serving and distributed inference challenges.
  • Hands-on experience with inference frameworks such as vLLM, llm-d, SGLang, TensorRT-LLM, Modular MAX, or similar.
  • Expert-level proficiency in Python or GoLang; familiarity with gRPC and running critical services on Kubernetes at scale.
  • Strong distributed systems and architecture skills, including profiling, load balancing, and memory/GPU optimization.
  • Customer-facing engineering experience with clear communication skills and the ability to translate business SLAs into technical solutions.
Preferred Qualifications
  • Prior Forward Deployed Engineering, AI inference architecture, or technical consulting experience supporting production AI systems.
  • Experience collaborating with GPU vendors, infrastructure providers, or model vendors on benchmarking, optimization, or launch readiness.
  • Preference for engineers who deliver production-ready code, low-latency container images, and deployment blueprints.
Location & Travel
  • This job is located in Bengaluru, India.
  • Ability to travel up to 30% for customer engagements, workshops, conferences, and internal collaboration.
  • Must consistently overlap with North American business hours, including availability until at least noon Eastern Time.
Skills

Technical Leadership Distributed Systems System Design Performance Optimization Architecture Proficiency Customer-facing Communication Troubleshooting Collaboration Scalability Engineering GPU/Accelerator Optimization

Experience Level

Senior

Employment Type

Full Time, Permanent

  • Travel up to 30%
  • Opportunity to contribute to open-source projects
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Forward Deployed Engineer I (AI Inference)
Senior Forward Deployed Engineer I (AI Inference)

DigitalOcean • Bengaluru

On-site
INR 4,500,000 - 7,000,000
Forward Deployed Engineer I
Forward Deployed Engineer I

Hiringeye Solutions • Bengaluru

Hybrid
INR 400,000 - 700,000
Senior Forward Deployed Engineer I (AI Infra)
Senior Forward Deployed Engineer I (AI Infra)

DigitalOcean • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Forward Deployment Engineer Devops and AI Development
Forward Deployment Engineer Devops and AI Development

PwC • Hyderabad, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Forward Deployed Engineer
Forward Deployed Engineer

B Capital • Gurugram District

On-site
INR 900,000 - 1,300,000
Forward Deployed Engineer
Forward Deployed Engineer

Leena AI • Gurugram District

On-site
INR 3,000,000 - 6,000,000
Senior Technical Leader Forward Deployed Engineer Fde Globallogic Bengaluru
Senior Technical Leader Forward Deployed Engineer Fde Globallogic Bengaluru

Vibehackers • Bengaluru

Hybrid
INR 3,200,000 - 5,200,000
Flexible work schedules
Paid time off
Professional development
+4
Forward Deployed Engineer GCC – Hyderabad
Forward Deployed Engineer GCC – Hyderabad

Summit Consulting Services • Hyderabad

On-site
INR 2,500,000 - 4,200,000
AI_DevOps Engineer
AI_DevOps Engineer

RIA Advisory • Pune District

Hybrid
INR 1,000,000 - 1,500,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000