MTS 2, AI Platform

eBay

Bengaluru

Vor Ort

INR 4.000.000 - 7.000.000

Vollzeit

vor 43 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Mach aus dieser Rolle ein Bewerbungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

eBay is seeking an experienced LLM Inference Engineer to remove compute bottlenecks in production inference. You will own end-to-end inference delivery, optimize GPUs and runtimes, and partner with data science and product teams to meet performance and availability SLOs.

You’ll work across vLLM, Triton, CUDA, and TensorRT, driving measurable improvements in latency and throughput on real traffic, while shaping scalable AI systems.

Qualifikationen

  • 5+ years of development experience with scalable systems.
  • Experience deploying LLM inference services in production.
  • Strong Python and Go or Rust skills.
  • Experience with PyTorch, vLLM, or TensorRT.
  • Knowledge of GPU architecture and performance profiling.
  • Expertise in LLM inference optimization techniques.
  • 3+ years optimizing AI/ML workloads.
  • Ability to deliver measurable production improvements.
  • Experience using autonomous AI coding agents to speed pipelines.
  • Root-cause analysis across model, runtime, network, infra.

Aufgaben

  • Own production inference: take models from handoff to production-grade serving, incl. release eng., capacity planning, cost optimization, incident response.
  • Tune inference performance: reduce latency and increase throughput on real production traffic.
  • Optimize runtimes and servers: scale across heterogeneous GPU fleets; optimize vLLM, Triton, schedulers, KV cache, batching, memory.
  • Benchmark and measure: build benchmarking suites, metrics, tooling for latency, throughput, GPU utilization, memory, cost.
  • Reliability and observability: improve monitoring, tracing, alerting; participate in postmortems to harden systems.
  • Apply and ship new optimizations: evaluate research; implement quantization, paging, kernel/runtimes improvements.
  • Partner with cross-functional teams: translate requirements into performance and availability SLOs.

Kenntnisse

Python
Go/Rust
LLM inference
Performance optimization
GPU/CPU profiling
ML frameworks
System deployment
Root-cause analysis
Autonomous AI coding agents
Throughput/latency tuning

Tools

vLLM
Triton
CUDA
TensorRT
Profiling tools

Jobbeschreibung

At eBay, we're more than a global ecommerce leader — we’re changing the way the world shops and sells. Our platform empowers millions of buyers and sellers in more than 190 markets around the world. We’re committed to pushing boundaries and leaving our mark as we reinvent the future of ecommerce for enthusiasts.

Our customers are our compass, authenticity thrives, bold ideas are welcome, and everyone can bring their unique selves to work — every day. We're in this together, sustaining the future of our customers, our company, and our planet.

Join a team of passionate thinkers, innovators, and dreamers — and help us connect people and build communities to create economic opportunity for all.

As an LLM Inference Engineer on our AI Platform team, you’ll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable—the “last mile” from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps, and requires strong intuition for how model architecture, runtimes, and hardware interact.

What You’ll Do
  • Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.
  • Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
  • Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e.g., schedulers, KV cache, batching, memory).
  • Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.
  • Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
  • Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e.g., quantization, paging, kernel/runtimes improvements).
  • Partner with cross-functional teams: Work with data science and product teams to translate business requirements into performance and availability SLOs.
What We’re Looking For
  • 5+ years of strong development experience
  • Experience deploying and operating LLM inference services in production.
  • Strong production coding skills in Python plus Go or Rust (systems-level implementation and debugging).
  • Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
  • Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
  • Solid understanding of LLM inference and optimization techniques: continuous batching, KV cache management, quantization, speculative decoding (nice-to-have), etc.
  • 3+ years hands-on experience in performance optimization and systems programming for AI/ML workloads.
  • Demonstrated ability to deliver measurable production improvements (e.g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
  • Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.
  • Demonstrated proficiency in applying autonomous AI coding agents to speed up software delivery pipelines. This includes advanced prompting and careful human-in-the-loop code review to improve development speed and code accuracy.
Additional Details

eBay is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, sexual orientation, gender identity, veteran status, and disability, or other legally protected status. If you have a need that requires accommodation, please contact us at talent@ebay.com. We will make every effort to respond to your request for accommodation as soon as possible. View our accessibility statement to learn more about eBay's commitment to ensuring digital accessibility for people with disabilities.

We use cookies to enhance your experience and may use AI tools for administrative tasks in the hiring process. To learn how we handle your personal data and use AI responsibly, please visit our Talent Privacy Notice, Privacy Center, and AI Hiring Guidelines.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

MTS 2, AI Platform Professional
MTS 2, AI Platform Professional

The Networker • Bengaluru

Vor Ort
INR 3.000.000 - 5.200.000
Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

Vor Ort
INR 4.000.000 - 7.500.000
MTS-2 Machine Learning Engineer
MTS-2 Machine Learning Engineer

eBay • Bengaluru

Vor Ort
INR 4.000.000 - 7.000.000
MTS-1 Machine Learning Engineer
MTS-1 Machine Learning Engineer

eBay • Bengaluru

Vor Ort
INR 3.500.000 - 7.000.000
Data Scientist, Knowledge Management
Data Scientist, Knowledge Management

eBay • Bengaluru

Vor Ort
INR 1.800.000 - 3.000.000
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)

Solutions By Text • Bengaluru

Vor Ort
INR 2.500.000 - 4.000.000
AI Engineer - AI & Automation
AI Engineer - AI & Automation

eBay Inc. • Bengaluru

Vor Ort
INR 3.000.000 - 6.000.000
MTS 2, Software Engineer - FullStack
MTS 2, Software Engineer - FullStack

eBay • Bengaluru

Vor Ort
INR 4.000.000 - 7.000.000
Maternal & paternal leave
Paid sabbatical
Financial security plans
MTS 2, Platform Reliability Engineer
MTS 2, Platform Reliability Engineer

eBay • Bengaluru

Vor Ort
INR 4.200.000 - 6.400.000
MTS 2, Backend Software Engineer
MTS 2, Backend Software Engineer

eBay • Bengaluru

Vor Ort
INR 4.000.000 - 7.000.000
Maternal & paternal leave
Paid sabbatical