Forward Deployed Inference Engineer

Free resume

München

Hybrid

EUR 53.000 - 59.000

Vollzeit

vor 2 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Tensordyne is seeking a Forward Deployed Inference Engineer to bridge customer workloads with high-performance AI inference on Tensordyne hardware and software. You will own the path from a customer workload or model request to a technical result and ensure reliable deployment outcomes.

In this remote, full‑time role you will profile, benchmark, and optimize models, collaborate with product and engineering teams, and help scale our AI inference platform for demanding workloads.

Qualifikationen

  • Profiling, benchmarking, and optimization of AI models and runtimes.
  • Experience with deploying models to production environments.
  • Ability to work with customers and translate requirements into measurable KPIs.
  • Hands-on with Python and PyTorch for model code changes.
  • Familiarity with inference stacks and novel accelerators.

Aufgaben

  • Work directly with customers on technical PoCs, integration, deployment, and debugging.
  • Translate requirements into measurable acceptance criteria for latency, throughput, and quality.
  • Profile-to-hardware accuracy: compare profiling with real deployment results and identify gaps.
  • Develop reusable tooling, documentation, benchmarks, and product improvements.
  • Collaborate with compiler, runtime, kernel, and system teams to improve results.
  • Engage with customers or external partners and support deployment on new accelerators or non-standard hardware.

Kenntnisse

Profiling & Benchmarking
Model Enablement
Optimization
Deployment & PoCs
Python
PyTorch
LLM inference
Kubernetes
Rust
Communication & teamwork
Problem solving

Tools

Kubernetes
Rust
vLLM
SGLang
CUDA

Jobbeschreibung

About Tensordyne

Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world’s most demanding generative AI workloads.

Our platform combines purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture into a tightly integrated system purpose built for large-scale AI inference. We work with hyperscalers, Neoclouds, frontier model developers, enterprises, and infrastructure partners operating at the leading edge of AI.

As Tensordyne moves from system development into silicon bring-up, customer validation, beta deployments, and production rollout, we are building the technical customer organization that will sit directly between our engineering teams and the companies deploying the platform.

Role summary

We are looking for a Forward Deployed Inference Engineer who combines deep AI systems expertise with strong customer instincts. This person will own the path from a customer workload or model request to a technical result and, where needed, to an optimized model running successfully on Tensordyne hardware and software.

The role sits at the intersection of model architecture, inference performance, systems optimization, developer tooling, and customer deployment. You will work hands-on with engineering while also acting as a technical bridge to Product, BizDev, Sales, and customers.

What you will do
  • Turn customer workloads into fast, credible performance answers through profiling & benchmarking. Define relevant KPIs, compare against competitive baselines, and keep our evaluation methodology current with external benchmarks.
  • Model enablement & optimization: convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify performance bottlenecks, and work with compiler, runtime, kernel, and system teams to improve results.
  • Deployment / forward engineering: work directly with customers and partners on technical PoCs, integration, deployment, and debugging; translate requirements into measurable acceptance criteria for quality, latency, throughput, and other relevant KPIs.
  • Track profiling-to-hardware accuracy by continuously comparing profiling/simulation results with actual hardware deployments, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, KV-cache management, or adjacent systems to the owning teams.
  • Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements.
  • Strong hands-on experience with AI models and inference systems, especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models.
  • Strong Python and PyTorch skills and the ability to understand and modify model code.
  • Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization.
  • Strong problem‑solving and communication skills, with the ability to drive ambiguous technical problems across team boundaries.
  • Proficiency in using AI‑powered developer tools (e.g., Claude Code, Cursor).
  • Experience with LLM serving and deployment stacks such as vLLM, SGLang, or similar systems.
  • Experience working directly with customers or external technical partners.
  • Experience bringing models up on new accelerators or non‑standard hardware, including performance debugging across framework/runtime/hardware boundaries.
  • Practical experience with production inference techniques or environments such as quantization, distributed inference, or Kubernetes.
  • Experience navigating and contributing to Rust codebases.
Tensordyne Values
  • Think big. Pursue ambitious technical and business goals.
  • Aim for excellence. Quality matters in everything we build and deliver.
  • Own it and get it done. Take responsibility and drive results.
  • Operate with integrity. Be direct, transparent, and respectful.
  • Win as a team. Make the people around you more effective.
  • Value different perspectives. The strongest teams challenge assumptions and bring diverse experience to difficult problems.

Tensordyne is an equal opportunity employer. We believe diverse teams are better equipped to solve complex problems and build exceptional technology. All qualified applicants will receive consideration for employment without regard to age, color, gender identity or expression, marital status, national origin, disability, protected veteran status, race, religion, pregnancy, sexual orientation, or any other characteristic protected by applicable laws, regulations, and ordinances.

Remote full-time en 53,000–59,000 EUR/yr

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Forward Deployed Inference Engineer
Forward Deployed Inference Engineer

Tensordyne • München

Vor Ort
EUR 90.000 - 135.000
Simulation & Modeling Engineer
Simulation & Modeling Engineer

Tensordyne • München

Remote
EUR 70.000 - 110.000
Jr Software Engineer - ML Runtime - Rust
Jr Software Engineer - ML Runtime - Rust

Meyandy LLC • München

Vor Ort
EUR 90.000 - 130.000
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

Tether Operations Limited • Deutschland

Remote
EUR 90.000 - 130.000
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Senior Applied Scientist, Efficient LLM Inference & Model Optimization

United States Digital Space LLC • Berlin

Vor Ort
EUR 90.000 - 150.000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

United States Digital Space LLC • Deutschland

Remote
EUR 120.000 - 190.000
Remote-first
International team
Cutting-edge AI research
+2
Field Application Engineer - AI Software & Hardware
Field Application Engineer - AI Software & Hardware

AI Chopping Block • München

Vor Ort
EUR 80.000 - 120.000
Senior Solutions Architect – Large Scale AI Inference
Senior Solutions Architect – Large Scale AI Inference

NVIDIA • Deutschland

Vor Ort
EUR 293.000 - 507.000
(Senior)Product Manager - AI Infrastructure
(Senior)Product Manager - AI Infrastructure

SpiNNcloud Systems • Dresden

Vor Ort
EUR 90.000 - 130.000
Forward Deployed Engineer
Forward Deployed Engineer

Taktile • Berlin

Hybrid
EUR 90.000 - 130.000
Equity
Top-of-market compensation
Self-development budget
+1