Member of Technical Staff - Inference & Hardware Optimization

Albs Labs GmbH

Freiburg im Breisgau

Hybrid

EUR 90.000 - 130.000

Vollzeit

Vor 6 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

A complete application in a minute — tailored resume and cover letter, ready to send.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Albs Labs GmbH is searching for a Staff / Senior IC to drive performance engineering across the real-time multimodal inference stack. You will work with the research team on architecture and quantization decisions, aiming to reduce latency, memory usage, and power while supporting diverse hardware targets.

You will contribute to streaming inference, KV cache management, and on-device models for speech, vision, and language, with a hands-on approach to kernel and runtime optimization.

Qualifikationen

  • Several years of hands-on experience in performance engineering for ML workloads.
  • Strong systems programming in C++, C, or Rust.
  • Experience with SIMD intrinsics (NEON, AVX), CUDA, or vendor NPU SDKs.
  • Understanding low-precision inference and quantization.
  • Ability to read profiling traces and vendor documentation.
  • Bonus: ML compilers, streaming inference, or on-device speech/vision models.

Aufgaben

  • Work across the inference stack end to end: execution engine, memory planning, scheduling, and low-precision kernels.
  • Collaborate with research on architecture and quantization decisions to optimize latency, memory and power.
  • Make real-time multimodal inference fast: streaming, KV cache management, time-to-first-token, and running speech, vision and language components.
  • Improve performance via operator fusion, data layouts, and threading; drop to SIMD, GPU, or NPU kernels when needed.
  • Support hardware targets from ARM CPUs to vendor NPUs as needed for customers.
  • Build benchmarking harness: per-layer latency profiling, memory traces, and power checks.

Kenntnisse

Performance engineering
C++ / Rust
Systems programming
SIMD (NEON / AVX)
CUDA / NPU SDKs
Low-precision inference

Tools

CUDA
Vendor NPU SDKs
SIMD intrinsics

Jobbeschreibung

About us

Albs is an AI research lab building real-time multimodal intelligence for machines, enabling them to see, hear, reason, and interact. We treat model architecture, inference, and runtime as one system. Our purpose-built models and optimized runtimes unlock the full potential of each device within defined limits for hardware cost, power consumption, and response time. Companies can adapt our technology to their own machines without building the underlying AI from scratch.

We founded Albs at the intersection of LLM architecture research and on-device engineering, and we are currently in stealth, but well funded, with dedicated compute for large-scale training and experimentation. Publishing is a core part of our research culture, and we contribute our work to top venues. We share more details about the company, the team and our backing in the first conversation.

Who we're looking for

This is a Staff / Senior IC role. We are looking for experienced engineer, typically with several years of industry experience or an equivalent track record. The exact scope of each role depends on your background: some people go deep on one part of the stack, others shape the technical direction of a whole area. We agree on scope together with you during the interview process. We welcome applications from all qualified candidates, regardless of gender, age, ethnic origin, religion, disability or sexual orientation.

What you'll work on
  • Work across our inference stack end to end: execution engine, memory planning, scheduling, and the low-precision kernels our models actually run on.

  • Work directly with the research team on architecture and quantization decisions, and bring latency, memory and power requirements into the model design.

  • Make real-time multimodal inference fast: streaming, KV cache management, time-to-first-token, and running speech, vision and language components side by side.

  • Squeeze the last few percent out of every target: operator fusion specific to our architecture, cache-aware data layouts and threading. Build on existing runtimes where they serve us, and drop down to SIMD, GPU or NPU kernels when profiling shows that's where the remaining performance is.

  • Bring up new hardware targets, from ARM CPUs and mobile GPUs to vendor NPUs, as we and our customers expand.

  • Build and extend our benchmarking harness: per-layer latency profiling, memory traces, power measurement and accuracy checks across all the hardware we support

What we're looking for
  • You have several years of hands-on experience in performance engineering for numerical or ML workloads, in industry or research.

  • You write strong systems code in C++, C or Rust.

  • You have optimized ML workloads on constrained hardware at the kernel or runtime level, with hands-on experience in SIMD intrinsics (NEON, AVX), CUDA, or vendor NPU SDKs.

  • You understand low-precision inference and quantization, and how they trade off speed against model accuracy.

  • You are comfortable reading profiling traces and vendor documentation, and working out where the cycles went.

  • Bonus: experience with ML compilers, streaming inference, or speech and vision models on device.

How We Work Together

We are a small, focused team of experts. You would join early, work directly with the founders, and help shape how we build. Fast iteration, short lines of communication, and in-person discussion matter a lot to us. Our culture is built around the office in Freiburg, Germany, with a default of three days a week on site. Alternatively, you can work remotely and join us on a regular cadence. We will discuss what works best for you during the interview process.

We hold ourselves to a high standard: the research must be rigorous, the understanding deep, and the product well crafted. Ideas are judged on their merit, not on who proposed them, credit belongs to the team, and no task is beneath anyone. We make ambitious bets and ship early instead of waiting for perfect conditions. Above all, we care about each other, because initiative only works when it comes with respect for the people around you.

What we offer

Expect a competitive base salary plus equity, and dedicated compute for your research and experiments. We give you the time and support to publish and present at top conferences, and flexibility in how you work, whether on site or remote with regular in-person visits. With flat hierarchies and short decision paths, good ideas move from discussion to experiment quickly. And yes, the coffee is excellent, and we regularly get together as a team outside of work. We go through compensation and all other details with you early in the process.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Member of Technical Staff - Data Engineering
Member of Technical Staff - Data Engineering

Albs Labs GmbH • Freiburg im Breisgau

Hybrid
EUR 120.000 - 180.000
Equity
Remote-friendly work
Support to publish at conferences
Member of Technical Staff - LLM Pre-Training
Member of Technical Staff - LLM Pre-Training

Albs Labs GmbH • Freiburg im Breisgau

Hybrid
EUR 120.000 - 180.000
Equity
Dedicated compute for research
Publish support and conferences
+1
Member of Technical Staff - Speech Language Models
Member of Technical Staff - Speech Language Models

Albs Labs GmbH • Freiburg im Breisgau

Hybrid
EUR 120.000 - 180.000
Equity
Flexible work model
Member of Technical Staff - LLM Post-Training
Member of Technical Staff - LLM Post-Training

Albs Labs GmbH • Freiburg im Breisgau

Hybrid
EUR 120.000 - 180.000
Equity
Remote-friendly
Conference support
Member of Technical Staff - Vision-Language Models
Member of Technical Staff - Vision-Language Models

Albs Labs GmbH • Freiburg im Breisgau

Hybrid
EUR 120.000 - 190.000
Equity
Flexible work policy
On-site three days a week in Freiburg
Founding ML Researcher
Founding ML Researcher

Base Compute • Berlin

Vor Ort
EUR 110.000 - 170.000
Founding team equity
Strong base salary
Founding AI Engineer
Founding AI Engineer

alago • München

Vor Ort
Confidential
Wettbewerbsfähige Gehälter
Beteiligung am Unternehmen
Zugang zu EGYM Wellpass
Senior Research Scientist
Senior Research Scientist

adaption • Berlin

Vor Ort
EUR 90.000 - 130.000
Flexible work
Travel stipend
Lunch stipend
+2
AI Research Engineer (Kernel & Inference Optimization) arbeitnow Jobgether Germany · 9/30/2026
AI Research Engineer (Kernel & Inference Optimization) arbeitnow Jobgether Germany · 9/30/2026

Primetime • Deutschland

Remote
EUR 110.000 - 160.000
Remote-first working environment
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Deutschland

Vor Ort
USD 120.000 - 180.000