Inference Optimization — Member of Technical Staff

Construct Labs

Berlin

Vor Ort

EUR 120.000 - 180.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Construct Labs in Berlin is seeking an Inference Optimization — Member of Technical Staff to push GPU efficiency for deployed models. You will identify novel hardware-informed quantization methods, write and tune Triton and CUDA kernels, and locate bottlenecks in compute, memory, and communication.

Great candidates may come from ML systems, compilers, mathematics, physics, competitive programming, security, or HPC, and share a habit of deep reasoning and a passion for optimizing end-to-end

Qualifikationen

  • Experience with model quantization techniques and hardware-aware optimization.
  • Ability to profile and optimize GPU workloads across compute, memory, and communication.
  • Strong foundation in machine learning systems and performance engineering.

Aufgaben

  • Find novel hardware-informed quantization methods across model architectures.
  • Write and tune Triton and CUDA kernels.
  • Identify bottlenecks across compute, memory, and communication.
  • Efficiently serving LoRAs and specialized weights for model adaptation.

Kenntnisse

Quantization methods
GPU profiling
Performance optimization
ML systems

Tools

Triton
CUDA
Profiling tools

Jobbeschreibung

Back to careers


Inference Optimization — Member of Technical Staff


Berlin, Full-time, In person


Who Are We

Models should learn from what happens after deployment. We're building the loops that make this possible.


We will be the default inference provider for specialized tokens, running models that continuously adapt to each customer and workload. That means serving many continuously adapting models with frontier-level performance and attractive economics.


The Work


  • Finding novel, hardware-informed quantization methods across model architectures

  • Writing and tuning Triton and CUDA kernels

  • Finding bottlenecks across compute, memory, and communication

  • Efficiently serving LoRAs, learned memory, specialized weights, and other forms of model adaptation


This is capture-the-flag for GPU efficiency: profile the system, find something everyone else missed, and prove the gain on real workloads.


Great candidates might come from ML systems, compilers, mathematics, physics, competitive programming, security, or HPC. The common thread is strong first-principles reasoning, an obsession with efficiency, and a habit of going deep on hard problems for the fun of it.


We work together in person from our office in Berlin.


Interview process


  • Two technical interviews

  • An onsite interview at our Berlin office

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Model Adaptation — Member of Technical Staff
Model Adaptation — Member of Technical Staff

Construct Labs • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Corporation • Berlin

Vor Ort
EUR 120.000 - 180.000
Applied Mathematician
Applied Mathematician

Circonomit • Berlin

Hybrid
EUR 90.000 - 140.000
Impact on factory planning
Ownership of engine in production
Equity potential (VSOP)
+4
Applied Operations Research Engineer
Applied Operations Research Engineer

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Hardware budget
Equity/VSOP
Founding ML Researcher
Founding ML Researcher

Base Compute • Berlin

Vor Ort
EUR 95.000 - 170.000
Founding team equity
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000
Applied Operations Research Engineer
Applied Operations Research Engineer

Circonomit GmbH • Köln

Hybrid
EUR 110.000 - 140.000
Hardware of your choice
AI tooling budget
sports membership
+1
Senior Software Engineer – TensorRT Edge-LLM
Senior Software Engineer – TensorRT Edge-LLM

NVIDIA • Deutschland

Hybrid
EUR 159.000 - 248.000
Systems Engineer - m/f/d
Systems Engineer - m/f/d

Linuxconfig • Berlin

Hybrid
EUR 90.000 - 140.000
Applied Operations Research Engineer
Applied Operations Research Engineer

Circonomit • Köln

Vor Ort
EUR 90.000 - 130.000
Impact on production
Ownership
Hardware of your choice
+4