Erhalte mehr Antworten von Arbeitgebern
Versende in nur wenigen Minuten einen passgenauen Lebenslauf.
Construct Labs in Berlin is seeking an Inference Optimization — Member of Technical Staff to push GPU efficiency for deployed models. You will identify novel hardware-informed quantization methods, write and tune Triton and CUDA kernels, and locate bottlenecks in compute, memory, and communication.
Great candidates may come from ML systems, compilers, mathematics, physics, competitive programming, security, or HPC, and share a habit of deep reasoning and a passion for optimizing end-to-end
Back to careers
Inference Optimization — Member of Technical Staff
Berlin, Full-time, In person
Models should learn from what happens after deployment. We're building the loops that make this possible.
We will be the default inference provider for specialized tokens, running models that continuously adapt to each customer and workload. That means serving many continuously adapting models with frontier-level performance and attractive economics.
This is capture-the-flag for GPU efficiency: profile the system, find something everyone else missed, and prove the gain on real workloads.
Great candidates might come from ML systems, compilers, mathematics, physics, competitive programming, security, or HPC. The common thread is strong first-principles reasoning, an obsession with efficiency, and a habit of going deep on hard problems for the fun of it.
We work together in person from our office in Berlin.