Stand out for this role — generate a tailored resume and cover letter in about a minute.
Remanence in Europe seeks a Member of Technical Staff - Infrastructure to own the compute and execution platform for training, inference and evaluation. You will ensure GPU resources are productive and research workflows are easy to operate and debug.
You will build and run GPU clusters, scheduling, networking and storage for distributed ML workloads; develop the execution platform for many concurrent environments; create data and artifact pipelines and improve observability, latency and
Remanence is pioneering the next era of enterprise AI by building intelligent systems that learn continuously from real-world execution. We transform complex enterprise workflows and business context into dynamic, interactive environments where AI agents can safely learn, adapt, and improve. By combining high-fidelity simulation environments with state-of-the-art training loops, we build specialized models that solve long-horizon, complex tasks with unmatched reliability. We're building the most talent-dense AI team in Europe to make this happen.
You'll own the compute and execution platform supporting training, inference, task generation, and evaluation. Your work will make GPU resources productive and research workflows straightforward to operate and debug.
You have strong systems fundamentals and can reason carefully about concurrency, resource contention, and failure recovery. You enjoy making complex infrastructure understandable and dependable for the people using it.
We welcome infrastructure specialists and exceptionally fast-learning generalists. Prior ML infrastructure experience is valuable; evidence of building and operating reliable systems matters.