Recevez plus de réponses des employeurs
Envoyez un CV adapté au poste en quelques minutes.
Kalray is seeking an AI Runtime / Low-Level Software Engineer to own the offload runtime for inference workloads on our custom RISC-V AI accelerator, interfacing with an x86 host. You will collaborate with AI, hardware, and compiler teams to drive the stack from emulation to production silicon.
You will develop and optimize the device lifecycle, scheduling, memory movement, and host-accelerator communication while maintaining robust CI tests and emulated environments.
Kalray is a European leader in hardware acceleration, with full-stack acceleration expertise: from silicon to complete system.
Our MPPA® (Massively Parallel Processor Array) architecture is the foundation of Kalray’s processor (30+ patent families, 15+ years of development) and acceleration cards that combine processing power, flexibility, and energy efficiency.
Our mission is to deliver open data-efficient hardware accelerators to power next generation of data-intensive, AI-driven systems and infrastructures. We offer off-the-shelf processors, acceleration cards, and specialized processor development.
With over 130 employees and presence in France and Romania, Kalray is backed by top-tier investors and publicly listed on Euronext Growth. You’ll be part of a pioneering team that is shaping the future of computing with cutting-edge processor architecture, software-defined solutions, and next-generation acceleration platforms.
We offer a fast-paced, inclusive, and collaborative environment where ambitious experienced professionals and young talents can thrive—while enjoying the positive and inspiring lanscape of the Alps or the Côte d’Azur. You can learn more about us on our website, follow us on LinkedIn.
As an AI Runtime / Low-Level Software Engineer, you will own the offload runtime layer that enables inference workloads to run efficiently on our custom RISC-V AI accelerator from an x86 host.
You will join our AI & Compute team, which is building a full-stack GenAI inference platform, from serving to silicon: LLM Serving → AI Compilation → Runtime / Offload → Optimized AI Kernel Libraries
You will develop the Low-Level software stack responsible for device lifecycle management, scheduling and workload dispatch, high-performance host-to-device communication, and the runtime APIs exposed to compiler and serving layers. You will build this stack end-to-end on an open software foundation and collaborate closely with hardware, compiler and AI serving teams to drive the solution from emulation to production silicon.
Your main responsibilities will include:
Technical skills:
Nice to have:
Profile:
KALRAY is committed to creating a diverse and inclusive environment, and we welcome applications from individuals of all backgrounds, identities, and experiences. We do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, skin color, national origin, gender, sexual orientation, age, marital status, disability status, or any other characteristic protected by law. Should you require any accommodations or adjustments throughout the interview process and beyond, please do not hesitate to let us know. We are committed to ensuring that all candidates have an equal opportunity to showcase their abilities and succeed in our organization.