Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
RIVR Technologies AG is seeking an expert in Vision-Language-Action models to lead multi-modal robotics research in Zürich. You will develop VLA models, imitation learning and transformer-based architectures to advance autonomous delivery robots in real-world environments.
You will supervise and mentor a team of software engineers, collaborate with RL teams, and help translate research into deployed edge solutions on hardware like Nvidia Jetson Thor.
RIVR, an Amazon company, is building Physical AI by deploying autonomous robots for real-world doorstep delivery. Operating daily in diverse urban environments, RIVR's robots continuously learn from and navigate the millions of scenarios encountered during deliveries. By owning the full stack from software.
Our fleet of delivery robots operates globally today, generating vast amounts of robotic real-world data. By utilizing state-of-the-art Vision-Language-Action (VLA) models, large-scale generalist models (like Transformers), generative AI, and similar methods, we can leverage this pool of data to significantly enhance its autonomy, navigation, and manipulation skills. In this role, you will develop multi-modal models that enable robots to autonomously generate actions from demonstrations, real-time sensor data, and natural language commands. We are seeking an expert in VLA models, imitation learning, and generative AI techniques with a deep knowledge of supervised, and self-supervised learning algorithms. If you are passionate about pushing the boundaries of AI we invite you to join us in shaping the future of intelligent robotics.