Senior RL Engineer - Ingénieur(e) principal(e) en apprentissage par renforcement

NBCUniversal

Quebec

On-site

CAD 90,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NBCUniversal seeks a Reinforcement Learning Engineer to design and operate high-fidelity simulation environments for autonomous agents, focusing on robust rewards, scalable policies, and real-world deployment readiness.

You will collaborate with ML engineers, build 2D/3D simulators (Unity/Unreal/Isaac), implement PPO/SAC, and bridge sim to reality with domain randomization to ensure reliable performance across complex multi-sensor setups.

Qualifications

  • Graduate degree in robotics, CS, AI, or related field with RL focus.
  • Proven experience as RL/Research Engineer in fast-paced environments.
  • Experience in robotics, game development, or aerospace contexts preferred.

Responsibilities

  • Cross-Functional coordination with ML engineers, annotations, and TPMs to define data and training requirements.
  • Design and maintain high-fidelity 2D/3D simulation environments using Unity/Unreal/Isaac Sim.
  • Design and tune reward functions to align agent behavior with goals and safety constraints.
  • Develop and optimize RL algorithms (PPO, SAC, Offline RL) for high-dimensional observation spaces.
  • Analyze sim-to-real gaps and apply domain randomization/adaptation techniques.

Skills

Python
Git
Unix shell
RL frameworks
Ray RLlib
Stable Baselines3
CleanRL

Education

Master’s or PhD in Robotics/CS/AI with RL focus

Tools

MuJoCo
Bullet
Unity
Unreal Engine
Isaac Sim

Job description

NBCUniversal is one of the world's leading media and entertainment companies. We create world-class content, which we distribute across our portfolio of film, television, and streaming, and bring to life through our global theme park destinations, consumer products, and experiences. We own and operate leading entertainment and news brands, including NBC, NBC News, NBC Sports, Telemundo, NBC Local Stations, Bravo, and Peacock, our premium ad-supported streaming service. We produce and distribute premier filmed entertainment and programming through our powerhouse film and television studios, including Universal Pictures, DreamWorks Animation, and Focus Features, and the four global television studios under the Universal Studio Group banner, and operate industry-leading theme parks and experiences around the world through Universal Destinations & Experiences, including Universal Orlando Resort, home to Universal Epic Universe, and Universal Studios Hollywood. NBCUniversal is a subsidiary of Comcast Corporation. Visit www.nbcuniversal.com for more information.

Our impact is rooted in improving the communities where our employees, customers, and audiences live and work. We have a rich tradition of giving back and ensuring our employees have the opportunity to serve their communities. We champion an inclusive culture and strive to attract and develop a talented workforce to create and deliver a wide range of content reflecting our world.

Rendez-vous sur www.nbcuniversal.com pour plus d’informations.

Notre impact repose sur l’amélioration des communautés dans lesquelles vivent et travaillent nos employés, nos clients et nos publics. Nous avons une riche tradition d’engagement social et veillons à ce que nos employés aient la possibilité de s’investir au sein de leurs communautés. Nous défendons une culture inclusive et nous nous efforçons d’attirer et de former une main-d’œuvre talentueuse afin de créer et de proposer un large éventail de contenus reflétant notre monde.

We are seeking a Reinforcement Learning Engineer with experience manipulating virtual environments to train autonomous agents. This role focuses on the design of robust simulation environments, reward structures, and policy architectures that can navigate complex, multi-sensor landscapes.

Key Responsibilities
  • Cross-Functional Coordination: Work with partner ML and Annotation engineers and TPMs to spec out data, simulation, and training requirements.
  • Environment Design: Build and maintain high-fidelity 2D/3D simulation environments (using tools like Unity, Unreal, or Isaac Sim) that serve as the training ground for RL agents.
  • Reward Engineering: Design and tune complex reward functions that align agent behavior with product goals and safety constraints.
  • Algorithm Implementation: Develop and optimize RL algorithms (e.g., PPO, SAC, or Offline RL) capable of handling high-dimensional 3D observation spaces.
  • Sim-to-Real Strategy: Analyze the "reality gap" and implement domain randomization or adaptation techniques to ensure models perform reliably in real-world scenarios.

Nous sommes à la recherche d’un(e) ingénieur(e) en apprentissage par renforcement ayant de l’expérience dans la création et l’exploitation d’environnements virtuels pour l’entraînement d’agents autonomes. Ce rôle consiste à concevoir des environnements de simulation robustes, des structures de récompense et des architectures de politiques capables d’évoluer dans des contextes complexes et multi-capteurs.

Vous jouerez un rôle clé dans le rapprochement entre simulation et performance réelle en développant des systèmes RL évolutifs et en garantissant un comportement fiable des agents dans des conditions variées.

  • Collaboration interfonctionnelle : Travailler avec les ingénieurs ML, les équipes d’annotation et les TPM afin de définir les besoins en données, en simulation et en entraînement.
  • Conception d’environnements : Développer et maintenir des environnements de simulation 2D/3D à haute fidélité à l’aide d’outils tels que Unity, Unreal ou Isaac Sim.
  • Ingénierie des récompenses : Concevoir et optimiser des fonctions de récompense afin d’aligner le comportement des agents avec les objectifs produit et les contraintes de sécurité.
  • Implémentation d’algorithmes : Développer et optimiser des algorithmes d’apprentissage par renforcement (ex. : PPO, SAC, RL hors ligne) adaptés à des espaces d’observation à haute dimension.
  • Stratégie sim-to-real : Réduire l’écart entre simulation et réalité à l’aide de techniques comme la randomisation de domaine et l’adaptation afin d’assurer des performances fiables en conditions réelles.
  • Education: Graduate degree (Master’s or PhD) in Robotics, Computer Science, AI, or a related field with a focus on Reinforcement Learning, Imitation Learning, or other Online Machine Learning fields.
  • Professional Experience: Proven experience as an RL Engineer or Research Engineer in a fast-paced environment.
  • Industry Context: Prior experience in industries with complex multi-disciplinary teams such as robotics, smart grids, precision agriculture, game development, or aerospace.
Technical Proficiency
  • Core Tools: Fluency with Python, Git, and the Unix shell.
  • RL Frameworks: Deep familiarity with frameworks like Ray Rllib, Stable Baselines3, or CleanRL.
  • Physics & 3D Engines: Experience with physics engines (MuJoCo, Bullet) or 3D game engines.
  • Ecosystem: Familiarity with collaborative tools such as Jira/Confluence, Slack, a Git server, and an experiment tracking framework.
Attributes
  • Strong Mathematical Background: Essential for understanding Markov Decision Processes (MDPs) and gradient-based optimization.
  • High Attention to Detail: Critical for debugging non-deterministic agent behaviors and ensuring environment parity.
  • Formation : Maîtrise ou Doctorat en robotique, informatique, intelligence artificielle ou domaine connexe avec une spécialisation en apprentissage par renforcement, imitation ou apprentissage en ligne.
  • Expérience : Expérience démontrée en tant qu’ingénieur(e) en apprentissage par renforcement ou en recherche dans un environnement dynamique.
  • Contexte industriel : Une expérience dans des secteurs multidisciplinaires tels que la robotique, les réseaux intelligents, l’agriculture de précision, les jeux vidéo ou l’aérospatiale est fortement valorisée.
Compétences techniques
  • Outils principaux : Excellente maîtrise de Python, Git et des environnements Unix.
  • Frameworks RL : Expérience avec des frameworks tels que Ray RLlib, Stable Baselines3 ou CleanRL.
  • Physique et simulation : Expérience avec des moteurs physiques (MuJoCo, Bullet) ou des environnements de simulation 3D.
  • Écosystème : Familiarité avec des outils collaboratifs tels que Jira, Confluence, Slack, les workflows Git et les plateformes de suivi d’expériences.
Qualités recherchées
  • Solides bases mathématiques : Bonne compréhension des processus de décision de Markov (MDP) et de l’optimisation basée sur le gradient.
  • Rigueur et précision : Capacité à déboguer des systèmes non déterministes et à assurer la cohérence et la précision des environnements de simulation.

As part of our selection process, external candidates may be required to attend an in-person interview with an NBCUniversal employee at one of our locations prior to a hiring decision. NBCUniversal's policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.

If you are a qualified individual with a disability or a disabled veteran and require support throughout the application and/or recruitment process as a result of your disability, you have the right to request a reasonable accommodation. You can submit your request to AccessibilitySupport@nbcuni.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior RL Engineer - Ingénieur(e) principal(e) en apprentissage par renforcement
Senior RL Engineer - Ingénieur(e) principal(e) en apprentissage par renforcement

NBCUniversal • Montreal (administrative region)

On-site
CAD 90,000 - 140,000
Staff MLOps Engineer | Ingénieur·e MLOps Staff
Staff MLOps Engineer | Ingénieur·e MLOps Staff

NBCUniversal • Quebec

On-site
CAD 110,000 - 160,000
ML Solutions Architect - Architecte de solutions ML
ML Solutions Architect - Architecte de solutions ML

NBCUniversal • Montreal (administrative region)

On-site
CAD 110,000 - 160,000
Senior Data Scientist - Scientifique des données principal(e) (niveau senior)
Senior Data Scientist - Scientifique des données principal(e) (niveau senior)

NBCUniversal • Quebec

Hybrid
CAD 90,000 - 130,000
Senior Multiplayer Programmer | Programmeur ou programmeuse multijoueur senior
Senior Multiplayer Programmer | Programmeur ou programmeuse multijoueur senior

NBCUniversal • Montreal (administrative region)

On-site
CAD 100,000 - 160,000
Senior Producer (Pillar) | Producteur(trice) principal(e) (Pilier)
Senior Producer (Pillar) | Producteur(trice) principal(e) (Pilier)

NBCUniversal • Montreal (administrative region)

On-site
CAD 110,000 - 160,000
Senior Technical Designer
Senior Technical Designer

NBCUniversal • Quebec

On-site
CAD 90,000 - 120,000
Associate Art Director - Characters | Directeur.trice Artistique Associé.e – Personnages
Associate Art Director - Characters | Directeur.trice Artistique Associé.e – Personnages

NBCUniversal • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Lead Environment Artist / Chef Artiste d’environnement
Lead Environment Artist / Chef Artiste d’environnement

NBCUniversal • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Senior Animator - Gameplay | Animateur ou animatrice senior, Jouabilité
Senior Animator - Gameplay | Animateur ou animatrice senior, Jouabilité

Impleo • Quebec

On-site
CAD 70,000 - 110,000