Research Engineer, Forge

Jobtailor

Paris

Sur place

EUR 70 000 - 100 000

Plein temps

Il y a 2 jours
Soyez parmi les premiers à postuler

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Jobtailor is seeking an experienced ML Engineer to translate real customer requirements into reliable training and deployment workflows, spanning model adaptation, post-training evaluation, data pipelines, and infrastructure.

You will build and improve post-training and evaluation workflows for CPT, SFT, RL, and distillation, develop tools for synthetic data generation, debugging, observability, and scalable systems across cloud and on‑prem environments.

Qualifications

  • Experience with large codebases, testing, CI, and deployment ownership.
  • Hands-on experience with LLM training or post-training pipelines.
  • Ability to debug distributed ML jobs and data issues.
  • Strong communication with technical and non-technical stakeholders.
  • Experience building ML infrastructure and data pipelines.

Responsabilités

  • Turn real customer requirements into reliable training and deployment workflows
  • Work end-to-end across model adaptation and post-training, evaluation, data, and infrastructure
  • Build and improve post-training and evaluation workflows for CPT, SFT, RL, and distillation
  • Develop tools and pipelines for synthetic data generation, data curation, training, evaluation, and deployment
  • Debug and harden large-scale ML systems, including distributed training, scheduling/execution, checkpointing, observability, and reproducibility
  • Improve the Forge codebase through clear APIs, tests, documentation, and maintainable abstractions
  • Advance the RL training stack, including high-throughput asynchronous rollout and scalable post-training systems
  • Ensure Forge deployment is seamless and adaptable across diverse clients, hardware, software stacks, cloud, and on-premises environments
  • Partner with researchers and infrastructure engineers to translate bottlenecks into concrete system improvements
  • Collaborate with scientists, engineers, product, and customer-facing teams to ship maintainable and trusted Forge projects

Connaissances

Python Engineering
PyTorch
JAX
Distributed Systems
Data Pipelines
Testing & CI
Operational Ownership
Clear Communication

Outils

FSDP
DeepSpeed
Megatron
SLURM
Ray
Kubernetes
Kueue
Karpenter
Skypilot

Description du poste

  • Turn real customer requirements into reliable training and deployment workflows
  • Work end-to-end across model adaptation and post-training, evaluation, data, and infrastructure
  • Build and improve post-training and evaluation workflows for CPT, SFT, RL, and distillation
  • Develop tools and pipelines for synthetic data generation, data curation, training, evaluation, and deployment
  • Debug and harden large-scale ML systems, including distributed training, scheduling/execution, checkpointing, observability, and reproducibility
  • Improve the Forge codebase through clear APIs, tests, documentation, and maintainable abstractions
  • Advance the RL training stack, including high-throughput asynchronous rollout and scalable post-training systems
  • Ensure Forge deployment is seamless and adaptable across diverse clients, hardware, software stacks, cloud, and on-premises environments
  • Partner with researchers and infrastructure engineers to translate bottlenecks into concrete system improvements
  • Collaborate with scientists, engineers, product, and customer-facing teams to ship maintainable and trusted Forge projects
Requirements
  • Strong Python engineering skills and experience working in large codebases, including testing, code review, CI, and operational ownership
  • Hands-on experience with PyTorch, JAX, or similar
  • Strong systems and infrastructure fundamentals
  • Experience with LLM training or post-training, including fine-tuning, RL, distillation, evaluation, and/or data pipelines
  • Excellent debugging skills in distributed jobs, data issues, quality regressions, and infrastructure failures
  • Clear communication with technical and non-technical stakeholders
  • High agency, low ego, and comfort in fast-moving, under-specified environments
  • Nice to have: distributed training experience with FSDP, DeepSpeed, Megatron, or similar
  • Nice to have: cluster/orchestration experience with SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, or similar
  • Nice to have: experience building reliable ML infrastructure, evaluation systems, or large-scale data processing pipelines
  • Nice to have: research experience in LLMs, agents, multimodal models, reasoning, code, or domain adaptation
  • Nice to have: open-source contributions, publications, or widely used internal tooling
  • Nice to have: experience training multi-billion-parameter models and on petabyte- and exabyte-scale datasets
  • Nice to have: ability to identify bottlenecks across the stack and drive improvements from first principles
Core Competencies

Demonstrates strong Python engineering skills and hands-on experience with PyTorch or JAX, focusing on building and improving ML workflows, including training, evaluation, and deployment. Capable of collaborating with cross-functional teams to enhance system performance and reliability in fast-paced environments.

Highest-signal resume keywords
  • Python Engineering
  • ML Workflow Development
  • Distributed Training
  • Post-Training Evaluation
  • Clear Communication
ATS Optimization Keywords
Hard Skills
  • Python
  • PyTorch
  • JAX
  • LLM Training
  • Debugging
  • Data Pipelines
  • Testing
  • Code Review
  • Operational Ownership
  • Infrastructure Fundamentals
Soft Skills
  • Clear Communication
  • High Agency
  • Low Ego
Industry Keywords
  • Model Adaptation
  • Synthetic Data Generation
  • Large-Scale ML Systems
  • Observability
  • Reproducibility
  • Maintainable Abstractions
  • Scalable Systems
  • Bottleneck Identification
Tools & Technologies
  • FSDP
  • DeepSpeed
  • Megatron
  • SLURM
  • Ray
  • Kubernetes
  • Kueue
  • Karpenter
  • Skypilot
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Backend Software Engineer – Core AI
Backend Software Engineer – Core AI

Jobtailor • Strasbourg

Sur place
EUR 90 000 - 120 000
AI Engineer
AI Engineer

Jobtailor • Paris

Sur place
EUR 60 000 - 90 000
Tech Lead AI
Tech Lead AI

Jobtailor • Nantes

Sur place
EUR 90 000 - 130 000
Staff Software Engineer
Staff Software Engineer

Startup Talents • France

À distance
EUR 60 000 - 90 000
Healthcare and life insurance
Retirement plan
Sponsored transportation
+3
Senior AI Engineer
Senior AI Engineer

Jobtailor • Paris

Sur place
EUR 90 000 - 120 000
AI and GenAI Architect Consultant
AI and GenAI Architect Consultant

Jobtailor • Rennes

Sur place
EUR 70 000 - 100 000
Technical Program Manager, Science Operations
Technical Program Manager, Science Operations

Jobtailor • Paris

Sur place
EUR 75 000 - 110 000
Technical Program Manager
Technical Program Manager

Jobtailor • Paris

Sur place
EUR 90 000 - 120 000
ML Infrastructure Engineer
ML Infrastructure Engineer

Visa Hunt • Paris

Hybride
EUR 85 000 - 135 000
Relocation package
Comprehensive medical insurance (Paris
Hardware and tools provided
+2
Technical Staff Member – Agent
Technical Staff Member – Agent

Jobtailor • Paris

Sur place
EUR 120 000 - 180 000