Frédéric JEHAN

Spécialité : Informatique

Laboratoire : LIPN

Directeur de thèse : Joseph Le Roux

Co-encadrant : Antoine Rozenknop

Titre de la thèse : Diffusion structurée pour un traitement moderne du langage naturel

Les modèles de langage, comme les nombreux chatbots aujourd’hui connus du grand public, reposent le plus souvent sur un algorithme auto-régressif : le texte est généré mot par mot, chaque mot étant produit à partir des mots précédents.

Un nouveau paradigme est récemment apparu, inspiré des modèles de génération d’images fondés sur la diffusion. Appliqué au langage, il consiste à entraîner un réseau de neurones à débruiter un texte dégradé, dont certains mots ont été remplacés au hasard ou retirés. Une fois entraîné, ce réseau devient capable de générer un texte entier à partir d’un texte aléatoire, sans aucune information initiale.

Cette approche a un vrai potentiel : elle permet de générer plusieurs parties du texte simultanément, donc plus rapidement. Mais même en parallèle, ces modèles génèrent les mots un par un, indépendamment les uns des autres, ce qui crée des incohérences et des répétitions que les modèles auto-régressifs ne produisent pas.

C’est ce problème que ma thèse tente de résoudre, en générant les mots par paquets plutôt qu’individuellement, pour améliorer la cohérence globale du texte. Cela introduit toutefois une nouvelle difficulté : réconcilier ces paquets, générés en parallèle, quand ils se chevauchent ou ne correspondent pas entre eux.

Thesis title : Diffusion for structure-aware modern natural language processing

Language models, such as the many chatbots now widely known to the general public, mostly rely on an auto-regressive algorithm: text is generated word by word, with each new word produced from the sequence of preceding words.

A new paradigm has recently emerged, inspired by image generation models based on diffusion processes. Applied to language, it consists in training a neural network to denoise a corrupted text, in which some words have been randomly replaced or simply removed. Once trained, this network becomes able to generate an entire text starting from a purely random text, containing no information at all.

This approach has real potential: it allows several parts of the text to be generated simultaneously, and therefore faster. But even in parallel, these models generate words one by one, independently of each other, which creates inconsistencies and repetitions that auto-regressive models do not produce.

This is the problem my thesis tries to address, by generating words in batches rather than individually, to improve the overall coherence of the text. This, however, introduces a new difficulty: reconciling these batches, generated in parallel, when they overlap or do not match each other.

Retour en haut