The Monte Carlo Transformer: a stochastic self-attention model for sequence prediction - IMT - Institut Mines-Télécom Accéder directement au contenu
Pré-Publication, Document De Travail Année : 2020

The Monte Carlo Transformer: a stochastic self-attention model for sequence prediction

Résumé

This paper introduces the Sequential Monte Carlo Transformer, an original approach that naturally captures the observations distribution in a transformer architecture. The keys, queries, values and attention vectors of the network are considered as the unobserved stochastic states of its hidden structure. This generative model is such that at each time step the received observation is a random function of its past states in a given attention window. In this general state-space setting, we use Sequential Monte Carlo methods to approximate the posterior distributions of the states given the observations, and to estimate the gradient of the log-likelihood. We hence propose a generative model giving a predictive distribution, instead of a single-point estimate.
Fichier principal
Vignette du fichier
smc_transformer_2020.pdf (561.13 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)

Dates et versions

hal-02896961 , version 1 (11-07-2020)
hal-02896961 , version 2 (12-12-2020)

Identifiants

Citer

Alice Martin, Charles Ollion, Florian Strub, Sylvain Le Corff, Olivier Pietquin. The Monte Carlo Transformer: a stochastic self-attention model for sequence prediction. 2020. ⟨hal-02896961v2⟩
221 Consultations
777 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More