FLORIAN BASTIN
All sessions

Course 01

Transformer

We start with the main NLP tasks and the architectures used before Transformers. We then look at why recurrent networks struggle with long dependencies, how attention works, and how it leads to the Transformer architecture.

On the programme
  1. 01Background on NLP and tasks
  2. 02Tokenization
  3. 03Embeddings
  4. 04Word2vec, RNN, LSTM
  5. 05Attention mechanism
  6. 06Transformer architecture