FLORIAN BASTIN
All sessions

Course 02

Transformer-based models and tricks

Attention on its own carries no information about the order of tokens. We look at how models represent position and how the approaches evolved up to RoPE.

The second half covers encoder models, BERT in particular, and how they are used for classification and search.

On the programme
  1. 01Attention approximation
  2. 02MHA, MQA, GQA
  3. 03Position embeddings (regular, learned)
  4. 04RoPE and applications
  5. 05Transformer-based architectures
  6. 06BERT and its derivatives