A while back someone figured out that neural nets needed history, and RNNs (recursive neural networks) were born. The current "time" in the network referenced previous "times." The problem was, RNNs kind of sucked.
The attention mechanism is a fancier version of the same idea. Instead of referring to a few previous times, you refer to ALL previous times with context-dependent weights.
Transformers were the end state of replacing recursive nets with attention in a much of different places.