The Many Minds of Transformers
What each attention head is really learning, and why running several in parallel gives models richer representations.
Originally published on Medium (~920 words). Read the full piece here: The Many Minds of Transformers.