The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning) – Jay Alammar – Visualizing machine learning one concept at a time.
5
Public whispers
5
Contributors
2026-07-19 09:44:31
First whispered
Public whispers on this page
Text Highlight2026-07-19 13:05:31
Original Highlight Excerpt
"ELMo actually goes a step further and trains a bi-directional LSTM"
Whisper Note
But doesn't training both directions just double the compute for marginal gains?
Text Highlight2026-07-19 12:56:31
Original Highlight Excerpt
"ELMo actually goes a step further and trains a bi-directional LSTM"
Whisper Note
So that's why ELMo feels smarter than the old uni-directional models.
Text Highlight2026-07-19 10:02:31
Original Highlight Excerpt
"BERT is basically a trained Transformer Encoder stack."
Whisper Note
Wait, so no decoder at all? That changes everything for me.
Text Highlight2026-07-19 09:53:31
Original Highlight Excerpt
"BERT is basically a trained Transformer Encoder stack."
Whisper Note
That post on Transformer really is a must-read first.
Text Highlight2026-07-19 09:44:31
Original Highlight Excerpt
"BERT is basically a trained Transformer Encoder stack."
Whisper Note
So it's just a fancy encoder with pretraining, got it.
Share this page's whispers
Short link
https://domwhisper.com/s/6101ece4184aEmbed snippet
<iframe src="https://domwhisper.com/embed/6101ece4184a" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>See what people are discussing on jalammar.github.io
Install DomWhisper to view live whispers as you browse, and join the discussion.
Get the extension