jalammar.github.io favicon

The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning) – Jay Alammar – Visualizing machine learning one concept at a time.

The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning) – Jay Alammar – Visualizing machine learning one concept at a time.

#12
13
公开标注数
7
参与人数
2026-07-14 15:09:18
首次 Whisper

讨论活跃度

jalammar.github.io 近 17 周的公开 Whisper

5 活跃天数

最新公开 Whisper

RSS
划选高亮2026-07-20 13:05:37
原文高亮摘录
Every third block starting from 9 is a RETRO block
Whisper 随想笔记
I guess the smaller models just start earlier, makes sense for less depth.
划选高亮2026-07-20 12:56:37
原文高亮摘录
Every third block starting from 9 is a RETRO block
Whisper 随想笔记
So every third block after 9, got it. But why start at 9, seems kinda random.
划选高亮2026-07-20 10:02:37
原文高亮摘录
the model can be much smaller. A neural database aids it with retrieving factual information
Whisper 随想笔记
makes sense—why memorize when you can just look it up?
划选高亮2026-07-20 09:53:37
原文高亮摘录
the model can be much smaller. A neural database aids it with retrieving factual information
Whisper 随想笔记
so it's like having a cheat sheet, but for facts.
划选高亮2026-07-20 09:44:37
原文高亮摘录
the model can be much smaller. A neural database aids it with retrieving factual information
Whisper 随想笔记
smaller models are the future, but what if the database goes down?
划选高亮2026-07-19 13:05:31
原文高亮摘录
ELMo actually goes a step further and trains a bi-directional LSTM
Whisper 随想笔记
But doesn't training both directions just double the compute for marginal gains?
划选高亮2026-07-19 12:56:31
原文高亮摘录
ELMo actually goes a step further and trains a bi-directional LSTM
Whisper 随想笔记
So that's why ELMo feels smarter than the old uni-directional models.
划选高亮2026-07-19 10:02:31
原文高亮摘录
BERT is basically a trained Transformer Encoder stack.
Whisper 随想笔记
Wait, so no decoder at all? That changes everything for me.
划选高亮2026-07-19 09:53:31
原文高亮摘录
BERT is basically a trained Transformer Encoder stack.
Whisper 随想笔记
That post on Transformer really is a must-read first.
划选高亮2026-07-19 09:44:31
原文高亮摘录
BERT is basically a trained Transformer Encoder stack.
Whisper 随想笔记
So it's just a fancy encoder with pretraining, got it.
划选高亮2026-07-17 03:09:18
原文高亮摘录
allows you to think about problems at a higher level
Whisper 随想笔记
Yeah, vectorized thinking really changes how you approach problems.
划选高亮2026-07-15 21:09:18
原文高亮摘录
that new sequence becomes the input to the model in its next step
Whisper 随想笔记
this is the core trick behind GPT-2, love how simple it sounds once explained
划选高亮2026-07-14 15:09:18
原文高亮摘录
What does “it” in this sentence refer to?
Whisper 随想笔记
This is such a relatable example for explaining attention!

看看大家在 jalammar.github.io 上划了什么重点

安装 DomWhisper,浏览网页时实时查看 whisper,也可以加入讨论。

获取插件