AI 学习路线 · 深入
Part 13 · Transformer 深入
站 2A 建立了直觉,这一站把手写代码跑通——从"懂概念"到"写得出"。
Section 1 · 自注意力再复习(从"看"到"讲")
Subsection 1 · 看完整版
- 打开 https://www.bilibili.com/video/BV1v3411r78R (李宏毅:自注意力机制和 Transformer 详解)看 30:00~60:00(注意力 + 多头部分)。
- 重点盯:Q、K、V 是什么、注意力分数怎么算。
Subsection 2 · 用一句话给"没学过的人"讲 写下来(照抄,理解着抄):
每个词都生成三个向量:Query(我要找谁)、Key(我是什么)、Value(我的内容)。 分数 = 我的 Q 和别人的 K 做点积 → 谁和我最搭,我就多关注谁的 V。 最后所有 V 加权平均 = 我这个词融合了全句信息的"新表示"。
过关:能对着 Q/K/V 讲一遍 = Section 1 完成。
Section 2 · 用代码算一次注意力(不靠库)
Subsection 1 · 跑手动实现
新建 attn_manual.py:
import torchimport torch.nn.functional as F# 3 个词,每个 4 维X = torch.tensor([ [1.0, 0.0, 0.0, 0.0], # 词A "我" [0.0, 1.0, 0.0, 0.0], # 词B "爱" [0.0, 0.0, 1.0, 0.0], # 词C "你"])# 简化:直接用 X 当 Q、K、V(真实模型会用可学习权重变换)Q = K = V = X# 1. 算注意力分数:Q 和所有 K 的点积scores = torch.mm(Q, K.T) # 3x3print("注意力分数:\n", scores)# 2. 归一化成权重(softmax,按行)weights = F.softmax(scores, dim=-1)print("\n注意力权重:\n", weights)# 3. 加权求和 Vout = torch.mm(weights, V)print("\n输出(每个词融合了全句信息):\n", out)- 预期看到:一个 3x3 的分数矩阵、权重矩阵、输出矩阵。
- 这就是注意力的全部内核,真实实现只是加了三组可学习权重 Wq/Wk/Wv 和多头。
过关:能跑通并解释每行干嘛 = Section 2 完成。
Section 3 · 多头注意力(为什么分"头")
Subsection 1 · 概念
多头 = 用多组 Wq/Wk/Wv 同时算多次注意力,每组关注不同关系例:头1 关注"语法搭配",头2 关注"指代关系",头3 关注"数字大小"最后把头拼起来 → 模型看得更全面Subsection 2 · 跑多头
新建 multi_head.py:
import torchimport torch.nn as nnclass MultiHeadAttention(nn.Module): def __init__(self, d_model=4, n_heads=2): super().__init__() self.n_heads = n_heads self.head_dim = d_model // n_heads # 每组头用自己的权重 self.wq = nn.Linear(d_model, d_model) self.wk = nn.Linear(d_model, d_model) self.wv = nn.Linear(d_model, d_model) self.out = nn.Linear(d_model, d_model) def forward(self, x): B, T, D = x.shape Q = self.wq(x).view(B, T, self.n_heads, self.head_dim).transpose(1, 2) K = self.wk(x).view(B, T, self.n_heads, self.head_dim).transpose(1, 2) V = self.wv(x).view(B, T, self.n_heads, self.head_dim).transpose(1, 2) scores = Q @ K.transpose(-2, -1) / (self.head_dim ** 0.5) weights = torch.softmax(scores, dim=-1) out = (weights @ V).transpose(1, 2).reshape(B, T, D) return self.out(out)x = torch.randn(1, 3, 4) # 1句话3个词,4维print(MultiHeadAttention()(x).shape) # torch.Size([1, 3, 4])- 预期看到:
torch.Size([1, 3, 4])(输入输出维度一致)。 - 这已经是真实多头注意力的骨架,跑通它你就"会写"了。
过关:多头代码跑通 = Section 3 完成。
Section 4 · 位置编码(模型怎么知道顺序)
Subsection 1 · 概念
注意力不分前后顺序("我打你"和"你打我"权重一样)→ 需要位置编码位置编码 = 给每个词的位置加一组特殊向量,让模型知道"谁在前谁在后"经典做法:sin/cos 公式(原论文)或 可学习的位置向量(GPT 用)Subsection 2 · 跑位置编码
新建 pos_enc.py:
import torchimport mathdef position_encoding(seq_len=4, d_model=8): pe = torch.zeros(seq_len, d_model) for pos in range(seq_len): for i in range(0, d_model, 2): pe[pos, i] = math.sin(pos / (10000 ** (i / d_model))) pe[pos, i + 1] = math.cos(pos / (10000 ** (i / d_model))) return pepe = position_encoding()print("位置0的编码:", pe[0])print("位置1的编码:", pe[1])print("不同位置编码不同 → 模型能区分顺序")- 预期看到:每个位置的编码向量不一样。
过关:能跑通并说"位置编码解决什么" = Section 4 完成。
Section 5 · Encoder vs Decoder(两种模型家族)
Subsection 1 · 记区别
Encoder-only 双向看全句(BERT)→ 适合理解/分类/嵌入Decoder-only 只往前看(GPT) → 适合生成(所有现代 LLM 都用它)Encoder-Decoder 编码+解码(T5/机器翻译)→ 复杂生成任务- 你用的 GPT 系列(含 DeepSeek)都是 Decoder-only。
Subsection 2 · 掩码自注意力(为什么"不能看未来")
新建 mask_attn.py:
import torchscores = torch.rand(4, 4) # 4个词的注意力分数# 掩码:只允许看自己和之前(上三角置为 -inf)mask = torch.triu(torch.full((4, 4), float("-inf")), diagonal=1)masked = scores + maskprint("掩码后(下三角可见,上三角=无效):\n", torch.round(masked, decimals=2))- 预期看到:右上三角全是 -inf。
- Decoder 生成时就是"一次只能看前面的词",保证生成不"作弊"。
过关:能说 Encoder/Decoder 区别 + 掩码作用 = Section 5 完成。
Section 6 · 解码策略(模型怎么选下一个词)
Subsection 1 · 跑三种解码
新建 decoding.py:
import torchimport torch.nn.functional as F# 假设模型给出 5 个候选词的分数logits = torch.tensor([2.0, 1.5, 0.1, -1.0, 3.5])# 1. 贪心:直接取分数最高的print("贪心:", logits.argmax().item())# 2. 采样(temperature):分数缩放后随机抽probs = F.softmax(logits / 0.8, dim=-1)sample = torch.multinomial(probs, 1).item()print("采样(温度0.8):", sample, "| 概率分布:", [round(p, 2) for p in probs])# 3. Top-P(核采样):只从累计概率到 p 的词里抽def top_p_sampling(logits, p=0.9): probs = F.softmax(logits, dim=-1) sorted_p, idx = torch.sort(probs, descending=True) cum = torch.cumsum(sorted_p, dim=-1) mask = cum - sorted_p > p sorted_p[mask] = 0 probs = sorted_p / sorted_p.sum() return torch.multinomial(probs, 1)print("Top-P 采样:", top_p_sampling(logits).item())- 预期看到:三种方法分别选出词。
- 面试常问:贪心稳但死板,采样多样但可能跑偏;生产常用 temperature + top_p 组合。
过关:能跑通并说出三方法取舍 = Section 6 完成。
Section 7 · 手写完整 Self-Attention 层
Subsection 1 · 跑完整代码
新建 my_attention.py(把 Section 2/3 合并成规范实现):
import torchimport torch.nn as nnimport torch.nn.functional as Fclass SelfAttention(nn.Module): def __init__(self, d_model=8): super().__init__() self.wq = nn.Linear(d_model, d_model) self.wk = nn.Linear(d_model, d_model) self.wv = nn.Linear(d_model, d_model) def forward(self, x): Q, K, V = self.wq(x), self.wk(x), self.wv(x) d = Q.shape[-1] scores = torch.mm(Q, K.T) / (d ** 0.5) weights = F.softmax(scores, dim=-1) return torch.mm(weights, V)x = torch.randn(4, 8) # 4个词,8维out = SelfAttention()(x)print("输出形状:", out.shape) # torch.Size([4, 8])- 预期看到:
torch.Size([4, 8])。
过关:能独立写出(或对照默写)这个 SelfAttention = Section 7 完成。
Section 8 · 手写一个迷你 GPT(Decoder 骨架)
Subsection 1 · 跑完整迷你 GPT
新建 mini_gpt.py:
import torchimport torch.nn as nnclass MiniGPT(nn.Module): def __init__(self, vocab=100, d_model=32, n_heads=4, n_layers=2): super().__init__() self.embed = nn.Embedding(vocab, d_model) self.pos = nn.Parameter(torch.randn(1, 50, d_model)) self.blocks = nn.ModuleList([ nn.TransformerEncoderLayer(d_model, n_heads, dim_feedforward=128, batch_first=True) for _ in range(n_layers) ]) self.head = nn.Linear(d_model, vocab) def forward(self, x): x = self.embed(x) + self.pos[:, :x.shape[1]] for b in self.blocks: x = b(x) return self.head(x) # 每个位置预测下一个词的分数model = MiniGPT()x = torch.randint(0, 100, (2, 10)) # 2句话,每句10个词logits = model(x)print("输出形状:", logits.shape) # torch.Size([2, 10, 100])print("第一句第一个位置预测的词表分数前3:", logits[0, 0, :3].detach().numpy())- 预期看到:
torch.Size([2, 10, 100])(每句话每位置给词表打分)。 - 这就是 GPT 的极简骨架:嵌入 + 位置 + N 层 Transformer + 预测词表分数。真实的 GPT 只是把它放大 + 用海量数据训练。
过关:迷你 GPT 跑通 = Section 8 完成。
Section 9 · 用 HuggingFace 做同样的生成(对照)
Subsection 1 · 跑 HF 生成
新建 hf_generate.py:
from transformers import pipeline# 加载一个很小的中文生成模型gen = pipeline("text-generation", model="uer/gpt2-chinese-cluecorpussmall")result = gen("今天天气", max_new_tokens=20, do_sample=True, temperature=0.8)print(result[0]["generated_text"])- 预期看到:模型续写出一小段中文。
- 对比 Section 8:HF 把"嵌入+层+head+采样"全封装好了。你现在明白它内部在干嘛了。
过关:能跑通并说"HF 内部就是我们手写的东西" = Section 9 完成。
Section 10 · 站 11 验收
勾选
- 能对着 Q/K/V 讲一遍注意力
- 手动注意力代码跑通(Section 2)
- 多头注意力跑通(Section 3)
- 位置编码跑通,说清作用(Section 4)
- 能说 Encoder/Decoder 区别 + 掩码(Section 5)
- 三种解码策略跑通并说取舍(Section 6)
- 手写 SelfAttention 跑通(Section 7)
- 手写迷你 GPT 跑通(Section 8)
- HF 生成对照跑通(Section 9)
写 300 字Part 11 总结:从概念到代码,你现在对 Transformer 的理解。
全勾选 = Part 13 通过 → 进入下一站 Part 14 · 开源模型与私有化部署,20 天)。
