返回学习路线

AI 学习路线 · 深入

Part 13 · Transformer 深入

站 2A 建立了直觉,这一站把手写代码跑通——从"懂概念"到"写得出"。


Section 1 · 自注意力再复习(从"看"到"讲")

Subsection 1 · 看完整版

  • 打开 https://www.bilibili.com/video/BV1v3411r78R (李宏毅:自注意力机制和 Transformer 详解)看 30:00~60:00(注意力 + 多头部分)。
  • 重点盯:Q、K、V 是什么、注意力分数怎么算。

Subsection 2 · 用一句话给"没学过的人"讲 写下来(照抄,理解着抄):

每个词都生成三个向量:Query(我要找谁)、Key(我是什么)、Value(我的内容)。 分数 = 我的 Q 和别人的 K 做点积 → 谁和我最搭,我就多关注谁的 V。 最后所有 V 加权平均 = 我这个词融合了全句信息的"新表示"。

过关:能对着 Q/K/V 讲一遍 = Section 1 完成。


Section 2 · 用代码算一次注意力(不靠库)

Subsection 1 · 跑手动实现 新建 attn_manual.py:

python
import torchimport torch.nn.functional as F# 3 个词,每个 4 维X = torch.tensor([    [1.0, 0.0, 0.0, 0.0],  # 词A "我"    [0.0, 1.0, 0.0, 0.0],  # 词B "爱"    [0.0, 0.0, 1.0, 0.0],  # 词C "你"])# 简化:直接用 X 当 Q、K、V(真实模型会用可学习权重变换)Q = K = V = X# 1. 算注意力分数:Q 和所有 K 的点积scores = torch.mm(Q, K.T)          # 3x3print("注意力分数:\n", scores)# 2. 归一化成权重(softmax,按行)weights = F.softmax(scores, dim=-1)print("\n注意力权重:\n", weights)# 3. 加权求和 Vout = torch.mm(weights, V)print("\n输出(每个词融合了全句信息):\n", out)
  • 预期看到:一个 3x3 的分数矩阵、权重矩阵、输出矩阵。
  • 这就是注意力的全部内核,真实实现只是加了三组可学习权重 Wq/Wk/Wv 和多头。

过关:能跑通并解释每行干嘛 = Section 2 完成。


Section 3 · 多头注意力(为什么分"头")

Subsection 1 · 概念

多头 = 用多组 Wq/Wk/Wv 同时算多次注意力,每组关注不同关系例:头1 关注"语法搭配",头2 关注"指代关系",头3 关注"数字大小"最后把头拼起来 → 模型看得更全面

Subsection 2 · 跑多头 新建 multi_head.py:

python
import torchimport torch.nn as nnclass MultiHeadAttention(nn.Module):    def __init__(self, d_model=4, n_heads=2):        super().__init__()        self.n_heads = n_heads        self.head_dim = d_model // n_heads        # 每组头用自己的权重        self.wq = nn.Linear(d_model, d_model)        self.wk = nn.Linear(d_model, d_model)        self.wv = nn.Linear(d_model, d_model)        self.out = nn.Linear(d_model, d_model)    def forward(self, x):        B, T, D = x.shape        Q = self.wq(x).view(B, T, self.n_heads, self.head_dim).transpose(1, 2)        K = self.wk(x).view(B, T, self.n_heads, self.head_dim).transpose(1, 2)        V = self.wv(x).view(B, T, self.n_heads, self.head_dim).transpose(1, 2)        scores = Q @ K.transpose(-2, -1) / (self.head_dim ** 0.5)        weights = torch.softmax(scores, dim=-1)        out = (weights @ V).transpose(1, 2).reshape(B, T, D)        return self.out(out)x = torch.randn(1, 3, 4)   # 1句话3个词,4维print(MultiHeadAttention()(x).shape)   # torch.Size([1, 3, 4])
  • 预期看到:torch.Size([1, 3, 4])(输入输出维度一致)。
  • 这已经是真实多头注意力的骨架,跑通它你就"会写"了。

过关:多头代码跑通 = Section 3 完成。


Section 4 · 位置编码(模型怎么知道顺序)

Subsection 1 · 概念

注意力不分前后顺序("我打你"和"你打我"权重一样)→ 需要位置编码位置编码 = 给每个词的位置加一组特殊向量,让模型知道"谁在前谁在后"经典做法:sin/cos 公式(原论文)或 可学习的位置向量(GPT 用)

Subsection 2 · 跑位置编码 新建 pos_enc.py:

python
import torchimport mathdef position_encoding(seq_len=4, d_model=8):    pe = torch.zeros(seq_len, d_model)    for pos in range(seq_len):        for i in range(0, d_model, 2):            pe[pos, i] = math.sin(pos / (10000 ** (i / d_model)))            pe[pos, i + 1] = math.cos(pos / (10000 ** (i / d_model)))    return pepe = position_encoding()print("位置0的编码:", pe[0])print("位置1的编码:", pe[1])print("不同位置编码不同 → 模型能区分顺序")
  • 预期看到:每个位置的编码向量不一样。

过关:能跑通并说"位置编码解决什么" = Section 4 完成。


Section 5 · Encoder vs Decoder(两种模型家族)

Subsection 1 · 记区别

Encoder-only   双向看全句(BERT)→ 适合理解/分类/嵌入Decoder-only   只往前看(GPT)  → 适合生成(所有现代 LLM 都用它)Encoder-Decoder 编码+解码(T5/机器翻译)→ 复杂生成任务
  • 你用的 GPT 系列(含 DeepSeek)都是 Decoder-only。

Subsection 2 · 掩码自注意力(为什么"不能看未来") 新建 mask_attn.py:

python
import torchscores = torch.rand(4, 4)   # 4个词的注意力分数# 掩码:只允许看自己和之前(上三角置为 -inf)mask = torch.triu(torch.full((4, 4), float("-inf")), diagonal=1)masked = scores + maskprint("掩码后(下三角可见,上三角=无效):\n", torch.round(masked, decimals=2))
  • 预期看到:右上三角全是 -inf。
  • Decoder 生成时就是"一次只能看前面的词",保证生成不"作弊"。

过关:能说 Encoder/Decoder 区别 + 掩码作用 = Section 5 完成。


Section 6 · 解码策略(模型怎么选下一个词)

Subsection 1 · 跑三种解码 新建 decoding.py:

python
import torchimport torch.nn.functional as F# 假设模型给出 5 个候选词的分数logits = torch.tensor([2.0, 1.5, 0.1, -1.0, 3.5])# 1. 贪心:直接取分数最高的print("贪心:", logits.argmax().item())# 2. 采样(temperature):分数缩放后随机抽probs = F.softmax(logits / 0.8, dim=-1)sample = torch.multinomial(probs, 1).item()print("采样(温度0.8):", sample, "| 概率分布:", [round(p, 2) for p in probs])# 3. Top-P(核采样):只从累计概率到 p 的词里抽def top_p_sampling(logits, p=0.9):    probs = F.softmax(logits, dim=-1)    sorted_p, idx = torch.sort(probs, descending=True)    cum = torch.cumsum(sorted_p, dim=-1)    mask = cum - sorted_p > p    sorted_p[mask] = 0    probs = sorted_p / sorted_p.sum()    return torch.multinomial(probs, 1)print("Top-P 采样:", top_p_sampling(logits).item())
  • 预期看到:三种方法分别选出词。
  • 面试常问:贪心稳但死板,采样多样但可能跑偏;生产常用 temperature + top_p 组合。

过关:能跑通并说出三方法取舍 = Section 6 完成。


Section 7 · 手写完整 Self-Attention 层

Subsection 1 · 跑完整代码 新建 my_attention.py(把 Section 2/3 合并成规范实现):

python
import torchimport torch.nn as nnimport torch.nn.functional as Fclass SelfAttention(nn.Module):    def __init__(self, d_model=8):        super().__init__()        self.wq = nn.Linear(d_model, d_model)        self.wk = nn.Linear(d_model, d_model)        self.wv = nn.Linear(d_model, d_model)    def forward(self, x):        Q, K, V = self.wq(x), self.wk(x), self.wv(x)        d = Q.shape[-1]        scores = torch.mm(Q, K.T) / (d ** 0.5)        weights = F.softmax(scores, dim=-1)        return torch.mm(weights, V)x = torch.randn(4, 8)          # 4个词,8维out = SelfAttention()(x)print("输出形状:", out.shape)   # torch.Size([4, 8])
  • 预期看到:torch.Size([4, 8])。

过关:能独立写出(或对照默写)这个 SelfAttention = Section 7 完成。


Section 8 · 手写一个迷你 GPT(Decoder 骨架)

Subsection 1 · 跑完整迷你 GPT 新建 mini_gpt.py:

python
import torchimport torch.nn as nnclass MiniGPT(nn.Module):    def __init__(self, vocab=100, d_model=32, n_heads=4, n_layers=2):        super().__init__()        self.embed = nn.Embedding(vocab, d_model)        self.pos = nn.Parameter(torch.randn(1, 50, d_model))        self.blocks = nn.ModuleList([            nn.TransformerEncoderLayer(d_model, n_heads, dim_feedforward=128, batch_first=True)            for _ in range(n_layers)        ])        self.head = nn.Linear(d_model, vocab)    def forward(self, x):        x = self.embed(x) + self.pos[:, :x.shape[1]]        for b in self.blocks:            x = b(x)        return self.head(x)   # 每个位置预测下一个词的分数model = MiniGPT()x = torch.randint(0, 100, (2, 10))     # 2句话,每句10个词logits = model(x)print("输出形状:", logits.shape)         # torch.Size([2, 10, 100])print("第一句第一个位置预测的词表分数前3:", logits[0, 0, :3].detach().numpy())
  • 预期看到:torch.Size([2, 10, 100])(每句话每位置给词表打分)。
  • 这就是 GPT 的极简骨架:嵌入 + 位置 + N 层 Transformer + 预测词表分数。真实的 GPT 只是把它放大 + 用海量数据训练。

过关:迷你 GPT 跑通 = Section 8 完成。


Section 9 · 用 HuggingFace 做同样的生成(对照)

Subsection 1 · 跑 HF 生成 新建 hf_generate.py:

python
from transformers import pipeline# 加载一个很小的中文生成模型gen = pipeline("text-generation", model="uer/gpt2-chinese-cluecorpussmall")result = gen("今天天气", max_new_tokens=20, do_sample=True, temperature=0.8)print(result[0]["generated_text"])
  • 预期看到:模型续写出一小段中文。
  • 对比 Section 8:HF 把"嵌入+层+head+采样"全封装好了。你现在明白它内部在干嘛了。

过关:能跑通并说"HF 内部就是我们手写的东西" = Section 9 完成。


Section 10 · 站 11 验收

勾选

  • 能对着 Q/K/V 讲一遍注意力
  • 手动注意力代码跑通(Section 2)
  • 多头注意力跑通(Section 3)
  • 位置编码跑通,说清作用(Section 4)
  • 能说 Encoder/Decoder 区别 + 掩码(Section 5)
  • 三种解码策略跑通并说取舍(Section 6)
  • 手写 SelfAttention 跑通(Section 7)
  • 手写迷你 GPT 跑通(Section 8)
  • HF 生成对照跑通(Section 9)

写 300 字Part 11 总结:从概念到代码,你现在对 Transformer 的理解。

全勾选 = Part 13 通过 → 进入下一站 Part 14 · 开源模型与私有化部署,20 天)。