Attention Is All You Need
Attention Is All You Need
Section titled “Attention Is All You Need”作者:Vaswani et al. (Google Brain, 2017)
引用:超过 100,000 次
抛弃 RNN 和 CNN,完全基于 Self-Attention 构建模型。三大优势:
- 并行计算:不使用循环,所有位置同时计算
- 短路径:任意两个位置只需 O(1) 步
- 可解释:注意力权重可视化
| 组件 | 值 |
|---|---|
| 512 | |
| 头数 | 8 |
| 64 | |
| FFN 隐藏层 | 2048 |
| Encoder/Decoder 层数 | 6 |
| Dropout | 0.1 |
| 优化器 | Adam () |
| 学习率 | warmup 4000 步后衰减 |
WMT 2014 英德翻译:
| 模型 | BLEU | 训练成本 |
|---|---|---|
| ConvS2S (之前最佳) | 26.4 | 高 |
| Transformer (base) | 27.3 | 1/4 |
| Transformer (big) | 28.4 | 1/3 |
论文原文名言
Section titled “论文原文名言”“We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.”
- Transformer 教程 — 论文的完整实现
- 注意力机制 — 核心机制详解