Tensor Parallelism, TP,是把单层内部的张量计算切分到多张 GPU 上。它不同于 Data Parallel 的“每张 GPU 都有完整模型”,也不同于 Pipeline Parallel 的“不同 GPU 负责不同层”。TP 关注的是同一层里的矩阵乘法、attention heads 或 MLP hidden dimension 如何拆分。
TP 主要解决单层参数和激活太大、单卡无法高效计算的问题,是大规模预训练和 Megatron-style 训练的核心组件。
线性层切分
Transformer 中大量计算来自线性层:
其中 ,。Tensor parallel 可以按列或按行切分 。
Column Parallel
按输出维度切分:
每张 GPU 计算:
最后 。如果下一层可以消费分片输出,就不必立即 AllGather。
Row Parallel
按输入维度切分:
每张 GPU 计算 partial output:
然后通过 AllReduce 求和:
Megatron-LM 的 MLP 常用 column-parallel + row-parallel 配对,以减少不必要通信。
Attention 与 MLP 切分
在 Transformer 中:
- attention 的 Q/K/V projection 可以按 heads 或 hidden dimension 切分;
- attention output projection 可以用 row parallel;
- MLP 的 up/gate projection 常用 column parallel;
- MLP 的 down projection 常用 row parallel。
Megatron-LM 的最小通信 Pattern
Megatron-LM 给出了一个经典的 Tensor Parallel 组合。它的目标不是让每一步都产生完整复制的 hidden states,而是让中间结果尽量保持分片,只在需要合并 partial output 时通信。
MLP:
Column Parallel Linear
-> local GeLU
Row Parallel Linear
-> AllReduce partial output
Self-Attention:
Column Parallel Q/K/V
-> local attention heads
Row Parallel output projection
-> AllReduce partial output因此,一个 Transformer layer 的主要通信可以概括为:
- forward:MLP 的 Row Parallel 输出 1 次 AllReduce,attention output projection 1 次 AllReduce;
- backward:对应位置各 1 次 AllReduce;
- 每个 layer 合计 2 次 forward AllReduce 和 2 次 backward AllReduce。
论文还使用两个 autograd-aware operator 表达 forward/backward 的通信位置:g 在 forward 执行 AllReduce、backward 保持 identity;f 在 forward 保持 identity、backward 执行 AllReduce。这个设计说明通信语义可以嵌入计算图,而不必由训练循环手工拼接。
Vocabulary Parallel 处理 input embedding 和 output projection。input embedding 沿 vocabulary dimension 分片;output projection 则将 vocabulary-sized projection 与 cross-entropy 融合,避免为了计算标量 loss 而 AllGather 完整 logits。
Vocabulary-Parallel Cross Entropy
当 LM Head 沿 vocabulary dimension 切分时,每个 TP rank 只持有形状约为 的 local logits。Cross Entropy 只需要 global max、global sum-exp 和真实 target 对应的 logit,因此可以用少量 collective 直接得到 loss:
local logits
-> AllReduce(MAX): global max
-> local exp sum
-> AllReduce(SUM): global sum-exp
-> target owner 提供 target logit
-> AllReduce(SUM): global target logit
-> cross entropyBackward 在每个 vocabulary shard 上计算 local ,不需要恢复完整 vocabulary gradient。这样避免了对 logits 的 AllGather,尤其适合大 vocabulary 和长 sequence。Megatron-Core 的实现入口是 tensor_parallel/cross_entropy.py;fused variant 还会合并部分计算和 collective,减少中间结果与 kernel launch。
这里仍需注意 loss normalization 属于另一维语义:Vocabulary Parallel 解决 维分片,DP/CP/packing 则可能改变有效 token 如何跨 ranks 聚合。得到每个 token 的正确 CE,不代表 batch-level reduction 已经自动正确。
读这部分源码或实现时,应该始终同时追踪三件事:当前 tensor 沿哪个维度分片、当前通信属于哪个 TP process group、通信完成后输出是 replicated 还是仍然 sharded。
这种设计让每个 GPU 只处理部分 heads 或部分 intermediate dimension,降低单卡计算和参数压力。
通信模式
TP 的通信发生在层内部,通常比 data parallel 更频繁。常见通信:
- AllReduce:合并 row-parallel partial outputs;
- AllGather:收集分片 activation;
- ReduceScatter:合并并分发结果;
- broadcast / scatter:准备输入分片。
因为 TP 通信频繁,最好放在同节点高速互联 GPU 上,例如 NVLink。跨节点做大 TP 往往通信开销很高。
Sequence Parallel
Sequence Parallelism 常与 tensor parallel 配合,把某些 activation 沿 sequence dimension 切分,减少 activation memory。它通常作用于 LayerNorm、Dropout、residual 等 token-wise 模块附近:这些模块按 token position 独立处理,不需要同时持有完整 sequence,但 LayerNorm / RMSNorm 仍需要该 token 的完整 hidden vector 或等价的正确聚合实现。
直观上:
- TP 切 hidden / heads / MLP dimension;
- sequence parallel 切 sequence dimension;
- 二者配合降低单卡 activation 和通信压力。
具体实现依赖框架,Megatron-Core 中 sequence parallel 是重要优化。需要注意,sequence parallel 不等于 Context Parallel:前者主要切分部分 activation layout,后者直接面向长上下文 attention 的 context / K/V 跨 shard 计算。
与其他并行的关系
TP 解决单层太大的问题;PP 解决层数/模型深度分布问题;DP/FSDP/ZeRO 解决数据并行冗余状态问题。
常见组合:
如果一个 1024 GPU 训练使用 ,则 。每个并行维度有不同通信组。
适用场景
TP 适合:
- 单层矩阵乘法太大;
- attention heads 很多;
- hidden size / MLP intermediate size 很大;
- 希望提高单 step 计算并行度;
- 大规模预训练中同节点 GPU 高速互联充足。
TP 不适合过小模型或低带宽跨节点环境,因为通信可能超过计算收益。
常见失败模式
- TP degree 过大:每卡矩阵太小,kernel efficiency 下降。
- 跨节点 TP:频繁通信导致严重瓶颈。
- 通信未重叠:AllReduce/AllGather 等待时间变长。
- 与 PP/DP 组网不匹配:GPU 拓扑没有按通信密集度布置。
- 实现细节错误:分片权重、checkpoint 合并和随机种子处理复杂。