torch-pipeline-parallelism

This skill provides guidance for implementing PyTorch pipeline parallelism for distributed training of large language models. It should be used when implementing pipeline parallel training loops, partitioning transformer models across GPUs, or working with AFAB (All-Forward-All-Backward) scheduling patterns. The skill covers model partitioning, inter-rank communication, gradient flow management, and common pitfalls in distributed training implementations.

查看源碼 machine-learning

maintainer

letta-ai

更新於 1/19/2026

星標

分支

quick start

Installation and usage

安裝

$ install --globalskills.sh

使用

安裝後，您可以通過在終端運行以下命令來使用此技能：

skills use torch-pipeline-parallelism