distributed-llm-pretraining-torchtitan

Name: distributed-llm-pretraining-torchtitan
Author: math-inc

Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.

स्रोत देखें framework-internals

maintainer

math-inc

अपडेट किया गया 3/19/2026

स्टार

1165

फोर्क

quick start

Installation and usage

इंस्टॉलेशन

$ install --globalskills.sh

उपयोग

इंस्टॉल करने के बाद, आप टर्मिनल में यह कमांड चलाकर इस स्किल का उपयोग कर सकते हैं:

skills use distributed-llm-pretraining-torchtitan