home/categories/machine-learning/orchestra-research-ai-research-skills-06-post-training-verl-skill-md
machine-learningdata-ai
verl-rl-training
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
maintainer
Orchestra-Research
Actualizado 1/29/2026
Estrellas
6563
Forks
515
quick start
Installation and usage
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
Instalación
$ install --globalskills.sh
Uso
Después de instalarlo, puedes usar este skill ejecutando el siguiente comando en tu terminal:
skills use verl-rl-training