home/categories/machine-learning/orchestra-research-ai-research-skills-06-post-training-verl-skill-md
machine-learningdata-ai
verl-rl-training
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
maintainer
Orchestra-Research
更新於 1/29/2026
星標
6563
分支
515
quick start
Installation and usage
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
安裝
$ install --globalskills.sh
使用
安裝後,您可以通過在終端運行以下命令來使用此技能:
skills use verl-rl-training