home/categories/llm-ai/atrawog-bazzite-ai-plugins-bazzite-ai-jupyter-skills-grpo-skill-md
llm-aidata-ai

grpo

Group Relative Policy Optimization for reinforcement learning from human feedback. Covers GRPOTrainer, reward function design, policy optimization, and KL divergence constraints for stable RLHF training. Includes thinking-aware reward patterns.

atrawog
maintainer
atrawog
更新日 1/12/2026
スター
0
フォーク
0
quick start

Installation and usage

Group Relative Policy Optimization for reinforcement learning from human feedback. Covers GRPOTrainer, reward function design, policy optimization, and KL divergence constraints for stable RLHF training. Includes thinking-aware reward patterns.

インストール
$ install --globalskills.sh
使い方

インストール後、ターミナルで以下のコマンドを実行してこのスキルを使用できます:

skills use grpo