All Tools
V
Fine-tuningFreeOpen Source
VERL
RL post-training framework for LLMs
Apache-2.0
ABOUT
RL post-training for LLMs usually means stitching trainers, rollouts, and GPU schedulers by hand. verl (HybridFlow) is a flexible RL framework with production recipes for PPO, GRPO, and more, so teams align models at scale without building a custom RL infrastructure.
INSTALL
pip install verlINTEGRATION GUIDE
1. Run PPO or GRPO post-training on an open LLM with vLLM rollouts
2. Scale RLHF jobs across GPUs without a custom trainer
3. Reproduce HybridFlow recipes for alignment and reasoning models
TAGS
pythonrlhfppogrpopost-trainingllmvllm