Agentic mid-training, reinforcement learning with reward verification (RLVR), scaling agent environments, interleaved agent reasoning with tools
Popular repositories Loading
-
qqr
qqr PublicForked from Alibaba-NLP/qqr
qqr is an RL training framework for open-ended agents.
Python 1
-
ariadna-training
ariadna-training PublicAriadna: Qwen3.5-4B CPT + SFT + GRPO training pipeline for ML concept curation
Python 1
-
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.


