Changdae Oh1,2,*,
Qi Zeng1,
Qi Qi1,
Andrey Zhmoginov,
Deren Lei1,
Yun He1,
Hoang Phan1,3,*,
Hangoo Kang4,
Azalia Mirhoseini4,
Sharon Li2
1Meta Superintelligence Labs 2University of Wisconsin–Madison 3New York University 4Stanford University
*Work done at Meta
Does RL post-training give a language-model agent new capabilities, or does it merely sharpen behaviors its base model already has? We compare 14 open base/post-trained checkpoint pairs from four model families on three agentic benchmarks (BFCL v4 multi-turn, WebShop and ACEBench), with 128 rollouts per task.
- Base models are capable agents. With a light text harness, pre-trained base models solve agentic tasks and, given enough rollouts, often reach more tasks (pass@k) than their post-trained counterparts. Post-training pushes tasks toward two extremes, always solved or never solved, trading solution coverage for sampling efficiency and consistency.
- Sharpening Tax. A diagnostic metric for the test-time scalability lost in post-training. It is positive in most of the 42 model-benchmark cases, can be estimated from a few rollouts, and correlates well with other metrics.
- PTGS. Posterior-tempered group sampling sets each prompt's rollout temperature from a Bayesian estimate of its difficulty. Applied during RL training, it pays a smaller tax than the fixed-temperature baseline while also improving single-shot accuracy.
This repository provides minimal working examples for understanding the project; it is not a reproduction package. Benchmark checkouts, model serving environments and RL training infrastructure are left to you. The metrics and PTGS run on CPU with numpy:
pip install -r requirements.txt
python -m ptgs # PTGS tempering rule, target ramp and one prompt's trajectory
python examples/ptgs_toy.py # PTGS vs. fixed-temperature RL on a toy task (~2 s)
python -m pytest # 70 testsCompute the Sharpening Tax from the example per-task outcomes:
from sharpening_tax.metrics import Pair
from sharpening_tax.outcomes import load_table
cells = load_table("data/outcomes_42cells.csv.gz") # {(benchmark, pair, arm): {task_id: (c, n)}}
tax = Pair(cells["webshop", "gemma-4-31B", "base"], cells["webshop", "gemma-4-31B", "rl"]).tax([128])
print(tax[128]["tax_A"], tax[128]["tax_S"]) # 4.49 0.036See sharpening_tax for collecting your own rollouts with the text harness,
and ptgs for adding PTGS to an RL trainer.
sharpening_tax |
Text harness for base models, rollout runners for BFCL / WebShop / ACEBench, and the pass@k, pass^k and Sharpening Tax metrics |
ptgs |
PTGS controller and the places it touches a group-based RL trainer |
data |
Example per-task success counts (14 base / post-trained pairs × 3 benchmarks, 128 rollouts per task) |
configs |
Example RL configuration with PTGS (a record, not a launcher) |
scripts |
vLLM serving recipe for each model family |
examples |
PTGS on a toy task |
If you find our work helpful, we would appreciate it if you could cite our paper:
@article{oh2026sharpening,
title = {Sharpening Tax in Post-Training},
author = {Oh, Changdae and Zeng, Qi and Qi, Qi and Zhmoginov, Andrey and Lei, Deren and
He, Yun and Phan, Hoang and Kang, Hangoo and Mirhoseini, Azalia and Li, Sharon},
journal = {arXiv preprint arXiv:2610.01509},
year = {2026}
}This project is licensed under CC BY-NC 4.0. See also THIRD_PARTY_NOTICES.md.
