Skip to content

About

Official codebase for the paper "Sharpening Tax in Post-Training"

Topics

Resources

Code of conduct

Contributing

Stars

25 stars

Watchers

0 watching

Forks

Repository files navigation

Sharpening Tax in Post-Training

Changdae Oh1,2,*, Qi Zeng1, Qi Qi1, Andrey Zhmoginov, Deren Lei1, Yun He1, Hoang Phan1,3,*,
Hangoo Kang4, Azalia Mirhoseini4, Sharon Li2

1Meta Superintelligence Labs   2University of Wisconsin–Madison   3New York University   4Stanford University
*Work done at Meta

Project page arXiv Twitter

Sharpening Tax in Post-Training

Does RL post-training give a language-model agent new capabilities, or does it merely sharpen behaviors its base model already has? We compare 14 open base/post-trained checkpoint pairs from four model families on three agentic benchmarks (BFCL v4 multi-turn, WebShop and ACEBench), with 128 rollouts per task.

  • Base models are capable agents. With a light text harness, pre-trained base models solve agentic tasks and, given enough rollouts, often reach more tasks (pass@k) than their post-trained counterparts. Post-training pushes tasks toward two extremes, always solved or never solved, trading solution coverage for sampling efficiency and consistency.
  • Sharpening Tax. A diagnostic metric for the test-time scalability lost in post-training. It is positive in most of the 42 model-benchmark cases, can be estimated from a few rollouts, and correlates well with other metrics.
  • PTGS. Posterior-tempered group sampling sets each prompt's rollout temperature from a Bayesian estimate of its difficulty. Applied during RL training, it pays a smaller tax than the fixed-temperature baseline while also improving single-shot accuracy.

Getting started

This repository provides minimal working examples for understanding the project; it is not a reproduction package. Benchmark checkouts, model serving environments and RL training infrastructure are left to you. The metrics and PTGS run on CPU with numpy:

pip install -r requirements.txt
python -m ptgs                    # PTGS tempering rule, target ramp and one prompt's trajectory
python examples/ptgs_toy.py       # PTGS vs. fixed-temperature RL on a toy task (~2 s)
python -m pytest                  # 70 tests

Compute the Sharpening Tax from the example per-task outcomes:

from sharpening_tax.metrics import Pair
from sharpening_tax.outcomes import load_table

cells = load_table("data/outcomes_42cells.csv.gz")  # {(benchmark, pair, arm): {task_id: (c, n)}}
tax = Pair(cells["webshop", "gemma-4-31B", "base"], cells["webshop", "gemma-4-31B", "rl"]).tax([128])
print(tax[128]["tax_A"], tax[128]["tax_S"])          # 4.49 0.036

See sharpening_tax for collecting your own rollouts with the text harness, and ptgs for adding PTGS to an RL trainer.

Repository

sharpening_tax Text harness for base models, rollout runners for BFCL / WebShop / ACEBench, and the pass@k, pass^k and Sharpening Tax metrics
ptgs PTGS controller and the places it touches a group-based RL trainer
data Example per-task success counts (14 base / post-trained pairs × 3 benchmarks, 128 rollouts per task)
configs Example RL configuration with PTGS (a record, not a launcher)
scripts vLLM serving recipe for each model family
examples PTGS on a toy task

Citation

If you find our work helpful, we would appreciate it if you could cite our paper:

@article{oh2026sharpening,
  title   = {Sharpening Tax in Post-Training},
  author  = {Oh, Changdae and Zeng, Qi and Qi, Qi and Zhmoginov, Andrey and Lei, Deren and
             He, Yun and Phan, Hoang and Kang, Hangoo and Mirhoseini, Azalia and Li, Sharon},
  journal = {arXiv preprint arXiv:2610.01509},
  year    = {2026}
}

License

This project is licensed under CC BY-NC 4.0. See also THIRD_PARTY_NOTICES.md.

About

Official codebase for the paper "Sharpening Tax in Post-Training"

Topics

Resources

Code of conduct

Contributing

Stars

25 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages