English | 中文
An experimental research project exploring whether structural modifications to the Transformer architecture can give rise to genuine cognitive capabilities: rule discovery, compositional reasoning, and analogical transfer.
This is honest, open-ended research. We report both successes and failures.
Standard Transformers are extraordinarily capable pattern-matchers, but arguably lack the structured inductive biases required for true reasoning and abstraction. This project investigates whether adding specific structural primitives — iterative computation, cross-example comparison, and associative memory — can close that gap.
We are not claiming to have solved reasoning. We are asking: can targeted architectural changes produce qualitatively different generalization behavior?
Built on a GPT-2-style backbone with three structural additions:
[Token Embedding + Positional Encoding]
↓
[Encoder Layers × 3] — standard Transformer blocks
↓
[CrossExampleComparator] — extracts invariants across demonstrations
↓
[AdaptiveIterativeCore] — PonderNet-style adaptive depth reasoning
↓
[Decoder Layers × 3] — with AnalogicalMemory key-value injection
↓
[LM Head]
| Component | Purpose | Key Design Choice |
|---|---|---|
AdaptiveIterativeCore |
Shared-weight Transformer blocks iterated N times, where N is learned (PonderNet) | Model decides its own "thinking depth" |
CrossExampleComparator |
Computes pairwise differences between windows; extracts invariant transformation patterns | Explicit difference computation, not just similarity |
AnalogicalMemory |
Input-modulated key-value memory — base templates + input-conditioned offsets | Dynamic, not fixed parameters |
Also includes a MiniCALMAutoencoder and EnergyHead (inspired by CALM, arXiv:2510.27688) for continuous-vector prediction mode, though results reported here use discrete mode.
We designed three tasks that each require a qualitatively different generalization capability.
- Train: All single-operation and 2-step compositions (30 combinations)
- Test: All 3-step compositions (125 combinations, never seen during training)
- Tests: Whether the model understands operation semantics and can compose them
- Train: 5 list-reduction rules:
max,min,first,last,sum - Test: 3 structurally distinct rules:
second_max,median,range - Tests: Whether the model can apply the meta-skill of "find the rule from examples" to unseen rule types
- Train: Within-domain relations (Domain A demos → Domain A query)
- Test: Cross-domain transfer (Domain A demos → Domain B query)
- Tests: Whether the model can extract an abstract relation and re-apply it in a new symbol space
All models: ~2.4M parameters, dim=128, 5000 training steps, seed=42.
| Train (1-2 steps) | Test (3 steps) | |
|---|---|---|
| RT | 100% | 4.6% |
| Vanilla | 100% | 6.2% |
Honest assessment: Both models fail. Neither architecture achieves compositional generalization. They memorize input-output mappings rather than understanding operation semantics. This is a known open problem.
| Training Rules | Unseen Rules | |
|---|---|---|
| RT | 99.8% | 97.7% |
| Vanilla | 62.1% | 21.9% |
Per-rule breakdown (Unseen Rules):
| Rule | RT | Vanilla |
|---|---|---|
second_max |
99.0% | 20.3% |
median |
99.5% | 35.4% |
range |
98.4% | 13.0% |
Assessment: A large, real difference. RT generalizes to structurally novel rules while Vanilla fails to even fully learn the training rules (note: sum only 5.2% for Vanilla). The cross-example comparator appears to be the key contributor. Caveat: the test rules belong to the same mathematical family as training rules (list sorting + selection), so this may not represent fully general abstraction.
Not completed (experiment interrupted). Future work.
PonderNet ponder cost increases with task complexity:
- T2 (harder): ponder grows from ~1.5 → ~3.2 over training
- The model learns to allocate more computation to more complex problems
This demonstrates the adaptive iteration mechanism is functioning as intended.
We are transparent about the limitations of this work:
- Single seed — No error bars. Results may not be statistically stable.
- Small scale — ~2.4M parameters, synthetic tasks only. No real NLP benchmarks.
- T1 failure — Compositional generalization is unsolved. Claimed "reasoning" must be qualified.
- T2 caveat — Test rules are mathematically similar to training rules. This may be in-distribution generalization, not true abstraction.
- No ablations — Component contributions (comparator vs. memory vs. iteration) have not been isolated.
pip install torch>=2.5.0 numpy>=1.24.0Or with Docker (GPU):
docker compose up -d
docker compose exec hp-jepa python cognitive_tests.py --steps 5000 --dim 128 --seed 42python cognitive_tests.py --steps 5000 --dim 128 --seed 42
# Single test:
python cognitive_tests.py --steps 5000 --tests t2
# Quick smoke test:
python smoke_test.py| Argument | Default | Description |
|---|---|---|
--steps |
5000 | Training steps per model per test |
--dim |
128 | Model hidden dimension |
--bs |
32 | Batch size |
--lr |
1e-3 | Learning rate |
--seed |
42 | Random seed |
--tests |
all |
Comma-separated: t1,t2,t3 or all |
--device |
auto | cuda or cpu |
rt-cognitive-benchmark/
├── rt_calm.py # RT architecture + Vanilla baseline
├── cognitive_tests.py # Three cognitive test suite
├── smoke_test.py # Quick sanity check
├── docker-compose.yml # GPU Docker environment
├── requirements.txt
└── README.md
- Multi-seed evaluation with statistical tests
- Standard benchmarks (SCAN, COGS, ARC)
- Ablation studies: isolate each component's contribution
- Larger scale (dim=256+, more training data)
- Hierarchical architecture with bidirectional information flow (the deeper goal)
If you find this useful, please star the repository. If you build on this work:
@misc{rt-cognitive-benchmark-2026,
title = {RT-Cognitive-Benchmark: Structural Inductive Biases for Rule Discovery in Transformers},
year = {2026},
url = {https://github.com/lebonbruce/rt-cognitive-benchmark}
}本项目是一个实验性研究,探索对 Transformer 架构进行结构级改造是否能产生真正的认知能力:规则发现、组合推理、类比迁移。
这是诚实的开放性研究。我们如实报告成功和失败。
标准 Transformer 是强大的模式匹配器,但它缺乏支撑真正推理和抽象所需的结构性归纳偏置。本项目研究:加入迭代计算、跨示例比较、联想记忆等结构原语,是否能在泛化行为上产生质变?
我们没有声称解决了推理问题。 我们在问:针对性的架构改变,能产生质变不同的泛化行为吗?
基于 GPT-2 风格骨架,加入三个结构组件:
| 组件 | 功能 | 关键设计 |
|---|---|---|
AdaptiveIterativeCore |
PonderNet 式自适应迭代推理核 | 模型自己决定"想几步" |
CrossExampleComparator |
跨示例差值计算,提取不变变换模式 | 显式差值,不只是相似度 |
AnalogicalMemory |
输入调制的键值记忆 | 动态偏移,不是固定参数 |
参数量约 2.4M,dim=128,5000 训练步,seed=42。
训练 1-2 步操作组合,测试 3 步组合(从未见过)。两个模型都约等于随机(5%)。
训练 max/min/first/last/sum,测试 second_max/median/range(从未见过的规则类型)。
| 训练规则 | 测试规则(未见过) | |
|---|---|---|
| RT | 99.8% | 97.7% |
| Vanilla | 62.1% | 21.9% |
- 单种子,没有误差棒
- 小规模模型,合成任务,未在真实 NLP benchmark 上验证
- T1 组合泛化完全失败,这是当前架构的根本限制
- T2 的测试规则与训练规则属于同一数学族,可能不是真正的泛化抽象
- 没有消融实验
pip install torch>=2.5.0 numpy>=1.24.0
python cognitive_tests.py --steps 5000 --dim 128 --seed 42或使用 Docker(GPU 环境):
docker compose up -d
docker compose exec hp-jepa python cognitive_tests.py --steps 5000 --dim 128