Skip to content

About

Exploring structural inductive biases for rule discovery and reasoning in Transformers. RT architecture with PonderNet adaptive iteration + cross-example comparison + analogical memory.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

RT-Cognitive-Benchmark

English | 中文


What Is This?

An experimental research project exploring whether structural modifications to the Transformer architecture can give rise to genuine cognitive capabilities: rule discovery, compositional reasoning, and analogical transfer.

This is honest, open-ended research. We report both successes and failures.


Motivation

Standard Transformers are extraordinarily capable pattern-matchers, but arguably lack the structured inductive biases required for true reasoning and abstraction. This project investigates whether adding specific structural primitives — iterative computation, cross-example comparison, and associative memory — can close that gap.

We are not claiming to have solved reasoning. We are asking: can targeted architectural changes produce qualitatively different generalization behavior?


Architecture: RT (Reasoning Transformer)

Built on a GPT-2-style backbone with three structural additions:

[Token Embedding + Positional Encoding]
        ↓
[Encoder Layers × 3]  — standard Transformer blocks
        ↓
[CrossExampleComparator]  — extracts invariants across demonstrations
        ↓
[AdaptiveIterativeCore]   — PonderNet-style adaptive depth reasoning
        ↓
[Decoder Layers × 3]      — with AnalogicalMemory key-value injection
        ↓
[LM Head]

Component Descriptions

Component Purpose Key Design Choice
AdaptiveIterativeCore Shared-weight Transformer blocks iterated N times, where N is learned (PonderNet) Model decides its own "thinking depth"
CrossExampleComparator Computes pairwise differences between windows; extracts invariant transformation patterns Explicit difference computation, not just similarity
AnalogicalMemory Input-modulated key-value memory — base templates + input-conditioned offsets Dynamic, not fixed parameters

Also includes a MiniCALMAutoencoder and EnergyHead (inspired by CALM, arXiv:2510.27688) for continuous-vector prediction mode, though results reported here use discrete mode.


Benchmark: Three Cognitive Tests

We designed three tasks that each require a qualitatively different generalization capability.

T1: Compositional Generalization

  • Train: All single-operation and 2-step compositions (30 combinations)
  • Test: All 3-step compositions (125 combinations, never seen during training)
  • Tests: Whether the model understands operation semantics and can compose them

T2: Rule Discovery

  • Train: 5 list-reduction rules: max, min, first, last, sum
  • Test: 3 structurally distinct rules: second_max, median, range
  • Tests: Whether the model can apply the meta-skill of "find the rule from examples" to unseen rule types

T3: Cross-Domain Analogical Transfer

  • Train: Within-domain relations (Domain A demos → Domain A query)
  • Test: Cross-domain transfer (Domain A demos → Domain B query)
  • Tests: Whether the model can extract an abstract relation and re-apply it in a new symbol space

Results

All models: ~2.4M parameters, dim=128, 5000 training steps, seed=42.

T1: Compositional Generalization

Train (1-2 steps) Test (3 steps)
RT 100% 4.6%
Vanilla 100% 6.2%

Honest assessment: Both models fail. Neither architecture achieves compositional generalization. They memorize input-output mappings rather than understanding operation semantics. This is a known open problem.

T2: Rule Discovery ✅

Training Rules Unseen Rules
RT 99.8% 97.7%
Vanilla 62.1% 21.9%

Per-rule breakdown (Unseen Rules):

Rule RT Vanilla
second_max 99.0% 20.3%
median 99.5% 35.4%
range 98.4% 13.0%

Assessment: A large, real difference. RT generalizes to structurally novel rules while Vanilla fails to even fully learn the training rules (note: sum only 5.2% for Vanilla). The cross-example comparator appears to be the key contributor. Caveat: the test rules belong to the same mathematical family as training rules (list sorting + selection), so this may not represent fully general abstraction.

T3: Cross-Domain Analogical Transfer

Not completed (experiment interrupted). Future work.


Notable Finding: Adaptive Computation

PonderNet ponder cost increases with task complexity:

  • T2 (harder): ponder grows from ~1.5 → ~3.2 over training
  • The model learns to allocate more computation to more complex problems

This demonstrates the adaptive iteration mechanism is functioning as intended.


Limitations

We are transparent about the limitations of this work:

  1. Single seed — No error bars. Results may not be statistically stable.
  2. Small scale — ~2.4M parameters, synthetic tasks only. No real NLP benchmarks.
  3. T1 failure — Compositional generalization is unsolved. Claimed "reasoning" must be qualified.
  4. T2 caveat — Test rules are mathematically similar to training rules. This may be in-distribution generalization, not true abstraction.
  5. No ablations — Component contributions (comparator vs. memory vs. iteration) have not been isolated.

Running the Experiments

Requirements

pip install torch>=2.5.0 numpy>=1.24.0

Or with Docker (GPU):

docker compose up -d
docker compose exec hp-jepa python cognitive_tests.py --steps 5000 --dim 128 --seed 42

Without Docker

python cognitive_tests.py --steps 5000 --dim 128 --seed 42

# Single test:
python cognitive_tests.py --steps 5000 --tests t2

# Quick smoke test:
python smoke_test.py

Arguments

Argument Default Description
--steps 5000 Training steps per model per test
--dim 128 Model hidden dimension
--bs 32 Batch size
--lr 1e-3 Learning rate
--seed 42 Random seed
--tests all Comma-separated: t1,t2,t3 or all
--device auto cuda or cpu

File Structure

rt-cognitive-benchmark/
├── rt_calm.py          # RT architecture + Vanilla baseline
├── cognitive_tests.py  # Three cognitive test suite
├── smoke_test.py       # Quick sanity check
├── docker-compose.yml  # GPU Docker environment
├── requirements.txt
└── README.md

Future Directions

  • Multi-seed evaluation with statistical tests
  • Standard benchmarks (SCAN, COGS, ARC)
  • Ablation studies: isolate each component's contribution
  • Larger scale (dim=256+, more training data)
  • Hierarchical architecture with bidirectional information flow (the deeper goal)

Citation

If you find this useful, please star the repository. If you build on this work:

@misc{rt-cognitive-benchmark-2026,
  title   = {RT-Cognitive-Benchmark: Structural Inductive Biases for Rule Discovery in Transformers},
  year    = {2026},
  url     = {https://github.com/lebonbruce/rt-cognitive-benchmark}
}


中文

项目介绍

本项目是一个实验性研究,探索对 Transformer 架构进行结构级改造是否能产生真正的认知能力:规则发现、组合推理、类比迁移。

这是诚实的开放性研究。我们如实报告成功和失败。


研究动机

标准 Transformer 是强大的模式匹配器,但它缺乏支撑真正推理和抽象所需的结构性归纳偏置。本项目研究:加入迭代计算、跨示例比较、联想记忆等结构原语,是否能在泛化行为上产生质变?

我们没有声称解决了推理问题。 我们在问:针对性的架构改变,能产生质变不同的泛化行为吗?


架构: RT (Reasoning Transformer)

基于 GPT-2 风格骨架,加入三个结构组件:

组件 功能 关键设计
AdaptiveIterativeCore PonderNet 式自适应迭代推理核 模型自己决定"想几步"
CrossExampleComparator 跨示例差值计算,提取不变变换模式 显式差值,不只是相似度
AnalogicalMemory 输入调制的键值记忆 动态偏移,不是固定参数

实验结果

参数量约 2.4M,dim=128,5000 训练步,seed=42。

T1: 组合泛化 — ❌ 两者均失败

训练 1-2 步操作组合,测试 3 步组合(从未见过)。两个模型都约等于随机(5%)。

T2: 规则发现 — ✅ RT 大幅领先

训练 max/min/first/last/sum,测试 second_max/median/range(从未见过的规则类型)。

训练规则 测试规则(未见过)
RT 99.8% 97.7%
Vanilla 62.1% 21.9%

局限性(诚实声明)

  1. 单种子,没有误差棒
  2. 小规模模型,合成任务,未在真实 NLP benchmark 上验证
  3. T1 组合泛化完全失败,这是当前架构的根本限制
  4. T2 的测试规则与训练规则属于同一数学族,可能不是真正的泛化抽象
  5. 没有消融实验

运行

pip install torch>=2.5.0 numpy>=1.24.0
python cognitive_tests.py --steps 5000 --dim 128 --seed 42

或使用 Docker(GPU 环境):

docker compose up -d
docker compose exec hp-jepa python cognitive_tests.py --steps 5000 --dim 128

About

Exploring structural inductive biases for rule discovery and reasoning in Transformers. RT architecture with PonderNet adaptive iteration + cross-example comparison + analogical memory.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages