Skip to content
View noctem-o's full-sized avatar

Highlights

  • Pro

Block or report noctem-o

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
noctem-o/README.md

George Gallagher

AI systems you can inspect, evaluate, adapt and control.

Agent runtimes · evaluation & adaptation · trustworthy infrastructure · local inference

Recension Research Greater Manchester, UK

Founder + engineer at Recension Research

Endophasia · Magpie · Deadbolt · Cogitator


I build experimental infrastructure for AI agents and AI-operated systems: ways to observe runtime behaviour, evaluate it reproducibly, preserve evidence, keep authority outside the model, and verify what happened afterwards.

The recurring question is simple:

How do increasingly capable agent systems remain inspectable, evidential and governable without pretending one mechanism solves everything?

My work keeps five jobs separate:

observe  →  evidence  →  policy  →  permission  →  verify
runtime     support      standing    authority      recompute

Current work

Project What it does Surface
Endophasia Runtime-neutral experimental substrate for coding agents: trajectory capture and comparison, replay, capability studies, governed steering, cross-runtime adapters and controlled evaluation. TypeScript · Node.js · agent harnesses · evals
Magpie Replayable epistemic memory with signed history, typed evidence, provenance, policy-defined standing and detached standing receipts. Rust · SQLite · Ed25519 · portable verification
Deadbolt Permission engine for AI-operated systems. Proposals remain untrusted until a typed route, policy, exact lease and confirmation authorize an effect. Rust · capability security · MCP · rollback
Cogitator Tamper-evident recorder and verifier for agent runs, with canonical witness events, pre-dispatch policy interception, deterministic replay, drift detection and recomputable BLAKE3 witness roots. Rust · BLAKE3 · verification · release assurance

Each project stands alone. Together they explore a division of responsibility:

agent runtime
     │
     ▼
Endophasia
observe · steer · compare
     │
     ├──── Magpie
     │     evidence · provenance · standing
     │
     └──── Deadbolt
           permission · effects · receipts

Cogitator
reproducible run verification

This is a division of responsibility, not a dependency graph. Observation is not authority, and a successful evaluation is not permission to replace a running system.

Recent empirical work

Endophasia has moved from architecture into repeated live agent studies. The current research separates three different claims:

exact replay               did the recorded session reproduce under recorded responses?
fixed-condition repeat     do repeated live runs under the same declared conditions agree?
perturbation invariance    does behaviour hold when an input that should not matter changes?
Study Evidence
Pinned-environment repeatability 240 live trials under a controlled Pi + local-model configuration, with every manipulation check passing.
Path sensitivity 360 trials across 30 working-directory paths, exposing task-dependent behavioural sensitivity to a nominally irrelevant perturbation.
Discriminating-task construction 72-trial screen + 60 fresh confirmation trials, selecting tasks that avoid saturated ~0% / ~100% success regimes for later candidate evaluation.
Completion-cap / test-time compute 192-trial main study + 40-trial pilot: raising the completion cap improved pooled success on the two discriminating tasks from 10/24 to 19/23, while also exposing why evaluation workspaces need stronger isolation.
Governed steering audit Re-verified 364 recorded interventions while hardening consumption semantics, STOP handling and authorization expiry.

The useful outcome is not a prettier benchmark number. It is a better experimental instrument: one that can expose where sampling, environment, control flow and harness policy actually change an agent trajectory.

The runtime surface now extends beyond one harness: the merged ACP v1 vertical slice exercised a real OMP session through the standard protocol, while the separate ACP v2 conformance study pins a Draft v2 baseline, tests negotiation, replay and permission semantics against deterministic agents, and deliberately admits no v2 capability merely from protocol compatibility. That v2 study is complete and parked pending upstream stabilisation rather than chased as a moving draft.

Operator Alpha and evaluation-driven adaptation

The immediate Endophasia milestone is Operator Alpha: make the existing instrument something a person can install and run through a distributable, versioned endo CLI before broadening the research surface.

Beyond that operator layer, EVOLVE treats prompts, cognition policies, harness components, tools, memory, models, environments and inference budgets as explicit candidate changes.

observe → propose → isolate → evaluate → compare → admit → promote

Current research directions include:

  • pluggable evaluation, environment, adaptation and optional training providers
  • harness and policy search
  • test-time compute allocation and stopping policies
  • held-out checks and explicit promotion gates
  • trajectory-preserving evidence for candidate comparison

RL is one possible adaptation mechanism. Training can produce a candidate; it does not decide what the evidence establishes, nor whether that candidate may replace anything.

Longer-horizon alignment research

Beyond the current Endophasia build, I am starting to frame a separate line of research around whether alignment-relevant epistemic behaviour can be studied as a developmental property rather than only as a post-training correction.

The current design keeps the claim narrow: separate epistemic-content effects from timing effects, use synthetic microworlds with exact gold outcomes, keep Magpie as a shadow evidence / provenance / standing surface rather than a truth oracle, and treat null or reversed results as first-class outcomes.

A second question is measurement: whether developmental state can be described more precisely than training step alone. What Do Language Models Learn and When? The Implicit Curriculum Hypothesis and its open ElementalTask tooling provide useful emergence-order and trajectory baselines. I treat those as instrumentation, not as evidence that an early epistemic curriculum is causally better.

Developmental epistemic training v0 is the substrate-side research contract and preregistration draft. developmental-epistemics is the MIT-licensed, concept-only longer-horizon repository.

Engineering surface

Languages: Rust · TypeScript/Node.js · Python · Bash/PowerShell
Agent/runtime: Pi · MCP · ACP · OpenAI-compatible APIs · local agent harnesses
Inference: llama.cpp · vLLM · SGLang · TensorRT-LLM · CUDA · GGUF · quantisation · long-context serving
Systems: Linux · Arch · Nix · systemd · containers/virtualisation · networking · low-level debugging
Delivery: GitHub Actions · reproducible environments · conformance testing · golden fixtures · cargo-dist · release verification

I use frontier and local models as a structured engineering workforce: architecture-first decomposition, bounded implementation briefs, maker/checker separation, adversarial review and human-controlled promotion.

Working principles

observation  != inference
evidence     != standing
proposal     != permission
evaluation   != promotion
simulation   != real execution

These projects are experimental research software. I try to keep claims narrow, make failure states visible, test hostile cases, and document what each system does not establish.


Make capability legible before making it larger.

recensionresearch.co.uk

Pinned Loading

  1. endophasia endophasia Public

    A harness-neutral experimental substrate for observing, steering, comparing and evolving coding-agent systems, with explicit evidence, replay and capability boundaries.

    TypeScript

  2. magpie magpie Public

    Local-first, replayable memory for AI systems. Signed history, provenance, deterministic replay, and explicit policies for what evidence is allowed to conclude.

    Rust

  3. deadbolt deadbolt Public

    A local-first permission engine for AI-operated systems. Typed authority, explicit consent, bounded execution, receipts, and rollback.

    Rust

  4. cogitator cogitator Public

    Tamper-evident AI agent audit harness. Cryptographic witness chain, pre-call policy interception, and byte-stable replay. Prove what your agent did, tried to do, and was blocked from doing. Written…

    Rust