Skip to content

feat(memory): reliable rollout-memory pipeline, deterministic retention, and agent memory tools - #3351

Open
YodonTan wants to merge 13 commits into
GCWing:mainfrom
YodonTan:feat/memory-openbitfun
Open

YodonTan wants to merge 13 commits into
GCWing:mainfrom
YodonTan:feat/memory-openbitfun

Conversation

@YodonTan

Copy link
Copy Markdown
Contributor

Summary

  • Overhaul of the memory pipeline: root-cause fixes for rollout-memory never generating, stale memory never expiring, and silent failures (phase-2 cooldown bypass on new input, diff-filtered tool state, unified retention clock, lease heartbeat reclaim, internal phase-2 agent auto-approval, partial-timeout no longer counted as success).
  • New agent memory tools: MemorySearch / MemoryFetch / MemorySave (note + capture_session), exposed to the standard harness and claw mode, gated on the memories config.
  • Failure bookkeeping: stage-1 failure records now expire (memories.max_failure_report_days, default 14d, clamped 1..90), a clear-failures action, per-record failure timestamps, and a compact restyled status card.
  • Adaptive phase-2 consolidation timeout (1200s..3600s scaled by diff size and session count).
  • Merges origin/main (9e51317dd, 106 commits) — conflict resolutions: agents/mod.rs standard harness list (kept memory trio, accepted upstream TodoWrite removal), MemorySettingsSection.tsx seeding integration (kept i18n/notification wiring + upstream useConfigSeed), 9 capability artifacts regenerated.
  • Test-side: E2E failure-bookkeeping spec + fixture seeder (mock stage-1 failure rows), journey race fix (turn-keyed assistant-round wait), memory catalog copy aligned with governance rules.

Verification

  • agentic::memories unit tests: 104/104; agentic::agents 105; agentic::tools 566 (1 pre-existing Windows path-separator failure, untouched by this branch); memory_tool 35; openbitfun-tool-packs 11; config-contracts memories 3.
  • E2E on the merged tip (real app + real model): failure-bookkeeping spec 4/4 (clear action, 14-day expiry sweep, compact error cards, failure dates); memory journey 8/8 (extraction → consolidation → new-session injection → MemorySearch).
  • Capability surfaces: capabilities:generate/:check PASS, capabilities:test 16/16; check:core-boundaries PASS; i18n:audit 0 warnings; web type-check PASS; memory settings section tests 19/19.
  • Two read-only audits (27-file overlap classification via git merge-tree, 106-commit semantic interference review) guided the merge; findings and resolutions recorded in the PR discussion.

Known follow-ups (explicitly out of scope here)

  • MiniApp agent workspaces now enter the phase-1 scan input set (upstream 43eaca438): decision is to exclude them in a follow-up change.
  • Unregistered workspace paths degrade phase-1 on-demand capture to "storage unavailable" silently (upstream session-storage port tightening): follow-up will surface this loudly.
  • tests/e2e has 118 standing tsc errors (pre-existing typing debt; one import-path nit in the memory journey spec).

Tant added 13 commits September 29, 2026 00:07
Port the memory overhaul from the pre-rename branch onto the 1.0.x base as one
commit, expressed with the current crate, type, and citation naming.

Ported fixes:
- phase2: bypass the success cooldown when new stage-1 input or meaningful
  workspace changes exist; never wipe inputs on an empty selection; honest
  empty-selection handling.
- workspace diff: exclude gitignored and tool-state files (`.openbitfun/`, the
  legacy `.bitfun/`, `.git/`) so no-op runs stay cheap and tool state never
  reaches the consolidation prompt.
- sync: prune rollout summaries only when their stage-1 row is gone; the
  keep-set is every live row, not the current selection.
- retention: one unused-days clock for all rows, a one-day floor, a 365-day
  age-gate clamp, and one aggregate log line for age-gated sessions.
- phase1: requeue a retry-exhausted session after a 24h quarantine, on-demand
  `run_once_for_session`, and a user-message transcript fallback.
- phase2 job: reclaim a stale lease through heartbeat liveness
  (`phase2_has_live_run`) and treat an incomplete consolidation as not
  successful, keeping the baseline and the pending input for the next pass.
- status and usage APIs: `status_snapshot`, `record_memory_usage`, and
  `list_live_rollout_summary_identities`.
- internal consolidation agent: auto-approve permission mode for the
  unattended turn, still bounded by its tool allowlist and memory-root path
  policy.
- prompts: deleting the dependent memory content is mandatory, a renamed
  rollout summary is not evidence loss, and no "evidence unavailable"
  placeholder may remain.
- docs: `docs/features/agent-memory.md` and its zh-CN counterpart.

Adaptations to the new base:
- The shared store now owns `MemoryRow`/`MemoryJobRow`
  (`openbitfun-services-core::memory_store`), so the duplicate record
  definitions are gone and the phase-2 success-cooldown predicate lives in
  `memories::db` instead: the shared method still requires the last
  consolidation watermark to equal the evaluated one, which would disarm the
  cooldown after every idle completion.
- Session storage is resolved by workspace ID on the scan path; the
  path-based helper used by the on-demand run returns `Option` and the run
  releases its claim instead of reading a directory that is not session
  storage.
- The internal phase-2 request carries the registered memory-root
  `workspace_id`, and the request builder keeps the auto-approve context.
- The memory-root registration guard is now an identity check on the resolved
  workspace record, because activity tracking is workspace-ID based; it stops
  dialog activity from adding the memory root to the recent-workspace list.
- Memory file and directory constants come from the shared memory store;
  `SKILLS_DIR_NAME` stays in the memory workspace module.
… 1.0.x

- MemorySearch / MemoryFetch / MemorySave built-in tools (Basic group, Direct exposure, use/generate gates, host-local only, usage recording on fetch/search)
- product-operation registry rows for get_memory_status / run_memory_pipeline_now (+ get_memory_paths graduated to Agnostic) and regenerated registries/catalogs
- desktop memory_api commands and generate_handler entries
- settings: truthful memory defaults, retention disclosure + 1-day floor guard, status card + run-now driven by the page config, MemorySettingsPage tests
Port the l1-memory-pipeline journey (settings -> capture_session -> on-demand
pipeline -> injection -> MemorySearch) onto the current base, adapting it to the
surfaces that changed upstream:

- Settings: MemoriesConfig -> MemorySettingsPage under Settings > AI > Memory
  (page id ai.memory, locale namespace settings/memory.json). The single
  memories-enabled-toggle replaces the old generate/use pair; advanced rows sit
  behind actions.expandAdvanced. Status-card testids are preserved.
- Permission: the old PermissionRequestPanel family is gone. Approval asks now
  render as ChatInputApprovalBand inside the composer, so PermissionPanel targets
  data-openbitfun-component="permission-request-panel" plus the
  chat-input-approval-* controls.
- Status attributes: data-bf-* no longer exists; helpers read
  data-openbitfun-status / data-openbitfun-state.
- Nav/Scene: session creation moved to the workspace-scoped
  nav-workspace-new-session-btn with an overflow-menu fallback; scene tabs are
  selected by data-scene-tab-id.
- Cold start: the nav entry can take up to a minute to appear on a fresh
  profile, so the journey waits for it explicitly instead of racing a 10s probe.
- Harness: stdout/stderr mirroring into app-stdout-<ts>.log, a
  OPENBITFUN_E2E_MOCHA_TIMEOUT knob, DOM-probe helpers, and a loud-skip seeding
  rule for the sandbox profile (ai section only, reasoning presets stripped).

Verified locally: 8/8 journey steps pass against a freshly built debug binary.
…ip reasons, ignore divergence, root identity)
…and compact error display

Phase 2 consolidation ran on a fixed 20 minute budget, so a large memory
workspace timed out silently. The budget now scales with the workload the
turn must read: 20 minutes floor, plus one second per 200 B of
phase2_workspace_diff.md and one second per KiB of MEMORY.md, clamped to
60 minutes. The observed ~200 KB diff with a ~393 KB index lands at ~44
minutes and a <10 KB diff stays at ~21 minutes. The measured inputs and the
computed budget are logged at agent start, and the pure formula plus the
workload reader are unit tested.

The failed-extraction list only cleared itself through the 24 hour
quarantine requeue. The Memory settings page now offers a confirmed "clear
failed records" action: MemoryDatabase::clear_failed_stage1_jobs deletes
only memory_stage1/* rows in error state, keeping the global phase-2 job,
successful rows, and every stage1_outputs row intact. Because the claim
path inserts a fresh job when no row exists, cleared sessions become
eligible for extraction again on the next run.

Both error blocks on the page were plain body-size inline text, so raw
provider JSON and timeout reports read like settings values. They now
render as compact tinted blocks one step below the description size, with
the provider payload tail stripped from the visible line, the session id
and retry-exhausted badge on a header line, and the full message kept on
the title attribute.

Verified: openbitfun-core memories + agentic_api tests, desktop command
compiles, MemorySettingsPage and settings scene suites, capabilities
generate/check, i18n audit, typography/theme/appearance audits.
Failure bookkeeping must not outlive its usefulness. A stage-1 extraction
that failed longer ago than the new `memories.max_failure_report_days`
window (default 14 days, clamped to 1..90) is now deleted, so a
retry-exhausted dead end such as a since-replaced model cannot stay listed
forever. The sweep runs at startup and before every on-demand pass, ahead
of phase-1 claims, so an expired row is gone before the quarantine requeue
would reconsider its attempt budget. Only per-session stage-1 error rows
expire: the global phase-2 job keeps its last_error as current-run
diagnostics, and stage1_outputs is never touched.

The status snapshot now reports when a failure happened
(`failed_at_unix_secs`, from finished_at or started_at), and the settings
page renders that age in the compact failure entry through the shared i18n
date helper, so a stale row is answerable at a glance.

Config and wire shapes are additive only: an older config without the key
deserializes to the default, and the extra DTO field is serde-default.
# Conflicts:
#	docs/interactive-capabilities/capabilities.json
#	docs/interactive-capabilities/technical/product-control-open-audit.json
#	docs/interactive-capabilities/technical/tauri-command-map.json
#	src/crates/assembly/core/src/agentic/tools/agent-tool-exposure.md
#	src/crates/contracts/product-domains/src/generated/product-control-catalog.json
#	src/web-ui/src/app/global-search/generated/interactive-capabilities.json
#	src/web-ui/src/app/scenes/settings/pages/ai/MemorySettingsSection.tsx
#	src/web-ui/src/app/scenes/settings/settingsRegistry.ts
#	src/web-ui/src/infrastructure/api/generated/productControl.ts
#	src/web-ui/src/locales/en-US/settings/memory.json
#	src/web-ui/src/locales/zh-CN/settings/memory.json
#	src/web-ui/src/locales/zh-TW/settings/memory.json
Conflict resolutions (11 conflicts: 2 hand-written + 9 generated)
- src/crates/assembly/core/src/agentic/agents/mod.rs: standard_harness_tools() keeps the
  memory trio (MemorySearch/MemoryFetch/MemorySave) and accepts upstream 49e5281's
  TodoWrite default removal; no TodoWrite remains in the harness list.
- src/web-ui/src/app/scenes/settings/pages/ai/MemorySettingsSection.tsx: three hunks -
  imports keep upstream's useConfigSeed plus our ConfigMessage; the component head keeps our
  useI18n/useNotification wiring (formatDate/formatNumber/notifyInfo) together with upstream's
  seed-based state (hasLoaded/loading from seed); loadData keeps our split model fetch and
  upstream's hasLoaded.current = true bookkeeping.
- The 9 generated capability artifacts were resolved one-sided (no marker surgery) and then
  regenerated from the merged sources.

Regeneration
- pnpm run capabilities:generate -> "Generated 22 features + 21 settings; audited 673 Tauri
  commands and 3810 UI interaction candidates across 373 files".
- pnpm run capabilities:check -> PASS (sources and generated artifacts agree).

Verification
- cargo test -p openbitfun-core --no-default-features --features "agent-runtime,git" --lib
  agentic::memories -> 104 passed, 0 failed.
- cargo test -p openbitfun-core --no-default-features --features "agent-runtime,git"
  agentic::agents -> 105 passed; --lib agentic::tools -> 566 passed, 1 pre-existing
  Windows-only path-separator failure in skills::registry (file untouched by this merge).
- cargo test -p openbitfun-config-contracts memories -> 3 passed.
- cargo test -p openbitfun-tool-packs (package of src/crates/execution/tool-provider-groups)
  -> 11 passed.
- pnpm run check:core-boundaries (incl. agent-harness:check) -> PASS.
- pnpm run lint:web -> 0 errors / 10 pre-existing warnings; src/web-ui type-check -> PASS;
  MemorySettingsSection.presentation.test.ts -> 2 passed; settingsRegistry.test.ts -> 30
  passed; pnpm run i18n:audit -> PASS.
- cargo build -p openbitfun-desktop (debug) -> PASS.
- pnpm --dir tests/e2e run test:l1:memory-failure -> 4/4 passed.
- OPENBITFUN_E2E_MOCHA_TIMEOUT=900000 pnpm --dir tests/e2e run test:l1:memory -> 8/8 passed.

Known failures that pre-date this merge and are unchanged by it (identical on 39ad71d)
- scripts/interactive-capabilities.test.mjs: 'the public contract is a compact
  feature-and-settings manual' and 'every interaction-only control has an individually
  reviewable reason and evidence' fail on the memory feature's own catalog entries
  (feature.ai-assistant:memory-tools contains the literal "evidence";
  setting.ai.memory:pipeline-status's structured reason omits its item title).
- src/web-ui/src/infrastructure/config/components/MemorySettingsPage.test.tsx imports the
  non-existent ./MemorySettingsPage module.
Both `capabilities:test` failures came from our memory copy in
`src/shared/interactive-capabilities/catalog.json`:

- `feature.ai-assistant:memory-tools`: the agent example contained the
  literal word "evidence", which the public-catalog assertion rejects
  globally. Reworded to "read the file MEMORY.md points at".
- `setting.ai.memory:pipeline-status`: the interaction reason omitted the
  item title, so the entry was not individually reviewable. The zh/en
  reasons now open with the title and depend on it.

Regenerated the catalog projections for the new graph digest
(a042a92c -> 88a95ef7).

Also relocate the orphaned memory settings behavioural test. Its module,
`MemorySettingsPage.tsx`, was deleted upstream, so the test could no
longer resolve its import and failed the collected web-ui suite. The
replacement module preserved every seam the test targets, so its 17
behavioural assertions move next to `MemorySettingsSection.tsx`
unchanged; the duplicate-free coverage was not dropped.
The product identity audit rejects the retired token outside the legacy
migration boundary, so the memory feature must not re-know it:

- resolve the memory tool-state directory from the product identity owner
  (`hidden_data_directory()`) instead of a hardcoded literal, which also
  makes the filter correct for a build with a different data namespace;
- stop probing retired config roots in the memory e2e seeder and keep the
  documented explicit override for a pre-rename profile;
- state only the current tool-state directory in both memory docs.

Legacy name knowledge stays in `openbitfun-legacy-migration`.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant