Skip to content

[pull] main from danny-avila:main - #260

Merged
pull[bot] merged 22 commits into
innFactory:mainfrom
LibreChat-AI:main
Sep 13, 2026
Merged

pull[bot] merged 22 commits into
innFactory:mainfrom
LibreChat-AI:main

Conversation

@pull

@pull pull Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

danny-avila and others added 22 commits September 12, 2026 05:45
* 🧭 feat: Remember Attached Workspace Choices

* Guard Workspace Defaults Across Navigation and Version Restore

* Order Workspace Preference Imports

* fix: Harden agent workspace defaults

* fix: Type workspace validation schema

* test: Cover safe execution projections

* fix: Bind workspace defaults to attached environments

* fix: Preserve unchanged workspace bindings

* fix: Scope workspace preferences to active agents

* fix: Preserve workspace selection state

* test: Target workspace discovery status

* fix: Keep workspace reset available

* fix: Resolve workspace defaults per environment

* fix: Collect per-environment workspace defaults

* fix: Reconcile Deleted Agent Graph Caches

* fix: Reject Stale Workspace Bindings
* 🛡️ fix: Prevent ReDoS in skill and artifact parsing

* 🧵 fix: Normalize artifact fences in linear time

* 🧠 fix: Bound artifact normalization memory
* fix: Preserve Directory Identity Ownership

* fix: Fail Closed on Identity Conflicts

* fix: Resolve Existing Sharing Principals Safely
* fix: Prefer storage keys for image encoding

* fix: Sanitize assistant image errors

* fix: Sanitize agent image storage errors

* fix: Move blob image encoding into API package

* fix: Complete image encoder TypeScript migration

* fix: Keep request-boundary error logs diagnosable

getSafeErrorMetadata reduces an error to { type: 'Error' } and an optional
status. That is the right bound where an error may echo submitted content, and
it stays in place at getS3FileStream, getAzureFileStream and
getFirebaseFileStream, where a signed URL reaches the log.

Applied at ResumableAgentController's initialization catch and the chatV1
generic branch it blinded the whole route: a missing key, an agent that fails to
load and a database error all logged as { type: 'Error' }, with no message,
name, code or stack.

getSafeErrorText keeps the error's own description and stack, reduces every URL
to its origin so an object path and signature cannot ride along, and strips
bearer credentials. It returns text so the two handlers log it inside the
message and pass no winston metadata, which is what stops an SDK error's own
enumerable properties (request.url, config, response.headers) from being
promoted onto the record.

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
…#15863)

A durable user connection captures the tool-cache publication generation before
resolving credentials, so a rotation from another replica fences the catalog it
publishes. A silent OAuth refresh performed while that same connection is being
built publishes a rotation of its own, and the build then fails the lease check
it just invalidated: the connection holding the newest credentials is torn down
and the tool call that triggered it errors with "Publication lease is no longer
current".

invalidateCachedTools now reports the generation it wrote, the authorization
publication carries it back to its caller, and a connection build adopts what
its own refresh published. A foreign rotation still moves the stored generation
past that value, so the fence keeps failing the build it was meant to fail.
* 🔎 feat: Filter and Sort the Sidebar Chats List

The sidebar could only list every unarchived chat by last activity. Archived chats
lived behind a settings modal, sorting was not exposed at all, and the only filter
was a bookmark dropdown squeezed next to the search field.

A single control beside the Chats heading now carries the status (active or
archived), the sort field and direction, and the bookmark filter, each scoped to
that list alone: Projects, Pinned and Favorites are untouched by it.

* 🧹 fix: Answer Review Round on Sidebar Chat Filters

Rows derive their archive action from `conversation.isArchived`, which the list
projection now carries, instead of from the sidebar's filter mode: the archive view
keeps the unarchived pins beside it, and its previous page's rows outlive a status
switch, so the list mode offered to restore chats that were never archived.

Mutations reconcile the archived list caches they can now reach — rename, pin,
tags, project assignment, shared links — and archiving refetches the archived
query instead of inserting into page one, which dropped the row whenever the
active sort placed it later.

Title grouping keeps the order the server paged, heads each group with the title's
first code point, and the archive's legacy group falls back to `createdAt` the way
the cursor orders it. A failed list request renders an error with Retry rather
than "No chats yet", the status setter sanitizes a malformed stored sort, and
signing into another account no longer inherits the previous one's filters.

* 🧵 fix: Keep Sidebar Cache Writes Inside the Chosen Sort

Title groups are emitted as contiguous runs, so a `#` heading that the server's
order places either side of a letter no longer merges and drags a row across the
cursor boundary the next page continues from.

The live cache writers only reproduce newest-first, which is the one order they can
place a row in without the server's keyset. A title, created-at or ascending variant
now takes the field update in place and is refetched instead of having rows moved to
or seeded at the front of page one.

Archive and restore refetch every loaded page of the lists they touch — the archive
preserves `updatedAt`, so a restored chat can belong on a later active page — while
inactive variants are left stale rather than eagerly refetched. Archiving everything
still refetches them all, because every cached list is wrong afterwards. Renaming or
deleting a bookmark now rewrites the sidebar's selected tags, which otherwise kept
filtering by a name no menu could uncheck.

* 🗂️ fix: Treat the Archived List as a Live Conversation Cache

* 🔍 fix: Refetch List Variants Whose Membership the Client Cannot Decide

A cached search result is evaluated by the server, so a newly created row may or may
not belong in it. Skipping those variants left a mounted search list missing a row
it should show — forking from an archived search result being the case that exposed
it. The insert decision is now three-valued: insert into a variant whose membership
and order the client can reproduce, skip one the row provably does not belong to,
and refetch anything undecidable.

Forks, duplicates and imports inherit the source's archive state, so their
invalidations cover the archived list too, and the delete path reads `chatProjectId`
from archived caches as well.

* 🔎 fix: Reconcile Searched List Variants on Every Write

* 📣 fix: Announce Archive and Restore From the App Live Region

* 🧹 fix: Drop Unmounted Lists on Archive-All and Fill a Missing Archive State

Archive-all refetched every loaded page of every cached list, mounted or not, so a
session that had visited several sort, bookmark or search combinations paid a burst of
requests for lists nobody was looking at. Mounted variants still refetch all pages;
the rest are dropped, which reconciles them more thoroughly than a refetch and costs
nothing until something mounts them.

A row's own `isArchived` decides whether its menu offers Archive or Restore, and this
branch is what adds that field to the list projection — so mid-rollout, an archived
row from an older backend would offer Archive and submit a no-op. The list query now
fills a missing value from what the variant asked for, leaving an explicit one alone.

* 🧭 fix: Give Callers One Way to Reconcile Both Conversation Lists

Callers that cannot say what changed — a stream recovered after a disconnect, a
schedule whose run moved a chat, a project that took its chats' fields with it — were
invalidating the active list by hand, so a mounted archive kept a row the server had
already changed. `invalidateConversationLists` reaches both prefixes and is now the
only way those paths reconcile; `findConversationListQueries` shares its key list, so
adding a prefix is one edit rather than a sweep.

* 🔤 fix: Follow the Server Order, Index the Created Sort, and Tell the Truth When Empty

* 🎯 fix: Identify the Open Chat When a Restore Resolves

A restore leaves the chat open, so its success callback writes the new archive state
into the active conversation — but the callback captured that conversation when the
menu item was clicked. Open another chat before the request lands and the identity
check still matched the captured row, replacing the conversation the user had just
opened while the URL kept pointing at it.

The write is now a functional update that compares `conversationId` against current
state, so a stale callback leaves the open chat alone. Reading the conversation value
is no longer needed either, which also stops every row's menu from re-rendering on
conversation updates.

* 🧷 fix: Recheck the Open Route When an Archive Resolves
* 🪪 style: Give Native Tools a Painted Verified Mark

The native badge on a tool card was lucide's BadgeCheck filled with status-success-strong: green, the same hue the card already uses for selected, and stroked in currentColor, so the white check color also outlined the badge and left a white halo on any surface that is not the card's own white — hover, selection, dark mode.\n\nVerifiedIcon paints the badge instead of stroking it and carries the check as the only stroked path, and status-verified gives provenance its own blue across all four themes.

* 🩹 fix: Hold the Verified Mark at 3:1 on Every Card State

The dark fill was tuned against surface-secondary, but ToolCard repaints to surface-tertiary (#2f2f2f) on hover, where #0b74d4 fell to 2.85:1 — below the non-text floor, and invisible to the spec, which asserted the panel only. The fill moves to #1a7fd8 and the contract now covers every background the card takes. Both of the mark's relationships are graphical objects under WCAG 1.4.11, so the check is held at 3:1 rather than AA: no hue clears 3:1 on the hover surface and 4.5:1 under a white check at once.\n\nA theme that predates the token names status-success-strong and not status-verified, and would otherwise wear LibreChat's stock blue beside its own text-on-status and card surfaces. resolveTheme and the legacy mapColors adapter now carry the mark onto the fill it had before the token existed, following the six derived-token fallbacks already there.

* ✅ test: Exercise the Verified Mark in the Tool Library

Four scenarios in the mock harness, each named by the behavior it observes: the mark a native card wears, its silhouette under the dark hover repaint that the retune fixes, a pre-token deployment palette carrying it on the theme's own success fill, and a third-party MCP server card carrying no mark at all. The contrast assertions read the rendered colors rather than the tokens, so a regression in the component or in theme application fails them.

* 🧪 fix: Teach the Marketplace Dialog Spec About VerifiedIcon

The suite hand-lists the @librechat/client surface it mocks, so ToolCard's new icon resolved to undefined and every render in the file threw.

* 🧪 fix: Reach the Agent Builder on a Narrow Viewport

The scenario run covers mobile emulation, where the side rail is collapsed and the panel opens from the header's Control Panel toggle; the helper only knew the rail. The MCP scenario also minted its access token before the app had loaded, so the refresh request had no origin to resolve against.

* 🧪 fix: Run the Verified-Mark Scenarios on a Touch Viewport

Mobile emulation has no side rail: the panel switcher lives inside a sidebar that starts off-canvas, so the helper opens the sidebar and picks the Agent Builder entry from the switcher's menu. A touch viewport also has no hover state, so the hover scenario measures whatever background the card actually takes there and still holds the mark to 3:1 against it.

* 🩹 fix: Keep Pre-Token Themes on the Fill the Mark Inherited

The first fallback only fired when a theme restated status-success-strong, so a partial palette that customizes surfaces or text-on-status and inherits the green still landed on the stock blue. resolveTheme now derives the omission for any custom theme from the resolved success fill, and the legacy adapter takes the mode's bundled palette from ThemeProvider so it can do the same for a theme that names neither.\n\nThe scenario helper picked its path from a one-second probe for the desktop rail, which turns a slow first paint into the mobile branch; it now decides on viewport width and keeps the suite's full assertion timeout in each branch. The mock project is Desktop Chrome, so a touch scenario asks for a touch context explicitly rather than leaving that layout unexercised in CI.

* 🧪 fix: Match the App's Inclusive Layout Breakpoint

useSidebarState switches on (max-width: 768px), so a 768px-wide context renders the sidebar switcher while the helper's strict comparison went looking for the desktop rail.

* 🩹 fix: Scope the Verified Fallback to Themes That Painted the Mark

Two rounds pulled the same predicate in opposite directions: keying it on a theme restating status-success-strong missed palettes that only repaint the check or the card, and keying it on any custom color at all gave a new partial theme a green mark identical to the selected cue. Both are answered by asking what the theme actually painted — the fill the mark wore, the text-on-status check it carries, or the surfaces ToolCard takes — which is the same shape as the series-scale fallback above it. A theme that repaints none of those keeps the bundled blue.

* 🧪 fix: Wait for the Shell Before Reading the Sidebar State

isVisible() answers immediately, so a slow first paint read as an already-open drawer and the helper then waited on a switcher sealed inside it. It now waits for whichever of the opener and the switcher the shell paints first, which is the same assertion timeout the desktop branch gets.

* 📝 fix: Keep the Token Catalog Entry Catalog-Sized

The entry restated the icon's construction, the contrast rule and the pre-token fallback, all of which live on the token's own doc comment and in semanticTokens.spec.ts; the catalog around it is one line per role.

* 🔦 fix: Let a Crashed Lighthouse Run Be Retried

The audit shells out to Lighthouse three times and fails the lane if any single process exits nonzero, so a NO_NAVSTART — Lighthouse losing the trace of a page load, whose own message is "run Lighthouse again" — sinks a pull request that has nothing wrong with it. That is what happened on this branch: runs 1 and 2 wrote reports, run 3 crashed, and the identical head passed on rerun.\n\nEach run now gets a second attempt before the failure propagates. The repeat is per run, so the median is still taken over three completed reports, and a page that genuinely cannot be traced fails both attempts and still fails the lane.

* 🧪 fix: Read the Drawer State From Its Inert Marker

The readiness wait accepted the panel switcher, which is mounted whether or not the drawer is open — translated off-canvas and inert — so a shell that painted the sidebar before the chat header could satisfy it with the drawer closed and leave the helper clicking an off-viewport control. The drawer's own inert attribute is the state, so the helper waits for the drawer to attach and reads it there, with no timing probe on either side.

* 🧪 fix: Open the Drawer Before Trusting the Restored Form

A narrow context that had already selected the builder restores the form inside the closed drawer, and a translated element passes a visibility check, so the early probe accepted a form sitting under an inert drawer and handed it to the caller. The layout decision now comes first and the drawer is opened before anything inside it is read.
…#15866)

* 🧪 test: Restore Two Rotted packages/api Integration Suites

MCPReinitRecovery.integration.test.ts mocked @librechat/data-schemas down to
its logger, so all eight tests threw once MCPServersRegistry started calling
scopedCacheKey (#15028). The mock now spreads the real module under the logger
override; ServerConfigsDB stays mocked, so the suite still never touches Mongo.

secrets.integration.spec.ts drove langfuse.secretKey through the generic admin
config API, which has stripped the langfuse section since #14108 moved it to
the dedicated Langfuse connection API. The generic-write matrix drops that row,
and a langfuse block covers the current contract against a real Config
collection: the stored secret stays encrypted at rest, generic reads redact it,
generic patch, tombstone and field-delete writes leave it untouched, and a full
base-config replacement carries it forward.

* 👷 ci: Run Every packages/api Integration Suite in CI

test:ci excludes every path containing "integration.", but the jobs meant to
run those suites selected by folder: the cache workflow only
src/{cache,cluster,mcp,stream}, the agents workflow only
src/agents/*.integration.spec.ts. Thirteen suites ran in no job, including
every *.integration.test.ts.

test:integration selects *.integration.(spec|test).ts anywhere under src and
replaces test:agents-integration in the (renamed) Integration Tests workflow.
The cache scripts also accept .test.ts, stream selects by suffix alone, and
core picks up src/middleware. Every suite test:ci excludes is now selected by
exactly one CI script.

* 🧪 test: Read the Tool Snapshot in the Concurrent Reinit Simulation

The concurrent reinitialize test copied reinitMCPServer but still called
fetchTools(), which returns [] when tools/list fails. Concurrent reinspections
of a recovered server can each succeed, and a later write replaces the app
connection an earlier request was already handed, so that request's tools/list
fails. The test failed in 9 of 25 iterations on Linux, and in the first CI run
that included the suite. Since #14686 production reads the snapshot and keeps
cached tools when it is incomplete; the simulation now does the same, and still
requires every complete snapshot to list both tools and at least one to complete.

A deterministic test pins the property that tolerance relies on, without
depending on the race: a newer config write replaces the held app connection,
whose snapshot is then incomplete, while the replacement lists the real tools.
…15867)

* 🪢 fix: Recover a Failed MCP Server Once When Reinspections Overlap

A server stored as an inspection-failed stub is recovered by the first
request that inspects it successfully, but every request that read the
stub before that write inspected again and wrote again. Each later write
bumped `updatedAt` past the connection made after the earlier one, so the
connection was replaced and a caller still holding it listed no tools. A
request that reached `reinspectServer` after the write hit the
"not in a failed state" guard and reported the server as still
unreachable.

- Share one in-flight reinspection per stored entry and effective
  allowlists in each process; DB entries are keyed per user and tenant.
- Replace the stub only while it is still the inspected stub: a new
  optional `replaceStub` compare-and-set on the in-memory and Redis
  aggregate-key stores (one Lua script), so replicas never overwrite a
  recovery another replica stored.
- Resolve a reinspection that finds the server already recovered to the
  stored config, and move the reinitialize decision into
  `MCPServersRegistry.recoverServerConfig` so `reinitMCPServer` connects
  with the recovered config instead of the stub it read.
- Pin the interleavings in `MCPReinitRecovery.integration.test.ts` against
  a real in-process MCP server, and require every concurrent reinitialize
  flow there to recover with a complete tool list.

* 🪢 fix: Key Shared Reinspections by the Stub They Inspect

A registry re-initialization can replace a failed stub while an older
reinspection of it is still running. The shared flight was keyed by
entry and allowlists only, so a request that read the newer stub joined
the older inspection and inherited its outcome without the newer config
ever being inspected.

`reinspectServer` now reads the entry before joining and keys the flight
by the stub's `updatedAt`. A flight whose compare-and-set loses to a
newer failed stub joins the inspection of that stub instead of reporting
the server unreachable, and a store that leaves the inspected stub in
place rejects rather than joining its own flight.

* 🪢 fix: Settle a Reinspection Against the Entry Stored After It

A reinspection that did not replace its stub only looked at storage when
its compare-and-set lost. Two paths reached a stale outcome:

- An inspection that failed rejected as unreachable even when another
  request had already stored a recovery, for example after capturing the
  stub before a slow allowlist resolution. Allowlists now resolve before
  the entry is read, and a failed inspection settles against the stored
  entry the same way a lost compare-and-set does.
- Returning a recovery that another replica stored left this replica's
  process-local `getAllServerConfigs` memo serving the stub until its TTL.
  Settling on a stored recovery now invalidates this replica's registry
  read caches, as the writing path does.

* 🪢 fix: Decide Reinspection Outcomes From Current Shared State

The Redis aggregate-key store answers `get` from a process-local
snapshot that can lag another replica's write by its TTL. A replica
whose inspection failed read that snapshot while settling, still saw the
stub, and reported the server unreachable although Redis held another
replica's recovery. The same stale read also sent a replica into a
needless inspection of a server another replica had already recovered.

- Add an optional `getCurrent` to the repository interface, implemented
  by the aggregate-key store as a read past its snapshot, and use it for
  the reinspection's first read and for settling.
- Inspect a newer stub within the flight that found it instead of
  joining another flight, so no flight ever waits on a second one.

* 🪢 fix: Join the Flight Already Inspecting a Newer Stub

A flight that lost to a newer failed stub inspected that stub itself even
when another request was already inspecting it. If that redundant
inspection failed transiently, the flight settled against the unchanged
newer stub and rejected its callers while the other flight went on to
store the recovery.

A flight now joins the flight for the newer stub, starting it when none
is pending, but only when the stub is strictly newer by `updatedAt`, so
every wait points forward and flights stay acyclic. A replacement that
is not newer, which only clock skew can produce, is still inspected
within the flight.

* 🪢 fix: Track Reinspection Waits Instead of Ordering Them by Clock

Joining only flights for strictly newer stubs kept waits acyclic by
relying on `updatedAt` order, which replicas with skewed clocks do not
guarantee. A replacement stored with an older timestamp was inspected
redundantly instead of joining the flight already inspecting it, and a
transient failure of that redundant inspection rejected callers the
other flight went on to recover.

Settling flights now record the flight they wait on and join any
in-flight reinspection for the stub they found, unless that flight
already waits on them directly or through others. Only a join that
would close a cycle falls back to inspecting within the settling flight.

* 🪢 refactor: Resolve the Reinitialize Config in packages/api

`reinitMCPServer` still decided in CJS what a failed-inspection config
recovers to and built the `unreachable` response itself. That decision
now lives in `resolveMCPReinitializeConfig` in packages/api, which takes
the registry as a dependency and returns either the config to connect
with or the typed result that ends the reinitialization. The CJS service
only passes request data in and returns that result.
…covery (#15865)

* 📇 fix: Stop an MCP Credential Refresh From Fencing Its Own Catalog Recovery

Passive catalog recovery captures the tool-cache generation before discovering a
cold server and discards the catalog when the stored generation has moved by the
time discovery returns. A silent OAuth refresh inside that same discovery
publishes a rotation of its own and then clears this replica's recovery state,
so the tracker drops the flight as superseded and the generation check rejects
it: the server is listed with no tools, and a concurrent catalog request starts
a second discovery instead of joining the first.

The authorization fence now reaches discovery through recovery's own
dependencies, which report each generation a discovery's refresh publishes. The
tracker adopts that generation while the flight is still running, so the flight
stays current and its outcome carries the generation it produced; a rotation or
clear from any other writer afterwards still supersedes it, and a refresh that
completes after its discovery settled no longer touches that settled state.

* 🪢 fix: Keep a Recovery Flight Joinable Through Its Own Credential Publication

A refresh's authorization fence writes the new shared generation, clears local
recovery state, and only then reports what it wrote. A catalog request that read
the new generation during that span found the flight still keyed to the old
generation, or its slot vacated by the clear, and replaced it with a second
discovery; the original flight was then discarded.

Discovery now runs its publication through a tracker bracket. While a flight's
own publication is in progress the tracker holds any clear for that server and
lets requests for the same config join the flight; when it ends, the flight
adopts the generation it wrote, or applies the held clear if the publication
failed or the flight settled first. The flight is registered before its
discovery starts, so a publication can never precede its own registration.

* ⏳ fix: Bound and Count What a Recovery Flight Holds During Its Own Publication

A recovery flight held every clear for its key and admitted any request for its
server for as long as its own credential publication stayed open. A clear from
another writer collapsed into the fence's own clear and was dropped on adoption,
leaving a backoff or reauthorization decision in place after a teardown; a
publication stalled in an unbounded storage write kept every later request
joined to the stalled flight; and a joiner that had already read a foreign
generation could keep tools under an older adopted generation when the final
generation read failed.

An open publication now holds clears and admits joins only while its flight is
in flight and for at most one discovery budget, after which requests take over
as before. The fence clears local recovery once when it completes, so any other
held clear still applies after the flight adopts its generation, and every held
clear applies when the publication fails or the flight settles first. An outcome
under a generation the request never observed stands only when the final shared
generation read confirms it.

* 🧬 fix: Correlate Catalog Recovery Clears With the Generation Their Fence Wrote

Local catalog-recovery clears carried only a user and server, so the tracker
could not tell a flight's own credential publication from any other writer. It
counted clears and bounded its hold with a time window instead, which left a
request that joined before the window expired waiting on a stalled publication,
and let a late fence clear delete the flight that took over after the window.

The authorization fence now hands local recovery the generation it wrote, and a
clear spares state already recorded under that generation; a clear without one
still clears. A flight holds the clears it receives while its own publication is
open, adopts the generation that publication wrote, and then applies each held
clear against it, so its own clear is spared while a teardown or another
publication's clear still applies. Retry-intent cleanup inside the fence is
bounded by the same per-attempt timeout as the generation write, since a leftover
intent only advances the generation again, so the publication itself can no
longer hold its joiners open and no time window is needed.

* 🔏 fix: Confirm an Adopted Catalog Generation for Requests That Read None

A request whose pre-discovery generation reads both failed observed no
generation, so the confirmation rule exempted it. When that request's own
discovery refreshed credentials, adopted the generation its publication wrote,
and another replica rotated past it before a final read that also failed, the
catalog was served without ever being confirmed against shared state.

A flight's own result now records that its generation came from its own
publication, and that mark is never copied into retained state. An outcome
under a generation the request did not read needs final confirmation whenever
the request read a different generation or the generation was adopted; an outcome
whose generation the tracker merely carried over for a request that read nothing
keeps the existing fail-open behavior, so a cache outage still preserves retained
backoff and reauthorization state.

* 🧭 fix: Judge In-Flight Catalog Recovery by the Shared Generation, Not a Tagged Clear

A publication's local clear carries the generation it wrote, and the tracker
deleted any state recorded under a different one. Generations carry no order, so
a refresh that outlived its own flight could clear with an older generation and
delete a flight that had started under a newer rotation from another replica,
discarding a catalog that was still current.

A tagged clear now only signals that the shared generation advanced. A flight
still in flight under a known generation, or publishing its own credential
change, is left to the requests awaiting it, which already confirm their outcome
against the shared generation exactly as they do for a rotation from another
replica. Retained state and flights with no known generation are cleared unless
recorded under the clear's generation, and a clear without a generation still
clears unconditionally. The tracker no longer holds clears during a publication:
the refresh's own clear spares its flight directly, and the flight adopts the
generation it wrote when the publication completes.

The shared-discovery tests compared catalog maps through expect.objectContaining,
which does not compare Map contents; they now assert each request's catalog with
toEqual.

* 🧾 fix: Confirm a Catalog Flight That a Clear for Another Generation Spared

A tagged clear spares a flight still in flight, because an unordered generation
cannot say whether that flight predates the rotation. The spare was not recorded,
so when the request's final shared-generation read timed out, a flight that had
observed its own generation kept a catalog the clear may have superseded: an older
recovery whose refresh published after a replacement flight started could leave
that replacement serving tools no read ever confirmed.

A spared flight now records the generation each such clear carried and settles
against its final generation. A clear for its own generation, including its own
refresh's, changes nothing; any other marks the flight's result unconfirmed, keeps
it out of retained state, and requires every request awaiting it to confirm the
generation against the shared read before using its tools. A spared flight that
finishes with no known generation is discarded, as the clear would have done. A
tagged clear no longer spares retained state recorded under its generation.
* feat: add run-scoped file sharing for subagents

Delegate current-message files through authorized execution catalogs and publish selected immutable code-output versions through the existing file pipeline.

Requires the companion agents SDK context adapter before the dependency can be released and pinned.

* fix: preserve child tool artifacts and verify file-sharing lifecycle

Keep citations and other non-sandbox artifacts on their existing delivery path while retaining private code outputs and inspected images until publication.

Cover native and text delivery, nested and sibling authorization, concurrent runs, cancellation cleanup, and checkpoint recovery through the browser harness.

* test: align file-sharing fixtures and integration cleanup

* test: repair LibreOffice conversion fixtures

* test: verify deployment skill bytes across checkout formats

* fix: align subagent file budgets and refresh boundaries

* chore: update agents SDK to 3.8.6

* refactor: move code output processing into API workspace

* fix: enforce shared file budgets for parents and teams

* fix: declare the shared file encoder interface
* ✨ feat: Add Upstream Model Error Observability

* ♻️ refactor: Move Terminal Error Policy to TypeScript

* 🛡️ fix: Preserve Agent Cancellation Semantics
* ✨ feat: Surface Upstream Model Errors

* 🧪 test: Cover Upstream Error Rendering

* 🛡️ fix: Harden Upstream Error Fallbacks

* 🛡️ fix: Defer Upstream Error Fallbacks
* feat: lock code environments per conversation

* style: Sort code environment imports

* fix: Harden code environment decisions

* fix: Preserve code environment rollout compatibility

* fix: Harden code environment boundaries

* refactor: Centralize code environment policy

* test: Restore partial disconnect policy mock

* fix: Close code environment edge paths

* fix: Preserve implicit stateful code routes

* fix: Stabilize implicit code routing

* test: Align remote workspace contract

* chore: Sort code route import
…d Tokens (#15871)

* 🪂 fix: Release Stalled MCP Catalog Recovery Without Dropping Refreshed Tokens

Passive catalog recovery could hang forever when discovery hit an expired
OAuth access token and the refreshed tokens' authorization-fence write
stalled. The shared flight, every request coalesced onto it, and their
process-wide catalog slots waited with it until capacity was exhausted.

- Stop discovery's OAuth token wait at its deadline or abort without
  cancelling the shared token flow, so redeemed tokens still persist.
- Stop recovery waiting on a discovery once its budget plus
  `mcpSettings.catalogRecovery.discoverySettleGraceMs` (default 10s) has
  passed, releasing the flight, its joiners, and their slots with backoff.
- Add `waitUntilDeadline` for waits that must end without cancelling work.

* 🪂 test: Rediscover Once a Released Flight's Refresh Publishes

* 🪂 fix: Hold a Server's Recovery Until Its Released Discovery Settles

* 🪂 fix: Scope Released-Discovery Holds to Observed State Under a Process-Wide Limit

* 🪂 fix: Count Token Work a Discovery Hands Off at Its Deadline
* 🔎 fix: Enforce WEB_SEARCH Role Permission on Agents

Squash of #15822 onto dev, preserving authorship:

- 🔎 fix: Enforce WEB_SEARCH Role Permission on Agents
- 🧪 test: Cover the Runtime Tool Loader's Web Search Gate
- 🔒 fix: Gate Web Search on the OpenAI-Compatible Embedder Surface
- 🩹 fix: Authorize Web Search by the Execution Context's User
- 📝 docs: Note the Grant Memo's One-Subject-Per-Request Assumption
- 🧹 fix: Deny Web Search Before an Endpoint Default Can Grant It
- 🙈 fix: Withhold the Web Search Model Parameter From a Denied Role
- 🧱 fix: Strip Provider-Native Search an Endpoint Re-Enables After Denial
- 🔌 fix: Close the OpenRouter Plugin Path and Stop Pruning Hidden Parameters

* 🧭 fix: Gate Provider-Native Web Search on the Built Config

The web search role gate ran before the provider builder and only when the
agent's stored `web_search` was not `false`, so an endpoint's `addParams`
re-enabled native search for a denied role, and every OpenAI-compatible and
Responses request read the role whether or not search was in play.

The gate now reads the builder's output: when it carries a native search tool
or OpenRouter's `web` plugin, the `WEB_SEARCH` grant is resolved and the search
stripped on denial. The routes that reach `initializeAgent` with `runtime` and
no `req` pass a lazy `resolveWebSearchGrant` that joins their request-memoized
grants, forwarded through agent discovery to handoffs and subagents. The
embedder surface wires it whenever `getRoleByName` is supplied, independent of
`appConfig`.

---------

Co-authored-by: TomasPalsson <tomas2212042710@gmail.com>
* 🔭 feat: Conversation Trace Viewer

Adds a trace view for a user's own conversations: a summary, a pinned
waterfall overview with interval focus, a hierarchical record ledger and
a record inspector, opened from the chat header (or the mobile overflow
menu) and read through a provider-neutral seam with Langfuse as the
first reader.

- interface.traceViewer (enabled, showInputOutput, maxRecords,
  maxContentLength, requestsPerMinute), off by default
- /api/traces/:conversationId availability, records and record routes,
  re-proving conversation ownership and narrowing the Langfuse session
  to traces the user's own sampled responses produced
- Langfuse reader over the v2 Observations API with cursor paging, time
  bounds, typed failures and server-side input/output withholding
- forks no longer copy the source messages' trace sampling record

* 🛂 fix: Address Trace Viewer Review Round 1

- interface.traceViewer.requestTimeoutMs bounds each Langfuse round trip
- availability is one existence query across every sampled response
  instead of the oldest one, run alongside the ownership check
- tree items report their position among visible siblings
- point-in-time Langfuse events read as completed, not running
- reads prefer the project holding the most responses when a tenant
  connection arrived mid-conversation

* 🎯 fix: Leave Focus With a Navigation That Closes the Trace

* 🧭 fix: Keep Trace Reads on One Project and Cancel Abandoned Reads

- page cursors and record details stay on the project that served the
  first page; pages report an opaque sourceId
- Open in Langfuse only appears when the linked project served every
  loaded page (the session link now reports its destinationId)
- client and server abort Langfuse reads when the viewer closes or the
  client disconnects
- reads never wait on the central project lookup, which the request
  timeout does not govern
- the cost total is withheld when any model call with usage is unpriced

* 🕰️ fix: Refresh Trace Pages After a Turn, Honor Clock Settings, Show Refresh Failures

- a settled run invalidates the cached trace pages, so reopening shows
  the new turn instead of pages fresh for another 30 seconds
- trace durations and clock times format in the app language and the
  user's 12/24-hour setting, as message timestamps do
- a failed refresh with every page loaded surfaces an alert and a retry

* 🔐 fix: Only Server-Authored Responses Grant Trace Access

- the message-create route and imports strip langfuseSampled and
  langfuseDestinationIds, and trace reference queries ignore client-authored
  rows (isUserSubmitted) and user turns, so a forged row cannot claim a trace
- the Langfuse reader receives its destination resolver and HTTP client
  from the route instead of defaulting to process-global ones
- availability asks the client to check again while central's project id
  is still resolving, instead of caching a false answer
- listRecords tries the next readable project when legacy responses left
  the first one empty
- a settled run also invalidates cached record details
- running is read from a record's status, not a missing end time

* 🏢 fix: Scope Trace Reads to the Tenant and Reach Every Project Holding a Turn

- ownership and both trace reference queries carry the requester's tenant
  (tenantless is its own scope), so a duplicate user and conversation id in
  another tenant cannot grant or supply trace references
- a read visits, page by page, every project that could hold a turn no
  earlier project provably holds, so legacy turns in another project stay
  reachable after a newer destination ranks first
- the trace read limiter takes its store from the route
- transient availability failures retry instead of hiding the control

* 🛟 fix: Fail Over Between Projects, Refresh One Page, Search Visible Labels

- a fresh read fails over to another project holding the same turns when the
  preferred one fails; the cursor carries the failed sources so later pages
  keep the rebuilt plan, and detail reads fail over the same way
- destinations for one project collapse into one source
- the trace read limiter keys each user within their tenant
- refresh and a settled run keep only the newest page before refetching, so
  older pages are not replayed through the limiter; retry repeats the read
  that failed
- search matches the localized kind and status labels a row shows

* 🧭 fix: Read the Newest Turn First, Report Failures Behind Empty Projects, Fit Trimmed Views

Rank trace sources by the newest sampled response they could hold before coverage and preference, so a conversation that moved projects shows its latest turn first.

Throw a list failure when a failed project could hold a turn that no answering project provably holds, instead of rendering an empty trace. Detail reads keep probing projects after an empty answer, skip ones whose turns an answering project provably holds, and surface a failure that could hide the record.

Refresh and settled runs trim the cache to the newest page; the viewer now fits a focused interval to the remaining records or clears it, and drops a selection whose record left with the older pages.

* 🧵 fix: Page Traces by Turn, Include Failed Runs, Keep the Active Row Mounted

List reads now page by turn segments instead of whole projects. A segment is the newest unread turn and up to 49 older consecutive turns one project could hold. The request filters Langfuse to those turns' run and title traces, so pages stay newest first when a conversation's responses alternate between projects. Legacy turns without destination ids are asked of the other projects in the segment. A project that failed for newer turns is retried when it alone holds older turns. A page whose rows all fail the observation schema is an upstream failure, not an empty trace.

A failed turn's error row now stores the sampling fields of the run that failed, plus langfuseRunId when that run was created, so the trace viewer and feedback scores follow the run's trace. langfuseRunId is stripped from client message writes and imports, dropped by forks, and hidden from client message projections.

The ledger keeps its active tree item mounted while virtualization scrolls it out of the window, so aria-activedescendant always names a rendered row.

* 🕰️ fix: Merge Multi-Project Trace Pages by Start Time, Keep Literal Content

A segment that reads more than one project now runs every read before returning. A read that fills the record budget cuts the page at the oldest start time all reads loaded, and the continuation carries that start time as an inclusive bound instead of a Langfuse cursor. Every project continues from the same point, so a newer turn held by another project is never left behind an older one, and failover works on any page. A page drawn from several projects no longer names a single source. A continuation that would repeat its own bound fails instead of looping.

Input and output that are literally the strings {} or null are kept; only an empty object or JSON null value is dropped.

* 🔎 fix: Read Each Trace Segment From a Project Shown to Hold It

Recorded destination ids say where a trace was eligible to go, not where it landed. A list read now looks up the turns more than one readable project could hold before choosing where to read them. Each project, preferred first, is asked which of those traces it holds a root observation for, and each turn is read from the preferred project that holds it, or from the preferred eligible one while a run's root is not exported yet. Turns only one project could hold need no lookup.

Segments are consecutive turns from one project, paged by that project's own Langfuse cursor. Records that share a start time page correctly, and the page budget bounds a single read. A failed project is dropped for the request. A turn only it could hold raises its failure at the turn's own place in page order, instead of leaving an empty or partial trace. Cursors carry the segment's trace-set hash and start over when the segment changed or Langfuse rejects them. Detail reads ask every project until the record turns up.

The trace summary withholds its cost total when any generation has no cost, with or without usage.

* 🧷 fix: Evidence per Trace, Continuations Without Replays, Stable Turn Order

A turn's run and title run are exported separately, so each trace id is looked up and read from a project shown to hold it. A trace no project shows (a title run the turn never had, or a run whose root is not exported yet) joins an adjacent segment whose project could hold it. Only a run that a failed project alone could hold raises that failure.

A continuation now stays on its project. Its failure is returned as is, and a cursor the segment no longer matches, or one Langfuse rejects, returns invalid_request instead of replaying the segment's first page. The viewer answers invalid_request on an older page by reloading from the newest page, with a message saying the trace changed.

Sampled responses sort by createdAt then _id, so responses saved in the same millisecond keep one order across page requests.

* 🪪 fix: Scope Trace Reads to the Exporting User, Order Split Title Runs

Trace ids derive from message ids, which a request can influence, so a colliding id must not return someone else's trace. Every Langfuse list, probe and detail read now also filters by userId: the requester's internal id, plus the configured langfuse.trace.userIdField value when set. The server stamps that id on the trace at export. The handler passes the requesting user into the trace query.

When a turn's run and title run sit in different projects, the probes' root start times decide which is read first, so a title generated after the response is not left behind Load older. Cursors name the trace a segment starts at within its turn.

The failed-turn trace lookup moves into packages/api as getFailedTurnTraceFields. The agent controller only passes the error row id, the run id and whether the run was created.

* 🔒 fix: Authorize Trace Reads by the Internal User Id Only

A configured langfuse.trace.userIdField can export a non-unique value (name, username) or one that repeats across tenants (email, provider ids) as the trace userId, so it cannot prove who a trace belongs to. Reads now require the internal user id alone. A deployment that exports another field reports no trace and logs why once, instead of trusting that value.

A failed read of a title run that a probe found in that project now raises its failure instead of dropping the title. The page contract now states that pages are ordered by turn and a turn's records may continue on the next page.

* 📚 fix: Page Sampled Responses With the Trace, Share Ingress Stripping

List reads load sampled response references a segment at a time: up to 50 ending at the cursor's turn, plus one older response that names the next page. Pages of a long conversation no longer rematerialize every sampled response. getConversationTraceRefs gains through and limit, anchored by createdAt and _id.

The message-create route and conversation imports strip trace sampling fields through withoutTraceRefs instead of their own field lists. A userIdField the exporter ignores counts as the internal id. A root lookup whose cursor repeats fails that project instead of paging forever.

In the viewer, Refresh also rereads the open record's detail, and closing the inspector after a filter removed every row returns focus to the search field.

* ⏱️ fix: One Round Trip per Trace Reference Read, Detail Reads Scoped to Their Turn

Cursors carry the anchor response's order key (creation time, then _id). The bounded sampled-response read then runs alongside the first-message lookup, with no anchor lookup ahead of it. A key that no longer names its response returns nothing, and the reader answers the cursor as changed.

Detail reads take the turn the list attributed the record to (?message=). They load only that sampled response instead of the conversation's whole history, and ask Langfuse within that turn's traces. The route rejects a detail read without it, and the client sends it from the listed record.

A list read whose Langfuse cursor repeats now fails instead of handing the next page the same records.

* 🧭 fix: Reject Out-of-range Trace Cursors, Open Any Listed Turn, Mark Copies Unsampled

A page cursor whose base-36 time parses to a number outside the Date range reached Mongoose as an Invalid Date and surfaced as a 500 the viewer retried; it now resumes nothing, which the reader reports as invalid_request so the viewer reloads its newest page.

The detail read capped the message id at the Langfuse record-id bound while the app stores response ids at any length, so a listed turn could refuse to open; message ids only name a stored row and are accepted at any length. Ids that reach a Langfuse filter keep their caps.

Copied and client-authored rows now carry langfuseSampled: false instead of no record, so a feedback score on a fork, import or posted message never recomputes sampling for a trace that was never made.

* 🧪 test: Type a Recordless Row for the Unsampled-copy Spec

An object literal with none of the trace fields fails TypeScript's weak-type check against the helper's all-optional constraint; the spec now passes a typed row that declares the optional field, as every stored message type does.

* 🚰 fix: Release Failed Langfuse Response Bodies

A non-success Langfuse response was turned into a TraceReadError without consuming or cancelling its body, and Node's fetch returns a connection to its pool only once the body is released, so a sustained upstream failure left sockets checked out until garbage collection. The body is now cancelled before the status becomes an error.
* fix: Preserve Disabled Workspace Bindings

* fix: Resolve Restored Workspace State
* 🪪 fix: Adopt MCP OAuth Callback Tokens Without Re-Prompting or Losing the Lease

A chat request loading OAuth tokens joins the `mcp_get_tokens` flow a catalog
discovery left behind and polls it. The OAuth callback deletes that entry when
it stores the new tokens, so the poll ends with "Flow state not found", the
factory reports the tokens missing, drops the `mcp_oauth` flow the callback
just completed, and starts a second interactive flow. Once tokens arrive, the
connection leases the publication generation it captured before the wait; the
callback's publication moved it, so the connection is torn down as it connects
and the agent fails with AGENT_EXPECTED_MCP_TOOLS_UNAVAILABLE.

`createFlowWithHandler` now serves a completed result at once, monitors only
pending attempts and retained failures, and replaces a failure nobody retains;
a waiter whose flow disappears gets `FlowStateNotFoundError`. The factory
treats that as a credential change: it re-reads storage once, after letting
the connection build re-capture its generation. The authorization transaction
and a silent refresh attach the generation they publish to the tokens they
release, and a build adopts that generation for credentials obtained after
its capture.

* 🪪 fix: Claim Flow Attempts Atomically and Lease Under the Stored Generation

A connection now adopts a carried publication generation only while it is the
generation currently stored, instead of comparing the token's exchange time to
its capture, so a callback that exchanged before the build began and published
after it is still adopted and a retired generation still leaves the fence in
place. Installing a flow attempt is an atomic compare-and-set against the
attempt the caller observed, on Redis and on the default in-memory store, so
concurrent replacements of a failed attempt, or concurrent creations of an
absent one, run one handler while the rest join it. An aborted caller is
rejected before a cached result is served, and a joiner that aborts leaves the
shared flow to its owner and the other waiters.

* 🪪 fix: Return Generation-Tagged Callback Tokens and Route Flow Claims Through the Cluster Executor

The authorization transaction now returns the same generation-tagged tokens it
releases to the OAuth completion and the pending token waiters, so the callback
route's tool-flow completion carries the published generation to the connection
waiting on it. The Action OAuth and index-sync flow managers pass the
cluster-aware Redis script executor, so their flow claims and guarded
settlements route through slot-master selection and READONLY recovery instead
of the raw client.
* 🚀 chore: prepare v0.8.8

* 🚀 chore: prepare v0.8.8-rc3

* docs: mention Agent workspace defaults
@pull pull Bot locked and limited conversation to collaborators Sep 13, 2026
@pull pull Bot added the ⤵️ pull label Sep 13, 2026
@pull
pull Bot merged commit e18606e into innFactory:main Sep 13, 2026
23 of 24 checks passed

This branch had an error being deployed

1 failed deployment
publish — e18606e5 Deployed Sep 13, 2026 by pull[bot] via publish-npm #25
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants