[pull] main from danny-avila:main - #260
Merged
Merged
Conversation
* 🧭 feat: Remember Attached Workspace Choices * Guard Workspace Defaults Across Navigation and Version Restore * Order Workspace Preference Imports * fix: Harden agent workspace defaults * fix: Type workspace validation schema * test: Cover safe execution projections * fix: Bind workspace defaults to attached environments * fix: Preserve unchanged workspace bindings * fix: Scope workspace preferences to active agents * fix: Preserve workspace selection state * test: Target workspace discovery status * fix: Keep workspace reset available * fix: Resolve workspace defaults per environment * fix: Collect per-environment workspace defaults * fix: Reconcile Deleted Agent Graph Caches * fix: Reject Stale Workspace Bindings
* 🛡️ fix: Prevent ReDoS in skill and artifact parsing * 🧵 fix: Normalize artifact fences in linear time * 🧠 fix: Bound artifact normalization memory
* fix: Preserve Directory Identity Ownership * fix: Fail Closed on Identity Conflicts * fix: Resolve Existing Sharing Principals Safely
* fix: Prefer storage keys for image encoding
* fix: Sanitize assistant image errors
* fix: Sanitize agent image storage errors
* fix: Move blob image encoding into API package
* fix: Complete image encoder TypeScript migration
* fix: Keep request-boundary error logs diagnosable
getSafeErrorMetadata reduces an error to { type: 'Error' } and an optional
status. That is the right bound where an error may echo submitted content, and
it stays in place at getS3FileStream, getAzureFileStream and
getFirebaseFileStream, where a signed URL reaches the log.
Applied at ResumableAgentController's initialization catch and the chatV1
generic branch it blinded the whole route: a missing key, an agent that fails to
load and a database error all logged as { type: 'Error' }, with no message,
name, code or stack.
getSafeErrorText keeps the error's own description and stack, reduces every URL
to its origin so an object path and signature cannot ride along, and strips
bearer credentials. It returns text so the two handlers log it inside the
message and pass no winston metadata, which is what stops an SDK error's own
enumerable properties (request.url, config, response.headers) from being
promoted onto the record.
---------
Co-authored-by: Danny Avila <danny@librechat.ai>
…#15863) A durable user connection captures the tool-cache publication generation before resolving credentials, so a rotation from another replica fences the catalog it publishes. A silent OAuth refresh performed while that same connection is being built publishes a rotation of its own, and the build then fails the lease check it just invalidated: the connection holding the newest credentials is torn down and the tool call that triggered it errors with "Publication lease is no longer current". invalidateCachedTools now reports the generation it wrote, the authorization publication carries it back to its caller, and a connection build adopts what its own refresh published. A foreign rotation still moves the stored generation past that value, so the fence keeps failing the build it was meant to fail.
* 🔎 feat: Filter and Sort the Sidebar Chats List The sidebar could only list every unarchived chat by last activity. Archived chats lived behind a settings modal, sorting was not exposed at all, and the only filter was a bookmark dropdown squeezed next to the search field. A single control beside the Chats heading now carries the status (active or archived), the sort field and direction, and the bookmark filter, each scoped to that list alone: Projects, Pinned and Favorites are untouched by it. * 🧹 fix: Answer Review Round on Sidebar Chat Filters Rows derive their archive action from `conversation.isArchived`, which the list projection now carries, instead of from the sidebar's filter mode: the archive view keeps the unarchived pins beside it, and its previous page's rows outlive a status switch, so the list mode offered to restore chats that were never archived. Mutations reconcile the archived list caches they can now reach — rename, pin, tags, project assignment, shared links — and archiving refetches the archived query instead of inserting into page one, which dropped the row whenever the active sort placed it later. Title grouping keeps the order the server paged, heads each group with the title's first code point, and the archive's legacy group falls back to `createdAt` the way the cursor orders it. A failed list request renders an error with Retry rather than "No chats yet", the status setter sanitizes a malformed stored sort, and signing into another account no longer inherits the previous one's filters. * 🧵 fix: Keep Sidebar Cache Writes Inside the Chosen Sort Title groups are emitted as contiguous runs, so a `#` heading that the server's order places either side of a letter no longer merges and drags a row across the cursor boundary the next page continues from. The live cache writers only reproduce newest-first, which is the one order they can place a row in without the server's keyset. A title, created-at or ascending variant now takes the field update in place and is refetched instead of having rows moved to or seeded at the front of page one. Archive and restore refetch every loaded page of the lists they touch — the archive preserves `updatedAt`, so a restored chat can belong on a later active page — while inactive variants are left stale rather than eagerly refetched. Archiving everything still refetches them all, because every cached list is wrong afterwards. Renaming or deleting a bookmark now rewrites the sidebar's selected tags, which otherwise kept filtering by a name no menu could uncheck. * 🗂️ fix: Treat the Archived List as a Live Conversation Cache * 🔍 fix: Refetch List Variants Whose Membership the Client Cannot Decide A cached search result is evaluated by the server, so a newly created row may or may not belong in it. Skipping those variants left a mounted search list missing a row it should show — forking from an archived search result being the case that exposed it. The insert decision is now three-valued: insert into a variant whose membership and order the client can reproduce, skip one the row provably does not belong to, and refetch anything undecidable. Forks, duplicates and imports inherit the source's archive state, so their invalidations cover the archived list too, and the delete path reads `chatProjectId` from archived caches as well. * 🔎 fix: Reconcile Searched List Variants on Every Write * 📣 fix: Announce Archive and Restore From the App Live Region * 🧹 fix: Drop Unmounted Lists on Archive-All and Fill a Missing Archive State Archive-all refetched every loaded page of every cached list, mounted or not, so a session that had visited several sort, bookmark or search combinations paid a burst of requests for lists nobody was looking at. Mounted variants still refetch all pages; the rest are dropped, which reconciles them more thoroughly than a refetch and costs nothing until something mounts them. A row's own `isArchived` decides whether its menu offers Archive or Restore, and this branch is what adds that field to the list projection — so mid-rollout, an archived row from an older backend would offer Archive and submit a no-op. The list query now fills a missing value from what the variant asked for, leaving an explicit one alone. * 🧭 fix: Give Callers One Way to Reconcile Both Conversation Lists Callers that cannot say what changed — a stream recovered after a disconnect, a schedule whose run moved a chat, a project that took its chats' fields with it — were invalidating the active list by hand, so a mounted archive kept a row the server had already changed. `invalidateConversationLists` reaches both prefixes and is now the only way those paths reconcile; `findConversationListQueries` shares its key list, so adding a prefix is one edit rather than a sweep. * 🔤 fix: Follow the Server Order, Index the Created Sort, and Tell the Truth When Empty * 🎯 fix: Identify the Open Chat When a Restore Resolves A restore leaves the chat open, so its success callback writes the new archive state into the active conversation — but the callback captured that conversation when the menu item was clicked. Open another chat before the request lands and the identity check still matched the captured row, replacing the conversation the user had just opened while the URL kept pointing at it. The write is now a functional update that compares `conversationId` against current state, so a stale callback leaves the open chat alone. Reading the conversation value is no longer needed either, which also stops every row's menu from re-rendering on conversation updates. * 🧷 fix: Recheck the Open Route When an Archive Resolves
* 🪪 style: Give Native Tools a Painted Verified Mark The native badge on a tool card was lucide's BadgeCheck filled with status-success-strong: green, the same hue the card already uses for selected, and stroked in currentColor, so the white check color also outlined the badge and left a white halo on any surface that is not the card's own white — hover, selection, dark mode.\n\nVerifiedIcon paints the badge instead of stroking it and carries the check as the only stroked path, and status-verified gives provenance its own blue across all four themes. * 🩹 fix: Hold the Verified Mark at 3:1 on Every Card State The dark fill was tuned against surface-secondary, but ToolCard repaints to surface-tertiary (#2f2f2f) on hover, where #0b74d4 fell to 2.85:1 — below the non-text floor, and invisible to the spec, which asserted the panel only. The fill moves to #1a7fd8 and the contract now covers every background the card takes. Both of the mark's relationships are graphical objects under WCAG 1.4.11, so the check is held at 3:1 rather than AA: no hue clears 3:1 on the hover surface and 4.5:1 under a white check at once.\n\nA theme that predates the token names status-success-strong and not status-verified, and would otherwise wear LibreChat's stock blue beside its own text-on-status and card surfaces. resolveTheme and the legacy mapColors adapter now carry the mark onto the fill it had before the token existed, following the six derived-token fallbacks already there. * ✅ test: Exercise the Verified Mark in the Tool Library Four scenarios in the mock harness, each named by the behavior it observes: the mark a native card wears, its silhouette under the dark hover repaint that the retune fixes, a pre-token deployment palette carrying it on the theme's own success fill, and a third-party MCP server card carrying no mark at all. The contrast assertions read the rendered colors rather than the tokens, so a regression in the component or in theme application fails them. * 🧪 fix: Teach the Marketplace Dialog Spec About VerifiedIcon The suite hand-lists the @librechat/client surface it mocks, so ToolCard's new icon resolved to undefined and every render in the file threw. * 🧪 fix: Reach the Agent Builder on a Narrow Viewport The scenario run covers mobile emulation, where the side rail is collapsed and the panel opens from the header's Control Panel toggle; the helper only knew the rail. The MCP scenario also minted its access token before the app had loaded, so the refresh request had no origin to resolve against. * 🧪 fix: Run the Verified-Mark Scenarios on a Touch Viewport Mobile emulation has no side rail: the panel switcher lives inside a sidebar that starts off-canvas, so the helper opens the sidebar and picks the Agent Builder entry from the switcher's menu. A touch viewport also has no hover state, so the hover scenario measures whatever background the card actually takes there and still holds the mark to 3:1 against it. * 🩹 fix: Keep Pre-Token Themes on the Fill the Mark Inherited The first fallback only fired when a theme restated status-success-strong, so a partial palette that customizes surfaces or text-on-status and inherits the green still landed on the stock blue. resolveTheme now derives the omission for any custom theme from the resolved success fill, and the legacy adapter takes the mode's bundled palette from ThemeProvider so it can do the same for a theme that names neither.\n\nThe scenario helper picked its path from a one-second probe for the desktop rail, which turns a slow first paint into the mobile branch; it now decides on viewport width and keeps the suite's full assertion timeout in each branch. The mock project is Desktop Chrome, so a touch scenario asks for a touch context explicitly rather than leaving that layout unexercised in CI. * 🧪 fix: Match the App's Inclusive Layout Breakpoint useSidebarState switches on (max-width: 768px), so a 768px-wide context renders the sidebar switcher while the helper's strict comparison went looking for the desktop rail. * 🩹 fix: Scope the Verified Fallback to Themes That Painted the Mark Two rounds pulled the same predicate in opposite directions: keying it on a theme restating status-success-strong missed palettes that only repaint the check or the card, and keying it on any custom color at all gave a new partial theme a green mark identical to the selected cue. Both are answered by asking what the theme actually painted — the fill the mark wore, the text-on-status check it carries, or the surfaces ToolCard takes — which is the same shape as the series-scale fallback above it. A theme that repaints none of those keeps the bundled blue. * 🧪 fix: Wait for the Shell Before Reading the Sidebar State isVisible() answers immediately, so a slow first paint read as an already-open drawer and the helper then waited on a switcher sealed inside it. It now waits for whichever of the opener and the switcher the shell paints first, which is the same assertion timeout the desktop branch gets. * 📝 fix: Keep the Token Catalog Entry Catalog-Sized The entry restated the icon's construction, the contrast rule and the pre-token fallback, all of which live on the token's own doc comment and in semanticTokens.spec.ts; the catalog around it is one line per role. * 🔦 fix: Let a Crashed Lighthouse Run Be Retried The audit shells out to Lighthouse three times and fails the lane if any single process exits nonzero, so a NO_NAVSTART — Lighthouse losing the trace of a page load, whose own message is "run Lighthouse again" — sinks a pull request that has nothing wrong with it. That is what happened on this branch: runs 1 and 2 wrote reports, run 3 crashed, and the identical head passed on rerun.\n\nEach run now gets a second attempt before the failure propagates. The repeat is per run, so the median is still taken over three completed reports, and a page that genuinely cannot be traced fails both attempts and still fails the lane. * 🧪 fix: Read the Drawer State From Its Inert Marker The readiness wait accepted the panel switcher, which is mounted whether or not the drawer is open — translated off-canvas and inert — so a shell that painted the sidebar before the chat header could satisfy it with the drawer closed and leave the helper clicking an off-viewport control. The drawer's own inert attribute is the state, so the helper waits for the drawer to attach and reads it there, with no timing probe on either side. * 🧪 fix: Open the Drawer Before Trusting the Restored Form A narrow context that had already selected the builder restores the form inside the closed drawer, and a translated element passes a visibility check, so the early probe accepted a form sitting under an inert drawer and handed it to the caller. The layout decision now comes first and the drawer is opened before anything inside it is read.
…#15866) * 🧪 test: Restore Two Rotted packages/api Integration Suites MCPReinitRecovery.integration.test.ts mocked @librechat/data-schemas down to its logger, so all eight tests threw once MCPServersRegistry started calling scopedCacheKey (#15028). The mock now spreads the real module under the logger override; ServerConfigsDB stays mocked, so the suite still never touches Mongo. secrets.integration.spec.ts drove langfuse.secretKey through the generic admin config API, which has stripped the langfuse section since #14108 moved it to the dedicated Langfuse connection API. The generic-write matrix drops that row, and a langfuse block covers the current contract against a real Config collection: the stored secret stays encrypted at rest, generic reads redact it, generic patch, tombstone and field-delete writes leave it untouched, and a full base-config replacement carries it forward. * 👷 ci: Run Every packages/api Integration Suite in CI test:ci excludes every path containing "integration.", but the jobs meant to run those suites selected by folder: the cache workflow only src/{cache,cluster,mcp,stream}, the agents workflow only src/agents/*.integration.spec.ts. Thirteen suites ran in no job, including every *.integration.test.ts. test:integration selects *.integration.(spec|test).ts anywhere under src and replaces test:agents-integration in the (renamed) Integration Tests workflow. The cache scripts also accept .test.ts, stream selects by suffix alone, and core picks up src/middleware. Every suite test:ci excludes is now selected by exactly one CI script. * 🧪 test: Read the Tool Snapshot in the Concurrent Reinit Simulation The concurrent reinitialize test copied reinitMCPServer but still called fetchTools(), which returns [] when tools/list fails. Concurrent reinspections of a recovered server can each succeed, and a later write replaces the app connection an earlier request was already handed, so that request's tools/list fails. The test failed in 9 of 25 iterations on Linux, and in the first CI run that included the suite. Since #14686 production reads the snapshot and keeps cached tools when it is incomplete; the simulation now does the same, and still requires every complete snapshot to list both tools and at least one to complete. A deterministic test pins the property that tolerance relies on, without depending on the race: a newer config write replaces the held app connection, whose snapshot is then incomplete, while the replacement lists the real tools.
…15867) * 🪢 fix: Recover a Failed MCP Server Once When Reinspections Overlap A server stored as an inspection-failed stub is recovered by the first request that inspects it successfully, but every request that read the stub before that write inspected again and wrote again. Each later write bumped `updatedAt` past the connection made after the earlier one, so the connection was replaced and a caller still holding it listed no tools. A request that reached `reinspectServer` after the write hit the "not in a failed state" guard and reported the server as still unreachable. - Share one in-flight reinspection per stored entry and effective allowlists in each process; DB entries are keyed per user and tenant. - Replace the stub only while it is still the inspected stub: a new optional `replaceStub` compare-and-set on the in-memory and Redis aggregate-key stores (one Lua script), so replicas never overwrite a recovery another replica stored. - Resolve a reinspection that finds the server already recovered to the stored config, and move the reinitialize decision into `MCPServersRegistry.recoverServerConfig` so `reinitMCPServer` connects with the recovered config instead of the stub it read. - Pin the interleavings in `MCPReinitRecovery.integration.test.ts` against a real in-process MCP server, and require every concurrent reinitialize flow there to recover with a complete tool list. * 🪢 fix: Key Shared Reinspections by the Stub They Inspect A registry re-initialization can replace a failed stub while an older reinspection of it is still running. The shared flight was keyed by entry and allowlists only, so a request that read the newer stub joined the older inspection and inherited its outcome without the newer config ever being inspected. `reinspectServer` now reads the entry before joining and keys the flight by the stub's `updatedAt`. A flight whose compare-and-set loses to a newer failed stub joins the inspection of that stub instead of reporting the server unreachable, and a store that leaves the inspected stub in place rejects rather than joining its own flight. * 🪢 fix: Settle a Reinspection Against the Entry Stored After It A reinspection that did not replace its stub only looked at storage when its compare-and-set lost. Two paths reached a stale outcome: - An inspection that failed rejected as unreachable even when another request had already stored a recovery, for example after capturing the stub before a slow allowlist resolution. Allowlists now resolve before the entry is read, and a failed inspection settles against the stored entry the same way a lost compare-and-set does. - Returning a recovery that another replica stored left this replica's process-local `getAllServerConfigs` memo serving the stub until its TTL. Settling on a stored recovery now invalidates this replica's registry read caches, as the writing path does. * 🪢 fix: Decide Reinspection Outcomes From Current Shared State The Redis aggregate-key store answers `get` from a process-local snapshot that can lag another replica's write by its TTL. A replica whose inspection failed read that snapshot while settling, still saw the stub, and reported the server unreachable although Redis held another replica's recovery. The same stale read also sent a replica into a needless inspection of a server another replica had already recovered. - Add an optional `getCurrent` to the repository interface, implemented by the aggregate-key store as a read past its snapshot, and use it for the reinspection's first read and for settling. - Inspect a newer stub within the flight that found it instead of joining another flight, so no flight ever waits on a second one. * 🪢 fix: Join the Flight Already Inspecting a Newer Stub A flight that lost to a newer failed stub inspected that stub itself even when another request was already inspecting it. If that redundant inspection failed transiently, the flight settled against the unchanged newer stub and rejected its callers while the other flight went on to store the recovery. A flight now joins the flight for the newer stub, starting it when none is pending, but only when the stub is strictly newer by `updatedAt`, so every wait points forward and flights stay acyclic. A replacement that is not newer, which only clock skew can produce, is still inspected within the flight. * 🪢 fix: Track Reinspection Waits Instead of Ordering Them by Clock Joining only flights for strictly newer stubs kept waits acyclic by relying on `updatedAt` order, which replicas with skewed clocks do not guarantee. A replacement stored with an older timestamp was inspected redundantly instead of joining the flight already inspecting it, and a transient failure of that redundant inspection rejected callers the other flight went on to recover. Settling flights now record the flight they wait on and join any in-flight reinspection for the stub they found, unless that flight already waits on them directly or through others. Only a join that would close a cycle falls back to inspecting within the settling flight. * 🪢 refactor: Resolve the Reinitialize Config in packages/api `reinitMCPServer` still decided in CJS what a failed-inspection config recovers to and built the `unreachable` response itself. That decision now lives in `resolveMCPReinitializeConfig` in packages/api, which takes the registry as a dependency and returns either the config to connect with or the typed result that ends the reinitialization. The CJS service only passes request data in and returns that result.
…covery (#15865) * 📇 fix: Stop an MCP Credential Refresh From Fencing Its Own Catalog Recovery Passive catalog recovery captures the tool-cache generation before discovering a cold server and discards the catalog when the stored generation has moved by the time discovery returns. A silent OAuth refresh inside that same discovery publishes a rotation of its own and then clears this replica's recovery state, so the tracker drops the flight as superseded and the generation check rejects it: the server is listed with no tools, and a concurrent catalog request starts a second discovery instead of joining the first. The authorization fence now reaches discovery through recovery's own dependencies, which report each generation a discovery's refresh publishes. The tracker adopts that generation while the flight is still running, so the flight stays current and its outcome carries the generation it produced; a rotation or clear from any other writer afterwards still supersedes it, and a refresh that completes after its discovery settled no longer touches that settled state. * 🪢 fix: Keep a Recovery Flight Joinable Through Its Own Credential Publication A refresh's authorization fence writes the new shared generation, clears local recovery state, and only then reports what it wrote. A catalog request that read the new generation during that span found the flight still keyed to the old generation, or its slot vacated by the clear, and replaced it with a second discovery; the original flight was then discarded. Discovery now runs its publication through a tracker bracket. While a flight's own publication is in progress the tracker holds any clear for that server and lets requests for the same config join the flight; when it ends, the flight adopts the generation it wrote, or applies the held clear if the publication failed or the flight settled first. The flight is registered before its discovery starts, so a publication can never precede its own registration. * ⏳ fix: Bound and Count What a Recovery Flight Holds During Its Own Publication A recovery flight held every clear for its key and admitted any request for its server for as long as its own credential publication stayed open. A clear from another writer collapsed into the fence's own clear and was dropped on adoption, leaving a backoff or reauthorization decision in place after a teardown; a publication stalled in an unbounded storage write kept every later request joined to the stalled flight; and a joiner that had already read a foreign generation could keep tools under an older adopted generation when the final generation read failed. An open publication now holds clears and admits joins only while its flight is in flight and for at most one discovery budget, after which requests take over as before. The fence clears local recovery once when it completes, so any other held clear still applies after the flight adopts its generation, and every held clear applies when the publication fails or the flight settles first. An outcome under a generation the request never observed stands only when the final shared generation read confirms it. * 🧬 fix: Correlate Catalog Recovery Clears With the Generation Their Fence Wrote Local catalog-recovery clears carried only a user and server, so the tracker could not tell a flight's own credential publication from any other writer. It counted clears and bounded its hold with a time window instead, which left a request that joined before the window expired waiting on a stalled publication, and let a late fence clear delete the flight that took over after the window. The authorization fence now hands local recovery the generation it wrote, and a clear spares state already recorded under that generation; a clear without one still clears. A flight holds the clears it receives while its own publication is open, adopts the generation that publication wrote, and then applies each held clear against it, so its own clear is spared while a teardown or another publication's clear still applies. Retry-intent cleanup inside the fence is bounded by the same per-attempt timeout as the generation write, since a leftover intent only advances the generation again, so the publication itself can no longer hold its joiners open and no time window is needed. * 🔏 fix: Confirm an Adopted Catalog Generation for Requests That Read None A request whose pre-discovery generation reads both failed observed no generation, so the confirmation rule exempted it. When that request's own discovery refreshed credentials, adopted the generation its publication wrote, and another replica rotated past it before a final read that also failed, the catalog was served without ever being confirmed against shared state. A flight's own result now records that its generation came from its own publication, and that mark is never copied into retained state. An outcome under a generation the request did not read needs final confirmation whenever the request read a different generation or the generation was adopted; an outcome whose generation the tracker merely carried over for a request that read nothing keeps the existing fail-open behavior, so a cache outage still preserves retained backoff and reauthorization state. * 🧭 fix: Judge In-Flight Catalog Recovery by the Shared Generation, Not a Tagged Clear A publication's local clear carries the generation it wrote, and the tracker deleted any state recorded under a different one. Generations carry no order, so a refresh that outlived its own flight could clear with an older generation and delete a flight that had started under a newer rotation from another replica, discarding a catalog that was still current. A tagged clear now only signals that the shared generation advanced. A flight still in flight under a known generation, or publishing its own credential change, is left to the requests awaiting it, which already confirm their outcome against the shared generation exactly as they do for a rotation from another replica. Retained state and flights with no known generation are cleared unless recorded under the clear's generation, and a clear without a generation still clears unconditionally. The tracker no longer holds clears during a publication: the refresh's own clear spares its flight directly, and the flight adopts the generation it wrote when the publication completes. The shared-discovery tests compared catalog maps through expect.objectContaining, which does not compare Map contents; they now assert each request's catalog with toEqual. * 🧾 fix: Confirm a Catalog Flight That a Clear for Another Generation Spared A tagged clear spares a flight still in flight, because an unordered generation cannot say whether that flight predates the rotation. The spare was not recorded, so when the request's final shared-generation read timed out, a flight that had observed its own generation kept a catalog the clear may have superseded: an older recovery whose refresh published after a replacement flight started could leave that replacement serving tools no read ever confirmed. A spared flight now records the generation each such clear carried and settles against its final generation. A clear for its own generation, including its own refresh's, changes nothing; any other marks the flight's result unconfirmed, keeps it out of retained state, and requires every request awaiting it to confirm the generation against the shared read before using its tools. A spared flight that finishes with no known generation is discarded, as the clear would have done. A tagged clear no longer spares retained state recorded under its generation.
* feat: add run-scoped file sharing for subagents Delegate current-message files through authorized execution catalogs and publish selected immutable code-output versions through the existing file pipeline. Requires the companion agents SDK context adapter before the dependency can be released and pinned. * fix: preserve child tool artifacts and verify file-sharing lifecycle Keep citations and other non-sandbox artifacts on their existing delivery path while retaining private code outputs and inspected images until publication. Cover native and text delivery, nested and sibling authorization, concurrent runs, cancellation cleanup, and checkpoint recovery through the browser harness. * test: align file-sharing fixtures and integration cleanup * test: repair LibreOffice conversion fixtures * test: verify deployment skill bytes across checkout formats * fix: align subagent file budgets and refresh boundaries * chore: update agents SDK to 3.8.6 * refactor: move code output processing into API workspace * fix: enforce shared file budgets for parents and teams * fix: declare the shared file encoder interface
* ✨ feat: Add Upstream Model Error Observability * ♻️ refactor: Move Terminal Error Policy to TypeScript * 🛡️ fix: Preserve Agent Cancellation Semantics
* ✨ feat: Surface Upstream Model Errors * 🧪 test: Cover Upstream Error Rendering * 🛡️ fix: Harden Upstream Error Fallbacks * 🛡️ fix: Defer Upstream Error Fallbacks
* feat: lock code environments per conversation * style: Sort code environment imports * fix: Harden code environment decisions * fix: Preserve code environment rollout compatibility * fix: Harden code environment boundaries * refactor: Centralize code environment policy * test: Restore partial disconnect policy mock * fix: Close code environment edge paths * fix: Preserve implicit stateful code routes * fix: Stabilize implicit code routing * test: Align remote workspace contract * chore: Sort code route import
…d Tokens (#15871) * 🪂 fix: Release Stalled MCP Catalog Recovery Without Dropping Refreshed Tokens Passive catalog recovery could hang forever when discovery hit an expired OAuth access token and the refreshed tokens' authorization-fence write stalled. The shared flight, every request coalesced onto it, and their process-wide catalog slots waited with it until capacity was exhausted. - Stop discovery's OAuth token wait at its deadline or abort without cancelling the shared token flow, so redeemed tokens still persist. - Stop recovery waiting on a discovery once its budget plus `mcpSettings.catalogRecovery.discoverySettleGraceMs` (default 10s) has passed, releasing the flight, its joiners, and their slots with backoff. - Add `waitUntilDeadline` for waits that must end without cancelling work. * 🪂 test: Rediscover Once a Released Flight's Refresh Publishes * 🪂 fix: Hold a Server's Recovery Until Its Released Discovery Settles * 🪂 fix: Scope Released-Discovery Holds to Observed State Under a Process-Wide Limit * 🪂 fix: Count Token Work a Discovery Hands Off at Its Deadline
* 🔎 fix: Enforce WEB_SEARCH Role Permission on Agents Squash of #15822 onto dev, preserving authorship: - 🔎 fix: Enforce WEB_SEARCH Role Permission on Agents - 🧪 test: Cover the Runtime Tool Loader's Web Search Gate - 🔒 fix: Gate Web Search on the OpenAI-Compatible Embedder Surface - 🩹 fix: Authorize Web Search by the Execution Context's User - 📝 docs: Note the Grant Memo's One-Subject-Per-Request Assumption - 🧹 fix: Deny Web Search Before an Endpoint Default Can Grant It - 🙈 fix: Withhold the Web Search Model Parameter From a Denied Role - 🧱 fix: Strip Provider-Native Search an Endpoint Re-Enables After Denial - 🔌 fix: Close the OpenRouter Plugin Path and Stop Pruning Hidden Parameters * 🧭 fix: Gate Provider-Native Web Search on the Built Config The web search role gate ran before the provider builder and only when the agent's stored `web_search` was not `false`, so an endpoint's `addParams` re-enabled native search for a denied role, and every OpenAI-compatible and Responses request read the role whether or not search was in play. The gate now reads the builder's output: when it carries a native search tool or OpenRouter's `web` plugin, the `WEB_SEARCH` grant is resolved and the search stripped on denial. The routes that reach `initializeAgent` with `runtime` and no `req` pass a lazy `resolveWebSearchGrant` that joins their request-memoized grants, forwarded through agent discovery to handoffs and subagents. The embedder surface wires it whenever `getRoleByName` is supplied, independent of `appConfig`. --------- Co-authored-by: TomasPalsson <tomas2212042710@gmail.com>
* 🔭 feat: Conversation Trace Viewer
Adds a trace view for a user's own conversations: a summary, a pinned
waterfall overview with interval focus, a hierarchical record ledger and
a record inspector, opened from the chat header (or the mobile overflow
menu) and read through a provider-neutral seam with Langfuse as the
first reader.
- interface.traceViewer (enabled, showInputOutput, maxRecords,
maxContentLength, requestsPerMinute), off by default
- /api/traces/:conversationId availability, records and record routes,
re-proving conversation ownership and narrowing the Langfuse session
to traces the user's own sampled responses produced
- Langfuse reader over the v2 Observations API with cursor paging, time
bounds, typed failures and server-side input/output withholding
- forks no longer copy the source messages' trace sampling record
* 🛂 fix: Address Trace Viewer Review Round 1
- interface.traceViewer.requestTimeoutMs bounds each Langfuse round trip
- availability is one existence query across every sampled response
instead of the oldest one, run alongside the ownership check
- tree items report their position among visible siblings
- point-in-time Langfuse events read as completed, not running
- reads prefer the project holding the most responses when a tenant
connection arrived mid-conversation
* 🎯 fix: Leave Focus With a Navigation That Closes the Trace
* 🧭 fix: Keep Trace Reads on One Project and Cancel Abandoned Reads
- page cursors and record details stay on the project that served the
first page; pages report an opaque sourceId
- Open in Langfuse only appears when the linked project served every
loaded page (the session link now reports its destinationId)
- client and server abort Langfuse reads when the viewer closes or the
client disconnects
- reads never wait on the central project lookup, which the request
timeout does not govern
- the cost total is withheld when any model call with usage is unpriced
* 🕰️ fix: Refresh Trace Pages After a Turn, Honor Clock Settings, Show Refresh Failures
- a settled run invalidates the cached trace pages, so reopening shows
the new turn instead of pages fresh for another 30 seconds
- trace durations and clock times format in the app language and the
user's 12/24-hour setting, as message timestamps do
- a failed refresh with every page loaded surfaces an alert and a retry
* 🔐 fix: Only Server-Authored Responses Grant Trace Access
- the message-create route and imports strip langfuseSampled and
langfuseDestinationIds, and trace reference queries ignore client-authored
rows (isUserSubmitted) and user turns, so a forged row cannot claim a trace
- the Langfuse reader receives its destination resolver and HTTP client
from the route instead of defaulting to process-global ones
- availability asks the client to check again while central's project id
is still resolving, instead of caching a false answer
- listRecords tries the next readable project when legacy responses left
the first one empty
- a settled run also invalidates cached record details
- running is read from a record's status, not a missing end time
* 🏢 fix: Scope Trace Reads to the Tenant and Reach Every Project Holding a Turn
- ownership and both trace reference queries carry the requester's tenant
(tenantless is its own scope), so a duplicate user and conversation id in
another tenant cannot grant or supply trace references
- a read visits, page by page, every project that could hold a turn no
earlier project provably holds, so legacy turns in another project stay
reachable after a newer destination ranks first
- the trace read limiter takes its store from the route
- transient availability failures retry instead of hiding the control
* 🛟 fix: Fail Over Between Projects, Refresh One Page, Search Visible Labels
- a fresh read fails over to another project holding the same turns when the
preferred one fails; the cursor carries the failed sources so later pages
keep the rebuilt plan, and detail reads fail over the same way
- destinations for one project collapse into one source
- the trace read limiter keys each user within their tenant
- refresh and a settled run keep only the newest page before refetching, so
older pages are not replayed through the limiter; retry repeats the read
that failed
- search matches the localized kind and status labels a row shows
* 🧭 fix: Read the Newest Turn First, Report Failures Behind Empty Projects, Fit Trimmed Views
Rank trace sources by the newest sampled response they could hold before coverage and preference, so a conversation that moved projects shows its latest turn first.
Throw a list failure when a failed project could hold a turn that no answering project provably holds, instead of rendering an empty trace. Detail reads keep probing projects after an empty answer, skip ones whose turns an answering project provably holds, and surface a failure that could hide the record.
Refresh and settled runs trim the cache to the newest page; the viewer now fits a focused interval to the remaining records or clears it, and drops a selection whose record left with the older pages.
* 🧵 fix: Page Traces by Turn, Include Failed Runs, Keep the Active Row Mounted
List reads now page by turn segments instead of whole projects. A segment is the newest unread turn and up to 49 older consecutive turns one project could hold. The request filters Langfuse to those turns' run and title traces, so pages stay newest first when a conversation's responses alternate between projects. Legacy turns without destination ids are asked of the other projects in the segment. A project that failed for newer turns is retried when it alone holds older turns. A page whose rows all fail the observation schema is an upstream failure, not an empty trace.
A failed turn's error row now stores the sampling fields of the run that failed, plus langfuseRunId when that run was created, so the trace viewer and feedback scores follow the run's trace. langfuseRunId is stripped from client message writes and imports, dropped by forks, and hidden from client message projections.
The ledger keeps its active tree item mounted while virtualization scrolls it out of the window, so aria-activedescendant always names a rendered row.
* 🕰️ fix: Merge Multi-Project Trace Pages by Start Time, Keep Literal Content
A segment that reads more than one project now runs every read before returning. A read that fills the record budget cuts the page at the oldest start time all reads loaded, and the continuation carries that start time as an inclusive bound instead of a Langfuse cursor. Every project continues from the same point, so a newer turn held by another project is never left behind an older one, and failover works on any page. A page drawn from several projects no longer names a single source. A continuation that would repeat its own bound fails instead of looping.
Input and output that are literally the strings {} or null are kept; only an empty object or JSON null value is dropped.
* 🔎 fix: Read Each Trace Segment From a Project Shown to Hold It
Recorded destination ids say where a trace was eligible to go, not where it landed. A list read now looks up the turns more than one readable project could hold before choosing where to read them. Each project, preferred first, is asked which of those traces it holds a root observation for, and each turn is read from the preferred project that holds it, or from the preferred eligible one while a run's root is not exported yet. Turns only one project could hold need no lookup.
Segments are consecutive turns from one project, paged by that project's own Langfuse cursor. Records that share a start time page correctly, and the page budget bounds a single read. A failed project is dropped for the request. A turn only it could hold raises its failure at the turn's own place in page order, instead of leaving an empty or partial trace. Cursors carry the segment's trace-set hash and start over when the segment changed or Langfuse rejects them. Detail reads ask every project until the record turns up.
The trace summary withholds its cost total when any generation has no cost, with or without usage.
* 🧷 fix: Evidence per Trace, Continuations Without Replays, Stable Turn Order
A turn's run and title run are exported separately, so each trace id is looked up and read from a project shown to hold it. A trace no project shows (a title run the turn never had, or a run whose root is not exported yet) joins an adjacent segment whose project could hold it. Only a run that a failed project alone could hold raises that failure.
A continuation now stays on its project. Its failure is returned as is, and a cursor the segment no longer matches, or one Langfuse rejects, returns invalid_request instead of replaying the segment's first page. The viewer answers invalid_request on an older page by reloading from the newest page, with a message saying the trace changed.
Sampled responses sort by createdAt then _id, so responses saved in the same millisecond keep one order across page requests.
* 🪪 fix: Scope Trace Reads to the Exporting User, Order Split Title Runs
Trace ids derive from message ids, which a request can influence, so a colliding id must not return someone else's trace. Every Langfuse list, probe and detail read now also filters by userId: the requester's internal id, plus the configured langfuse.trace.userIdField value when set. The server stamps that id on the trace at export. The handler passes the requesting user into the trace query.
When a turn's run and title run sit in different projects, the probes' root start times decide which is read first, so a title generated after the response is not left behind Load older. Cursors name the trace a segment starts at within its turn.
The failed-turn trace lookup moves into packages/api as getFailedTurnTraceFields. The agent controller only passes the error row id, the run id and whether the run was created.
* 🔒 fix: Authorize Trace Reads by the Internal User Id Only
A configured langfuse.trace.userIdField can export a non-unique value (name, username) or one that repeats across tenants (email, provider ids) as the trace userId, so it cannot prove who a trace belongs to. Reads now require the internal user id alone. A deployment that exports another field reports no trace and logs why once, instead of trusting that value.
A failed read of a title run that a probe found in that project now raises its failure instead of dropping the title. The page contract now states that pages are ordered by turn and a turn's records may continue on the next page.
* 📚 fix: Page Sampled Responses With the Trace, Share Ingress Stripping
List reads load sampled response references a segment at a time: up to 50 ending at the cursor's turn, plus one older response that names the next page. Pages of a long conversation no longer rematerialize every sampled response. getConversationTraceRefs gains through and limit, anchored by createdAt and _id.
The message-create route and conversation imports strip trace sampling fields through withoutTraceRefs instead of their own field lists. A userIdField the exporter ignores counts as the internal id. A root lookup whose cursor repeats fails that project instead of paging forever.
In the viewer, Refresh also rereads the open record's detail, and closing the inspector after a filter removed every row returns focus to the search field.
* ⏱️ fix: One Round Trip per Trace Reference Read, Detail Reads Scoped to Their Turn
Cursors carry the anchor response's order key (creation time, then _id). The bounded sampled-response read then runs alongside the first-message lookup, with no anchor lookup ahead of it. A key that no longer names its response returns nothing, and the reader answers the cursor as changed.
Detail reads take the turn the list attributed the record to (?message=). They load only that sampled response instead of the conversation's whole history, and ask Langfuse within that turn's traces. The route rejects a detail read without it, and the client sends it from the listed record.
A list read whose Langfuse cursor repeats now fails instead of handing the next page the same records.
* 🧭 fix: Reject Out-of-range Trace Cursors, Open Any Listed Turn, Mark Copies Unsampled
A page cursor whose base-36 time parses to a number outside the Date range reached Mongoose as an Invalid Date and surfaced as a 500 the viewer retried; it now resumes nothing, which the reader reports as invalid_request so the viewer reloads its newest page.
The detail read capped the message id at the Langfuse record-id bound while the app stores response ids at any length, so a listed turn could refuse to open; message ids only name a stored row and are accepted at any length. Ids that reach a Langfuse filter keep their caps.
Copied and client-authored rows now carry langfuseSampled: false instead of no record, so a feedback score on a fork, import or posted message never recomputes sampling for a trace that was never made.
* 🧪 test: Type a Recordless Row for the Unsampled-copy Spec
An object literal with none of the trace fields fails TypeScript's weak-type check against the helper's all-optional constraint; the spec now passes a typed row that declares the optional field, as every stored message type does.
* 🚰 fix: Release Failed Langfuse Response Bodies
A non-success Langfuse response was turned into a TraceReadError without consuming or cancelling its body, and Node's fetch returns a connection to its pool only once the body is released, so a sustained upstream failure left sockets checked out until garbage collection. The body is now cancelled before the status becomes an error.
* fix: Preserve Disabled Workspace Bindings * fix: Resolve Restored Workspace State
* 🪪 fix: Adopt MCP OAuth Callback Tokens Without Re-Prompting or Losing the Lease A chat request loading OAuth tokens joins the `mcp_get_tokens` flow a catalog discovery left behind and polls it. The OAuth callback deletes that entry when it stores the new tokens, so the poll ends with "Flow state not found", the factory reports the tokens missing, drops the `mcp_oauth` flow the callback just completed, and starts a second interactive flow. Once tokens arrive, the connection leases the publication generation it captured before the wait; the callback's publication moved it, so the connection is torn down as it connects and the agent fails with AGENT_EXPECTED_MCP_TOOLS_UNAVAILABLE. `createFlowWithHandler` now serves a completed result at once, monitors only pending attempts and retained failures, and replaces a failure nobody retains; a waiter whose flow disappears gets `FlowStateNotFoundError`. The factory treats that as a credential change: it re-reads storage once, after letting the connection build re-capture its generation. The authorization transaction and a silent refresh attach the generation they publish to the tokens they release, and a build adopts that generation for credentials obtained after its capture. * 🪪 fix: Claim Flow Attempts Atomically and Lease Under the Stored Generation A connection now adopts a carried publication generation only while it is the generation currently stored, instead of comparing the token's exchange time to its capture, so a callback that exchanged before the build began and published after it is still adopted and a retired generation still leaves the fence in place. Installing a flow attempt is an atomic compare-and-set against the attempt the caller observed, on Redis and on the default in-memory store, so concurrent replacements of a failed attempt, or concurrent creations of an absent one, run one handler while the rest join it. An aborted caller is rejected before a cached result is served, and a joiner that aborts leaves the shared flow to its owner and the other waiters. * 🪪 fix: Return Generation-Tagged Callback Tokens and Route Flow Claims Through the Cluster Executor The authorization transaction now returns the same generation-tagged tokens it releases to the OAuth completion and the pending token waiters, so the callback route's tool-flow completion carries the published generation to the connection waiting on it. The Action OAuth and index-sync flow managers pass the cluster-aware Redis script executor, so their flow claims and guarded settlements route through slot-master selection and READONLY recovery instead of the raw client.
* 🚀 chore: prepare v0.8.8 * 🚀 chore: prepare v0.8.8-rc3 * docs: mention Agent workspace defaults
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )