Skip to content

[pull] main from danny-avila:main - #270

Merged
pull[bot] merged 27 commits into
innFactory:mainfrom
LibreChat-AI:main
Sep 21, 2026
Merged

pull[bot] merged 27 commits into
innFactory:mainfrom
LibreChat-AI:main

Conversation

@pull

@pull pull Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

lia-by-librechat Bot and others added 27 commits September 20, 2026 05:24
* 🐣 feat: Invoke Newly Authored Skills In the Same Run

* test: Pin the SDK skill tool text the authoring variant rewrites

* fix: Replace a stale skill definition instead of suppressing the duplicate

* fix: Pin Skill Invocation to the Doc Authored In the Run

---------

Co-authored-by: Lia <lia@librechat.ai>
Co-authored-by: Lia <lia@librechat.ai>
* 🥖 fix: Refresh MCP Credentials That Arrive Already Expired

storeTokens replaced a stated access-token lifetime with the 365-day default whenever it had already elapsed, so a credential that arrived expired was recorded as valid for a year. getTokens decides to refresh from that record, so it never did: recovery waited for the resource server to reject a request, and on connection establishment a second rejection of the refreshed tokens starts interactive OAuth instead.

A lifetime stated by the response is now persisted as it stands, including when it has elapsed, so the next read refreshes it silently. A JWT exp keeps its existing leniency because it comes from the provider clock, where skew can make a live credential look expired.

* style: Sort the jsonwebtoken Import in the Expiry Suite

Package imports are ordered by line length, so `keyv` precedes `jsonwebtoken`. Fixes the Static checks import-order gate.

* fix: Renew Expired OAuth Results Before Connection Readiness

Preserve finite zero lifetimes at exchange and storage. Both callback consumers re-read expired flow results through coordinated refresh after crossing and releasing the callback persistence fence. Cover the real SDK callback transaction, concurrent token waiter, zero-expiry updates, retryable failures, bounded recovery, cancellation, and the unchanged fast path.

---------

Co-authored-by: Lia <lia@librechat.ai>
* 🥢 fix: Exclude Prompt Inputs From Terminal Events

Terminal events carried prompt-building inputs into every place a FINAL is
kept: the runtime cache, the durable Redis job hash, Pub/Sub publication and
late/cross-replica replay. Neither terminal producer projected the payload, so
a 706 KiB attachment-heavy FINAL was stored and published at full size.

sanitizeMessageForTransmit removed fileContext and files[].text from
requestMessage at the controller, but never image_urls, and responseMessage
was not sanitized at all.

projectTerminalEvent is one shared, schema-aware projection applied at both
GenerationJobManager.emitDone() and publishTerminalClaim(), and on read where
a stored record may predate it. It excludes fileContext, image_urls and
embedded file bodies from requestMessage, responseMessage and every
runMessages entry, while preserving message content, attachment references and
display metadata, and the terminal/reconciliation protocol fields.

publishTerminalClaim compared publicationEvent !== finalEvent to decide
persistenceFailed. Projection allocates a new object, so that identity check
now compares against the projected intendedEvent; a successfully persisted
FINAL would otherwise have been reported as a persistence failure.

* fix: Preserve Output Attachment Text and Project Live Terminal Frames

Addresses three review findings at 9e5f6f8.

Output attachment text was excluded as if it were a prompt input. `files`
carries the user's uploads, whose `text` is the body extracted for the prompt,
but `attachments` holds resolved artifactPromises that BaseClient assigns onto
responseMessage, so its `text` is model-generated output rendered inline by
TextAttachment from `file.text ?? ''`. Excluding it was also unrecoverable:
useAttachmentPreviewSync polls the preview endpoint only while status is
'pending', and an absent status reads as 'ready', so a late or cross-replica
subscriber whose only source is the stored FINAL rendered a blank preview.
Each collection now carries its own exclusion set; storage bookkeeping (_id,
__v) is still excluded from both.

A live Pub/Sub FINAL from an older replica bypassed projection. The store-read
paths were covered, but during a rolling deploy an old generation owner
publishes an unprojected frame that a new subscriber cached on its runtime and
forwarded to the browser. queueDone is the one choke point every terminal
delivery to a subscriber passes through, so the projection now runs there.

The declared return type was the input type T, so the exported projected
contract was never enforced. FinalMessageFields carries an index signature,
which makes Omit alone a no-op, so ProjectedMessageFields now maps the
transient fields to `?: never` and overloads return the projected contract for
a FinalEvent, the exact type for a non-terminal event, and ServerSentEvent for
the union every manager call site passes.

* fix: Spell Out Transient File Fields for isolatedDeclarations

tsdown builds packages/api with --isolatedDeclarations, which rejects a spread
element in an inferred array type (TS9018), so composing TRANSIENT_FILE_FIELDS
from a shared bookkeeping tuple broke Build packages and every lane that
depends on a built package. tsc --noEmit does not use that flag and passed,
which is why this reached CI.

Both tuples are now spelled out, with a comment recording the constraint.

---------

Co-authored-by: Lia <lia@librechat.ai>
* fix: give same-batch code calls the files a skill just loaded

A batch that carries a skill call next to execute_code/bash ran both
against the code-session snapshot the graph took when it planned the
batch, so the sandbox had no /mnt/data/skills/<name>/ for that turn.
Code calls in a mixed batch now wait on the skill calls and merge their
uploaded refs into their own context before dispatch.

* test: pin the spec's own file ref kind so CodeEnvFile stays assignable

---------

Co-authored-by: Lia <lia@librechat.ai>
* refactor(redis): import script cache proposal from #16069

Squashed copy of PR #16069 at 7e99419, rebased onto current dev without changing the original branch.

Co-authored-by: Marco Beretta <81851188+berry-13@users.noreply.github.com>

* refactor(redis): cache independent claims without queueing streams

---------

Co-authored-by: Lia <lia@librechat.ai>
Co-authored-by: Marco Beretta <81851188+berry-13@users.noreply.github.com>
…e From Its End (#16113)

* ✂️ fix: Name Tool Rounds From the Trace, and Match a Cut Response From Its End

* 🧵 fix: Join a Cut Model Call's Round Only When One Is Waiting

* 🏷️ fix: Read Labels From the End Only for the Response the Limit Cut

* 🪪 fix: Say a Tool Round Came From the Conversation Only When It Did

* 🧾 fix: Carry Where a Tool Round's Calls Came From
Classify HTTP 408, 429, and 5xx before legacy body-message classifiers can invalidate client registration or start consent. Preserve typed temporary failures across refresh storage and verify retry with the existing grant against the real SDK test provider. Keep permanent grant errors and existing UI retry behavior intact.

Co-authored-by: Lia <lia@librechat.ai>
…16120)

* 💵 feat: Show Cost per Record, Step and Response in the Trace Ledger

* 💵 fix: Cost for Whole Responses Only, Hidden Records Included, and Said Aloud

* 💵 fix: Give the Oldest Loaded Response No Cost While Older Records Remain
* 🥊 test: Prove MCP Refresh Coordination Across Replica Processes

Fork real replica processes with production Mongo, Redis, encryption and OAuth adapters, and assert one redemption per rotation, peer adoption, and recovery after the authorizing replica is killed. A negative control with coordination disabled observes two redemptions, so a green coordinated run cannot pass on timing alone.

Adds a refreshGate hook to the shared OAuth test provider so concurrent refreshes can be held open. No production code changes.

* test: Establish Replica Overlap Instead of Assuming It

The coordinated assertion counted provider redemptions after a fixed delay, so a peer that arrived once the winner had already persisted would read a fresh credential, never contend, and still leave one redemption: the test would have passed while proving nothing. Both replicas now report entering the refresh path on a credential they observed expired, and getTokens own onRefreshSuccess/onTokensAdopted hooks report which side each took, so the pair (one redeemed, one adopted) is asserted rather than inferred from timing.

A worker that never signalled ready was also left running with live Mongo and Redis connections, which would outlive the suite and disturb the rest of the integration lane; boot failures now reap the child and report the workers own initialization error.

---------

Co-authored-by: Lia <lia@librechat.ai>
* 🧹 fix: Build the Legacy Meili Cleanup Index MongoDB Rejected

`meili_excluded_legacy_cleanup_v3` declared `_meiliCleanupVersion: { $exists: false }`
in its `partialFilterExpression`. MongoDB rewrites that to `$not`, which no partial
index accepts, so the build failed on every startup for Message and Conversation and
the index never existed.

The missing-version condition moves into the key, where a missing field indexes as
null and the legacy cleanup query reaches it on the same scan. The filter keeps the
two conditions a partial index can express, and the name moves to `_v4` so no
deployment that somehow holds the old spec hits an options conflict.

* 🎯 fix: Bound the Legacy Cleanup Index to Documents Awaiting Cleanup

The buildable filter admitted every excluded document carrying `_meiliIndex: false`,
which the schema writes on every new subagent message and conversation. Those
documents already carry the current `_meiliCleanupVersion`, so the cleanup query never
selects them and their entries never leave the index.

The unstamped state is now expressed as a null equality, which a partial filter
accepts and which matches an absent or null version. The legacy cleanup branch asks
the same way, so the planner can still reach the index, and a document leaves it as
soon as cleanup stamps the version.
…16037)

* ⚓ fix: Keep a Resumed Compaction Anchored to the Turn It Summarizes

A compaction submits no user turn: its user-message slot names the leaf it
summarizes up to. The resume paths adopted that identity-only projection as a
ROW, rewriting the answer being summarized into an empty, parentless user
message, which buildTree files as a phantom root: the thread folded and the
pane could return on another branch. Treat the slot as an anchor everywhere
rows are written, and let the anchor name the branch to restore.

* Recognize a compaction anchored on a user leaf

Compact runs on whatever leaf the branch ends with and canCompact does not restrict its author, so the anchor is a user message as often as an answer. Detecting it by not-user-created recognized only the assistant-leaf kind: the other was rebuilt as an ordinary turn, so the sync path merged the identity-only projection over the stored row and blanked the prompt still on screen, and a failed abort appended a second row under its id. Judge the anchor by the projection shape instead, and carry the resolved answer on the resumed submission so a flagless re-attach agrees.

---------

Co-authored-by: Lia <lia@librechat.ai>
* fix: bind preserved MCP API keys to server configuration

* fix: surface MCP API key rebinding errors

* test: use typed MCP OAuth re-entry error
* test: add deployed instance Playwright smoke profile

* test: normalize deployed smoke account state

* test: harden deployed smoke lifecycle

* test: secure deployed smoke cleanup
* 🫧 feat: Hold a Live Run's Activity at One Row

Fold the span a run is still writing into a single collapsed row from its first tool call, with a header that bubbles up the newest activity (streamed intent, generic running text, filled batch label, or the thought streaming after them) on a 500ms leading+trailing throttle. The settled rendering is unchanged.

* 🫧 fix: Name Every Live Line and Take the Fold Preference From the Host

Live header resolves failed/cancelled through resolveToolCallPhase, names trailing commentary, and only folds agents-shaped calls it can name. ContentParts takes foldLiveActivity from its host instead of reading Recoil; subagent panels opt out.

* 🫧 fix: Carry Background Stops and Nested Sign-Ins Into the Live Fold

* 🫧 feat: Preview Tail Reasoning One Sentence at a Time in the Live Row

* 🫧 fix: Keep Handoffs, Detached Failures and Status Announcements Out of the Fold's Blind Spots

* 🫧 fix: Keep the Live Line's Shimmer Off the Ticker's Animated Element

Found in Chromium: .shimmer declares animation/position/display, so on the same element it replaced the slide-out and the retired line never left. The sweep now rides an inner span, and the header button is block-level flex so a live row measures the same 28px as the settled row.

* 🧹 chore: Sort Imports in the Live Activity Helper

* 🫧 refactor: Decide a Live Line's Outcome With the Group Header's Resolver

The live header derived pass/fail on its own and re-learned, one review at a time, signals the cards already handle. getToolMeta moves out of ToolCallGroup into Content/outcome.ts and the live line reads it, which covers cancelled status attachments and memory-tool prose failures. A legacy Assistants call now ends a live span like a handoff, and the live disclosure is named by its current line alone (aria-labelledby) so the polite region's previous line stays out of its accessible name.

* 🫧 fix: Make the Live Row Agree With the Cards It Hides

Span-level outcome beside the newest line (an earlier call can fail while a later one runs), announcements with their own identity in one polite region that outlives the live header, the memory-error artifact and step-scoped attachment ownership in the shared resolver, one localized 'Failed: {{0}}' template for the card and the row, and a parity suite that compares the folded row with the REAL cards on unmocked outcome logic.

* 🫧 fix: Keep Live Subagents Unfolded and Give Reused Ids and CJK Sentences Their Own Lines

* 🫧 fix: Hold a Finished CJK Sentence and Honor Open Thinking in Live Folds

* 🫧 fix: Say "Running in Background" as the Card Does and Hold Reasoning Identity Across Whitespace

* fix: Harden live activity folding boundaries and lifecycle

* 🫧 feat: Surface a Span's Own Glyphs on Its Collapsed Header

A header stands for rows it hides, so it carries the most specific glyphs they show. The icon stack swaps the generic search glyph for the sites a web search read (from attachments, or the streamed results while live), and the settled phase header keeps the span's tool, MCP and site icons instead of a bare check; a failed phase keeps its warning glyph. Applied to the live row, the settled phase header and the tool group header. Source helpers move out of WebSearch into Content/sources.ts.

* fix: Preserve owned glyphs and outcomes across collapsed headers

---------

Co-authored-by: Lia <lia@librechat.ai>
* ⚡ perf: Split Streaming Markdown Blocks Incrementally

* test: cover streamed markdown block rendering

* fix: isolate streamed markdown splitter caches

* fix: satisfy markdown splitter static checks

* test: mock isolated markdown splitter

* test: isolate markdown splitter spies

* test: cover concurrent streamed markdown messages

* fix: preserve provisional streaming markdown boundaries

---------

Co-authored-by: Lia <lia@librechat.ai>
…16070)

* ⚡ perf: Exclude Agent Version History From Default `getAgent` Reads

getAgent now defaults its projection to { versions: 0 }; loadAgent/loadAddedAgent
and canAccessAgentFromBody use getAgentWithVersionCount so the version count stays
exact without transferring the unbounded versions array on every chat request.
Callers that need history (v1 update/revert handlers) request it explicitly.

* fix: tighten agent version projection types

* fix: type projected agent version counts

* fix: type default agent projection

* style: format agent projection overload

* fix: expose projected agent return type
* ⚡ perf: Preserve Message References in the Content Handler

* test: cover content handler message reconciliation

* fix: preserve thread metadata for incomplete content events

* fix: retain fallback response thread identity

* fix: select parent from response thread

* fix: retain untagged conversation messages

* fix: preserve history while resolving parent

* fix: refresh cached response metadata

* fix: honor synchronized parent metadata

* fix: scope parent fallback to response thread

* fix: avoid cross-thread parent fallback

* fix: keep content handler selection semantics
* ⚡ perf: Virtualize the Model Selector's Search Results

* fix: make model search accessibility metadata global

* test: stabilize model selector search scenarios

* test: assert model selector keyboard semantics

* test: simplify model selector keyboard assertion

* test: exercise keyboard pin navigation

* fix: preserve search result announcements

* test: follow active row pin through keyboard

* test: scope pin toggle assertion

* test: scope pin toggle to active row

* test: stabilize keyboard pin scenario

* test: isolate model selector favorites

* fix: preserve virtual search boundary navigation

* test: cover virtual search boundaries

* fix: own virtual boundary by logical position
…6143)

`useLazyHighlight` tokenized the same input twice on every mount that could
highlight immediately. The state initializer highlighted during render whenever
the grammars were already loaded, which holds for every card after the first in
a session, and then the effect highlighted the same input again because its
dedupe ref started empty and could not match on a fresh mount. Each card paid
two tokenizations, two hast-to-React conversions and an extra commit.

Card bodies also tokenized while collapsed. Every code, bash, read-file and
file-authoring card highlighted its content on mount, so opening a long agent
conversation highlighted code nobody had asked to see.

The tokens now carry the key they were produced from, so the effect recognizes
what the initializer already did, and each call site passes its content only
while its pane is open. ExecuteCode gains the raw fallback its three siblings
already had: with highlighting deferred to the pane opening, rendering only the
highlighted nodes would leave the pane empty until the grammars load.
…d Thoughts (#16145)

* 🫧 fix: Fold a Streaming Thought Into the Live Row Before Its First Tool Call

A thought at the live tail rendered as its own row with a multi-line peek until a tool call arrived, which then snapped it into the one-row fold. With a reasoning model that talks between calls that is a jump up and back down on every step. The span now folds from the thought's first character and previews it one sentence at a time, under a reasoning glyph until a tool gives it icons.

* 🫧 fix: Keep a Live Card's Key When Its First Tool Call Arrives

* 🫧 fix: Stop Unfolding the Whole Live Span for Every Code Call Without a Leading Intent

blocksLiveFold unfolded the entire span while a bash/code call with complete args, no output and no intent was running, so its card could show the sandbox-startup label. When the model writes intent after command, every code call qualifies: a 47-call run opened to its full height and shut again on each call. The row now reads the same sandbox signal for its newest call and the span stays one card. Also keeps tool-anchored card-key aliases across finalization so a reasoning-led card a reader opened stays open when the response settles.

* 🧹 chore: Sort Imports in ActivityPhaseGroup

* 🫧 fix: Preserve Reasoning Disclosures and Isolate Streaming Siblings

* 🧪 test: Keep Real Part Identity in Content Renderer Mocks

* 🫧 fix: Migrate Sandbox Startup Readers and Writer to Jotai

---------

Co-authored-by: Lia <lia@librechat.ai>
`style.css` pinned every `code` and `pre` element to
`Consolas, Söhne Mono, Monaco, Andale Mono, Ubuntu Mono, monospace !important`.
The repository ships none of those faces, so Windows rendered code in Consolas
and macOS in Monaco, which carries neither an italic nor a bold face for the
browser to use. `!important` also outranked the 21 `pre` and `code` elements
that ask for `font-mono` by class, so the self-hosted Roboto Mono the app
already bundles was never used for code anywhere.

Move the stack to `theme.fontFamily.mono`, where `sans` already lives, so
Tailwind's preflight styles the bare elements and the `font-mono` utility
carries the same value. The tail is ordered so the glyphs the bundled latin
subset omits keep Roboto Mono's advance width.

Co-authored-by: Lia <lia@librechat.ai>
@pull pull Bot locked and limited conversation to collaborators Sep 21, 2026
@pull pull Bot added the ⤵️ pull label Sep 21, 2026
@pull
pull Bot merged commit 0cc52cd into innFactory:main Sep 21, 2026
1 check passed
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants