Repository navigation
doltgres 1.4.0 - #314846
Merged
Merged
doltgres 1.4.0#314846
Conversation
chenrui333
approved these changes
Oct 2, 2026
Contributor
|
🤖 An automated task has requested bottles to be published to this PR. Caution Please do not push to this PR branch before the bottle commits have been pushed, as this results in a state that is difficult to recover from. If you need to resolve a merge conflict, please use a merge commit. Do not force-push to this PR branch. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Created by
brew bumpCreated with
brew bump-formula-pr.Details
release notes
Add
git-remote.max-history-commitsandgit-remote.reset-history-on-prunesettings so users can trade disk usage for faster incremental transfers. Support local/globaldolt config, environment overrides, and typed options for embedders.Defaults remain unchanged. A history limit of
0means unlimited; disabling prune-triggered resets allows obsolete data to remain reachable until the history limit is reached.Preserve a persistent Git fetch anchor during teardown so subsequent pulls can negotiate incremental transfers instead of re-downloading the database. Compare-and-swap protects the anchor from overlapping sessions while allowing private refs to be cleaned up.
With these optimizations,
dolt logis now anO(N)operation rather than anO(N^2).Optimizations:
dolt_branches,dolt_tags,dolt_log()once instead of for every commitCommitInfois passed as a pointerPotential improvements:
Before:
Fixes:
dolt logtakes really long time. dolthub/dolt#9315Newly persisted Git-backed tables and conjoin input reads bypassed NBS spooling, causing redundant Git reads and decompression. This change spools persisted tables from bytes already in memory and makes conjoin reuse existing source readers.
Adds regression coverage for chunk correctness, blob read counts, and spool cleanup. Full NBS tests and focused race tests pass.
Fixes #11916.
Upserts could rewrite rows without UPDATE privileges. Update GMS to require UPDATE authorization and add protocol regression coverage for denied and permitted upserts.
The new library keeps the interface - New, Lock, TryLock, LockWithTimeout, Unlock, Close, ErrLocked and ErrTimeout - so nothing here changes but the import line. The package is still named fslock, imported under that name from a path that no longer matches it.
These allow Doltgres to see transaction start, why a transaction ended (rollback, commit) and interact with savepoints as well.
Limits CI deferral and release comments to relevant label changes and draft/ready transitions, preventing duplicate notifications on commit pushes.
Two pre-existing bugs on the
ghostGen == nilpath, both only reachable outside the Dolt server (dbfactoryalways supplies a ghost store):HasManyreturnednilinstead of the absent set, hiding missing hashes. Regression test included.PersistGhostHasheshad its nil check inverted.When GC moves chunks from new to old gen, it publishes them into old gen and then, eventually, drops them from new gen. If the Reads check old gen first, and then fall back to new gen, they can miss a chunk which is actually retained --- the entire GC can run in between the return from the old gen Read and before the new gen Read.
Reading new gen first is the fix. If the chunk isn't in new gen and it does exist, it has guaranteed already been published to old gen by the time we ask for it.
An unlocked read samples |nbs.keeperFunc| under |nbs.mu| and then runs
against a table set with the lock released, so a read already in flight
when BeginGC installed its keeper could return chunks having taken no
read dependency on them, leaving the GC free to collect them. BeginGC now
sets |gcInstallPending| and waits for |outstandingReads| to drain before
installing, and every read path parks on that flag at the top of its
locked section. Parking there, rather than in |beginRead|, keeps a read
which spans the memtable and the table files from being split across the
barrier. A BeginGC whose context is cancelled during the drain clears the
flag and restores conjoin, so it leaves no reads parked and no conjoin
disabled.
|beginRead| now samples |nbs.keeperFunc| and |nbs.gcCycleCounter| itself
and registers every unlocked read, rather than each call site sampling
them separately and only reads which started during a GC being
registered. A read's keeper and cycle therefore come from the same
locked section which registers it, |nbs.outstandingReads| counts every
read running against a table set, and |endRead| is never nil. This is
the plumbing for making a keeper installation a barrier: the follow-up
drains those registrations before it installs a keeper. The costs are
that a read with no GC running reacquires |nbs.mu| once at completion,
where before it did not, and that EndGC now also waits out reads which
started before its GC began.
Now take
nbs.muwhen we broadcastgcCond, so that our waiter is definitely already parked if they need to see our wakeup.The NodeStore and the ValueStore cache have the structure where they are queried and, if there is a miss, a fetch is made the underlying ChunkStore. They are then populated with the result. GC runs a Purge operation on these caches. This makes all future read operations reach the ChunkStore layer during the GC operation, so that proper dependencies can be taken on them.
This change fixes a race where
Get -> Read -> Purge -> Putcould result in the cache carrying data throughout and after the GC which might have been read before the Purge operation itself. The caches now carry a purge count and they refuse to populate with values which were fetched based on a Get at a previous purge count.dolt fsckon the database directories anytime these tests fail.Prints boolean SQL results as 1 and 0 in text and CSV output to match MySQL.
Fixes dolt returns boolean values sometimes dolthub/dolt#6044
Rejects invalid destination branch names before pushing to remotes.
Fixes
pushshould validate branch name likecheckoutorbranchdolthub/dolt#5341Enables DOLT executable comments with the parser support in Support Dolt executable SQL comments dolthub/vitess#494.
Fixes Dolt should support it's own flavor of special comments. dolthub/dolt#3096
Adds jsonl SQL result formatting with one JSON object per row.
Fixes Allow jsonl result format dolthub/dolt#1844
ON UPDATE CURRENT_TIMESTAMPacross schemas and table alterations inEXTRAcolumnFix dolthub/dolt#11774
Blocked by dolthub/go-mysql-server#3862
Adds REPLICATE_WILD_DO_TABLE and REPLICATE_WILD_IGNORE_TABLE support for MySQL-to-Dolt binlog replication, including MySQL wildcard matching, precedence, case handling, validation, and SHOW REPLICA STATUS output.
Filter updates are applied atomically and synchronized with replica lifecycle operations.
Part of #11787
Depends on: dolthub/go-mysql-server#3885
Two types of checks are added. Simple checks are run on every index parse. They assert that prefixes are in sorted order, that ordinals are within range, that offsets (dervied from the lengths) are not too small. The more expensive checks only run when we add new table files to the store, typically as part of GC, a pull or receiving a push. These further assert that recorded ordinals are unique and that the hash of the suffixes table matches the file name, which is how the table file name is derived on all write paths.
These sanity checks are added to make Dolt more robust in the face of bugs in Dolt or unexpected behavior from I/O interactions. If these validation checks fail, Dolt has a chance to fail a write operation which is attempting to take a new dependency on the corrupted data.
Bumps google.golang.org/grpc from 1.83.1 to 1.83.2.
Release notes
Sourced from google.golang.org/grpc's releases.
Commits
030ee8bUpdate version to 1.83.2 (#9375)8668b69cherry-pick #9365 to v1.83.x (#9366)a3e952dcherry-pick #9346 to v1.83.x and update x/net dependency (#9369)58f8fd9Change version to 1.83.2-dev (#9337)This change allows for more flexible row merge behavior, in particular the ability to detect conflict on matching values and create rows dynamically based on the callers requirements.
DumboDB is the direct consumer of this change. This allows us to match Dumbo's standard compare-and-set query pattern, and to also enable collection specific merge modes.
There may be cleaner ways to implement this, but I opted to touch the common (dolt) code path as little as possible since this could have significant perf impacts. nil rowMergePolicy is the dolt way as a result.
doltgresql
Enable PostgreSQL UPDATE row counts through GMS, including unchanged rows and ON CONFLICT DO UPDATE. Add regression coverage for command tags, PL/pgSQL FOUND, RETURNING, triggers, and mixed conflict statements.
Depends on: Support PostgreSQL UPDATE row counts dolthub/go-mysql-server#3970
Fixes #3113
Fixes:
Lets all roles read from the public system tables, and corrects table and view permissions. Some tests are skipped due to pre-existing issues that are outside the scope of this PR.
Builds on:
Fixes:
SET CONSTRAINTSfails with an internal error dolthub/doltgresql#3365This doesn't add full support for constraints since that's a much larger change (requires deferrability), but it at least unblocks the non-deferring cases in the issue. We still block on deferring since it's a fundamental feature that we'd be incorrectly throwing away, causing divergent behavior versus the expectation (generally you only want deferring so certain queries would work, so we should fail before we can get to that point).
Fixes
Renders literals and cast column defaults in
information_schema.columnscorrectly. Also fixes negative literals inside casts being persisted as-5::TEXT, which re-parsed as-('5'::text)and broke inserts using that default.This implements
JSON_TABLE, which is required for:JSON_TABLEis refused with a syntax error atCOLUMNSdolthub/doltgresql#3331Implements missing array functions, slicing, subscript updates, FOREACH SLICE, and PostgreSQL-compatible array error codes while deferring custom bounds that require serialization changes.
Row constructors compared against subqueries now follow Postgres' NULL rules,
<-style row comparisons evaluate each field only once, and CHECK constraints containingROW(...)no longer fail on every insert. These were issues found in Fixed context-dependent record comparisons dolthub/doltgresql#3424 that are being fixed in a separate PR.Fixes:
Record comparisons are now context-dependent, which is a special case that is detailed specifically within the documentation. This also removes an analyzer rule that was previously added and is no longer necessary.
UPDATE OFforCREATE TRIGGERFixes:
This adds
UPDATE OFcolumn lists toCREATE TRIGGER. These triggers only fire when anUPDATEsets one of the listed columns, andpg_trigger.tgattrnow reports those columns. Also, renaming a listed column updates the trigger, and dropping one requiresCASCADE, which also drops the trigger, which were found while fixing this issue.XMLFixes:
xmltype does not exist, andxpath()is not found dolthub/doltgresql#3337This adds most of the
XMLsurface. The most notable exclusions are the*_to_xmlfamily of functions, since those would require a bit more work.Fixes #3386.
This implements multidimensional arrays. They're serialized by writing a flattened array with the dimensions appended at the end, so that existing array data can be read as-is. Array types use the new serialized version (
1) while all other non-array types use the old version (0), since the new version is primarily geared for array types, and creating new non-array types should retain their backwards compatibility. Deserialized types also carry their version, so reading an array type made on an older version simply restricts writing multidimensional data.Other fixes are present for regressions and test failures found along the way, with their relevant tests included in the test files.
Adds an opt-in server setting to accept and ignore unsupported locking statements.
This provides a partial workaround for the limitations in Missing PostgreSQL coordination primitives: FOR UPDATE SKIP LOCKED, LOCK TABLE IN SHARE ROW EXCLUSIVE dolthub/doltgresql#2600
The loop's expression is evaluated once into a record, which unnest then visits in order, assigning each element to the loop variable. An expression that does not yield an array, or that yields a null array, raises as PostgreSQL does. SLICE is not yet supported.
Fixes
UNNESTfails with an internal error dolthub/doltgresql#3366Allow
ORDER BYclauses to explicitly specify PostgreSQL's default NULL placement withASC NULLS LASTandDESC NULLS FIRST. Regenerate the parser and add coverage for implicit and explicit NULL ordering, multi-column sorts, window ordering, and ordered aggregates.Resolve common-type inference for
VALUESrows that combine an unknown NULL literal with a typed value, with coverage for allASC/DESCandNULLS FIRST/NULLS LASTcombinations.Fixes #3388
Depends on: go-mysql-server #3898
AssignTriggers runs long before GMS pads an INSERT's rows out to the table's schema, so it could only attach the triggers below that projection. NEW then held just the columns the INSERT named, in the order it named them, and reading a later one panicked on a short row. A new rule lifts the triggers above the projection once it exists.
An EXCEPTION reported P0001 whatever its USING ERRCODE clause said, and a notice reported that clause's source text, quotes and all. Both now resolve it to the SQLSTATE it names, and a generated RAISE can carry one through the same option.
The function was declared strict, so it returned null for a null argument. Its type is known regardless of its value, which is what PostgreSQL reports.
pg_query_go declares it int4 whatever the expression yields, so a CASE over any other type failed to assign.
A PL/pgSQL IF, WHILE, EXIT WHEN, or CASE whose condition evaluated to NULL crashed the query handler on a failed type assertion.
Referencing a RECORD variable, or a trigger's OLD or NEW, without accessing a field panicked. ApplyBindings now recognizes such a reference and renders the record from its fields. When the binding is interpolated into a statement, the record is modeled as a single-row derived table selected by its whole-row reference, so a function such as to_jsonb() sees its field names; RAISE interpolates the record's text representation instead.
Referencing a record that has never been assigned reports an error. IoOutput now returns an error for pseudo-types instead of reaching the registry with their placeholder I/O function.
A declaration's default was previously either a bare parameter name or the source text itself, which was parsed through the declared type's input function when evaluated. We now compile defaults to a query at CREATE time, like the right-hand side of an assignment. They are then evaluated when the declaration runs. We continue storing the source text for the default definition. Older releases still read and apply the defaults which they previously supported.
Whole-row comparisons now work in triggers, and other bare whole-row references report PostgreSQL's column-count errors.
WHEN (old.* IS DISTINCT FROM new.*)makes every UPDATE of its table fail dolthub/doltgresql#3336Adds the regnamespace type with its casts and to_regnamespace.
CHECK constraints that call functions no longer reject every row.
Named NOT NULL, DEFAULT, and UNIQUE column constraints are now accepted.
Implements information_schema.triggers.
information_schema.triggersis empty while the trigger exists and fires dolthub/doltgresql#3330information_schema.columns now reports generation_expression for generated columns.
generation_expressionininformation_schema.columnsis NULL for a generated column dolthub/doltgresql#3328Built-in functions are executable without an explicit EXECUTE grant.
Implements convert_from and decode.
convert_from()is not found dolthub/doltgresql#3326Casting a bpchar to text, varchar, or name now trims trailing spaces.
Default, generated column, and CHECK expressions now keep the parentheses that change their meaning.
Generated column expressions now survive ALTER TABLE ADD PRIMARY KEY.
Compound simple-query messages containing
COPY FROM STDINstopped executing after the COPY operation completed, so later statements were skipped and the server sentReadyForQuerytoo early. COPY data was also committed independently instead of remaining within the compound query’s implicit or explicit transaction.This change retains the active simple-query execution while COPY waits for client data, resumes the remaining statements after
CopyDone, and preserves the enclosing transaction’s commit and rollback behavior.CopyFailnow aborts the suspended query, rolls back implicit work, and marks explicit transactions failed with SQLSTATE57014.The regression replayer also incorrectly combined all recorded COPY data into one stream and resent it for every
CopyInResponse. It now preserves COPY stream boundaries and sends each recorded input exactly once, in statement order.go-mysql-server
ALTER rewrites starting from persisted schema could reclassify literal defaults as expressions, causing datetime COLUMN_DEFAULT values to gain quotes after subsequent column changes. Resolve defaults in the current schema before rewriting, preserving literal metadata and existing multi-column ALTER behavior.
Adds shared engine coverage for metadata, SHOW CREATE, and default insertion. All 50 Dolt default-value BATS cases, Dolt MODIFY tests, and relevant GMS suites pass; expectations were checked against MySQL 8.4.11.
Addresses the Lambda BATS failure in Dolt #11994.
Add an engine override to support PostgreSQL UPDATE row count behavior, where all matched rows are counted. Also simplifies INSERT ON DUPLICATE UPDATE handling.
Addresses UPDATE metadata semantics incorrect dolthub/doltgresql#3113.
Doltgres PR: Report PostgreSQL UPDATE row counts dolthub/doltgresql#3468.
Returns NULL with a warning for INET_NTOA arguments outside the IPv4 range.
Fixes
INET_NTOAcorrupts valid IPv4 values above signed 32-bit dolthub/dolt#11510Validates twelve-hour STR_TO_DATE inputs and normalizes midnight and noon consistently with MySQL.
Fixes Dolt ignores the PM marker in STR_TO_DATE dolthub/dolt#11544
Reinitializes disposed regular expressions when they are reused across window partitions.
Fixes Dolt returns a wrong result for REGEXP_LIKE across partitions dolthub/dolt#11550
Clamps unsigned MEDIUMINT serialization to 16777215.
Stores binary collation metadata explicitly instead of deriving it from collation names.
Fixes Add explicit metadata flag on
Collationfor binary collations dolthub/go-mysql-server#3838Recomputes stored generated columns and their dependents when MODIFY or CHANGE alters their definitions.
Fixes ALTER TABLE MODIFY COLUMN does not recompute a STORED generated column, leaving rows and indexes stale dolthub/dolt#11781
Returns NULL with a warning when CONVERT USING utf8mb4 receives malformed UTF-8.
Fixes Emit
ER_INVALID_CHARACTER_STRINGon malformed multi-byte character strings instead of returningErrCollationMalformedStringdolthub/go-mysql-server#3837Reinterprets BIGINT UNSIGNED values as signed integers when evaluating CAST AS SIGNED.
Fixes Incorrect CAST from BIGINT UNSIGNED to SIGNED dolthub/dolt#11906
Preserves negation on residual NOT IN predicates when sibling subqueries become anti joins.
Fixes Planner drops the NOT from the first of two ANDed NOT IN (subquery) predicates when its subquery is a UNION — query returns exactly the excluded rows dolthub/dolt#11913
Accepts binary and numeric arguments to STR_TO_DATE using MySQL-compatible string conversion.
Fixes STR_TO_DATE() with BLOB argument errors with "invalid type: []uint8" dolthub/dolt#11919
Cancels blocked grouping producers when the grouping consumer returns an error.
Allows column-free expressions and correlated subqueries that reference only grouped columns under ONLY_FULL_GROUP_BY.
Tracks COUNT DISTINCT allocations through the session memory manager after the grouping cancellation fix in #3954.
Validates temporal cast precision before constructing the target type.
Preserves correlated subquery column dependencies when grouping over joins.
Returns an empty iterator when column statistics are queried without initialized session privileges.
TIMEwith precisionImplement conversion for numeric types to
TIMEincluding precision parameters.Also includes conversion logic for
booleantypes and small fixes for string conversions.InternalDecimalTypeforCONVERTwith decimal argumentsFix regression from fix panic for
CASTandCONVERTto types with invalid precision/scale dolthub/go-mysql-server#3939dolt bump in [no-release-notes] Update CustomValueValidator package name and NewConvert call dolthub/dolt#11915
doltgres bump in [no-release-notes] Update CustomValueValidator package name dolthub/doltgresql#3421
Fixes a bug where subqueries could bypass permissions in some instances (test cases are now returning errors), and also adds an interface for:
INSERT ... ON DUPLICATE KEY UPDATE could rewrite rows without UPDATE privileges. Require UPDATE authorization and return the MySQL-compatible error code when access is denied.
CASTandCONVERTto types with invalid precision/scaleIdeally,
CAST/CONVERTshould only require the destinationsql.Typeto make the proper conversion.However, a significant portion of our conversion and comparison logic is so dependent on existing hackiness that it would require an entire rewrite to properly implement.
As a result,...
View the full release notes at https://github.com/dolthub/doltgresql/releases/tag/v1.4.0.