Skip to content

Merge in v4.1.4 - #24

Merged
hughcapet merged 26 commits into
multisitefrom
multisite-v4.1.4
Jul 16, 2026
Merged

hughcapet merged 26 commits into
multisitefrom
multisite-v4.1.4

Conversation

@hughcapet

Copy link
Copy Markdown

No description provided.

hughcapet and others added 26 commits May 5, 2026 11:35
When data directory does not exist, `datadir_major == 0` and `bin_major
!= datadir_major`. Patroni used to erroneously return version 0 in such
cases, hence version specific features could not work correctly.

ref:
https://www.postgresql.org/message-id/CAFh8B=nP+xpSAgA6vMvoLAiowo3zQHtecZn-3VzXNruvx7cUMw@mail.gmail.com
Change 'timeout' to '--timeout' to match the other options.
Current etcd raises Unknown error when etcd leader is lost while updating the lease.
Fixed upstream in etcd-io/etcd#21671, but for older versions we want to override the reported error code.

Close patroni#3407
Create `PatroniLogger` before loading `Config` to capture early log messages.

Previously it was created inside `AbstractPatroniDaemon.__init__()`. This caused some INFO messages to be lost.
)

systemd 257+ requires MONOTONIC_USEC alongside RELOADING=1 for Type=notify-reload services. Without it, `systemctl reload` hangs indefinitely.

Close patroni#3594
…ni#3599)

Instead of logging at startup that systemd integration is not supported, check for NOTIFY_SOCKET and warn only when actually running under systemd without the python-systemd package installed.
- Fix for patroni#3576: `ensure_clean_shutdown()` corrupts WAL chain when
starting a replica restored from an external backup
- Check for `backup_label` file at the start of
`_handle_crash_recovery()` — if present, skip single-user crash recovery
and let PostgreSQL handle it during normal startup
- Add rn
- Update version
- Update pyright
to avoid FileNotFoundError: [Errno 2] No such file or directory when pkg
is installed but socket does not exist

plus, only try to import package/use systemd.daemon.notify() is env var
is set
Remove mem address from the CaseInsensitiveDict/CaseInsensitiveSet repr
and fix deep_compare to properly handle it
…atroni#3609)

After getting the version from connection, remove all auth params not
applicable if they were accidentally picked from for environment
Summary:
- Add usage documentation for patronictl demote-cluster.
- Add usage documentation for patronictl promote-cluster.
- Document each command's synopsis, behavior, parameters, and a basic
example.
It is achieved by modifying _get_local_timeline_lsn() method so that it
uses pg_controldata as a fallback when postgres is running but not yet
accepting connections and remove not is_starting restrictions in ha.py

Close patroni#3650
`--scheduled` was missing from the `patronictl switchover`
documentation, and other whitespace cleanup
### What
`CaseInsensitiveDict.keys()` returned normalized lowercase keys, while
iteration and `items()` kept the last-seen casing. Tiny follow-up to
patroni#3613.

### Repro
```python
from patroni.collections import CaseInsensitiveDict

values = CaseInsensitiveDict({"a": "b", "A": "B", "MixedCase": 1})
print(list(values))
print(list(values.keys()))
print(list(values.items()))
```

Before, `keys()` printed `["a", "mixedcase"]`.
After, it prints `["A", "MixedCase"]`, same vibe as iteration and
`items()`.

### Why
The class docs say `keys()` should expose the case-sensitive keys.
Returning internal normalized keys is a small footgun for normal mapping
API callers.
)

While I was reviewing the api.py to review prometheus metrics I realised
`patroni_postgres_timeline` is declared as `counter` in the `/metrics`
endpoint. Prometheus counters must only ever increase, but the metric's
own HELP text documents that it drops to `0` when PostgreSQL is not
running: "Postgres timeline of this node (if running), 0 otherwise."

Every PostgreSQL restart — via `patronictl restart`, a config change
requiring restart, or crash recovery — produces the scrape sequence `1 →
0 → 1`. Prometheus interprets the drop as a counter reset and applies
reset compensation:

```bash
  increase(patroni_postgres_timeline[5m]) = 2   (should be 0)
  rate(patroni_postgres_timeline[5m])     > 0   (should be 0)
```

**Any alert on timeline changes such as
`increase(patroni_postgres_timeline[5m]) > 0` fires on every routine
restart rather than only on actual failovers.**

The metric represents current state, not a cumulative count. Change its
declared type from `counter` to `gauge`.

To prevent false positive timeline alerts on the prometheus side we can
update the type of the metric.
…atroni#3634)

Documentation has an added "s" to the "retry_timeout" parameter

Co-authored-by: Greg Clough <gclough@drwuk.com>
In many places we are explicitly using `connect_timeout=3` and `options='-c statement_timeout=2000'`, which are essentially a default.
Therefore, it is logical to use them as default when setting `ConnectionPool.conn_kwargs` or when calling `get_connection_cursor()` function.

We explicitly pass `options` and `connect_timeout` only in case we want to override defaults.
In case of statement timeout we will use cached role fallback to avoid
demoting primary.

Besides that forcibly set `pg_stat_statements.track` to `none` to avoid
expensive GC calls.

Close patroni#3646 and patroni#3637
Besides that fixed bug in SlotsHandler.copy_slot_items().

Close patroni#3641
…ni#3666)

they are guaranteed to be local connections opened by superuser
to avoid `Error: No CtlPostgresqlRole.REPLICA among provided members`
RNs, change version, update pyright version
@hughcapet hughcapet changed the title Multisite v4.1.4 Merge in v4.1.4 Jul 16, 2026
@hughcapet
hughcapet marked this pull request as ready for review July 16, 2026 09:16
@hughcapet
hughcapet merged commit 8455b25 into multisite Jul 16, 2026
43 of 46 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants