| title | First backup and restore | |||
|---|---|---|---|---|
| description | Take a full backup against a sandbox PostgreSQL and restore it into a fresh data dir, gated by pg_verifybackup. | |||
| tags |
|
Backs up a real PostgreSQL deployment, restores it into a sandbox data dir on the same host, and confirms restorability with
pg_verifybackup. About 10 minutes against a small database; the commands scale unchanged to a 100 TB cluster.
This tutorial walks the round-trip every operator should run before
trusting any backup tool — including this one. You finish with a
restored data directory you can pg_ctl start against, and the
verifier's stamp on the manifest.
This tutorial deliberately covers just the base-backup path —
no continuous WAL. The base backup is one of two pieces; in
production you also run pg_hardstorage wal stream 24/7 so that
PITR can roll forward between backups. See the
PITR walkthrough for the full picture.
If you have not installed yet, do Getting started first.
- A reachable PostgreSQL 15+ instance. The
postgressuperuser will do for the tutorial; in production you would use a dedicated replication role (see getting-started). - 2 GB free disk for the sandbox repo + restored data dir.
pg_hardstoragev0.2 or later on$PATH.pg_verifybackupfrom the matching PG client tools (ships withpostgresql-client-17etc.).
A throwaway PostgreSQL in Docker works for the tutorial:
docker run -d --name hs-tutorial-pg \
-e POSTGRES_PASSWORD=postgres \
-p 5432:5432 \
postgres:17Add some data so the round-trip proves something:
PGPASSWORD=postgres psql -h 127.0.0.1 -U postgres -c \
"CREATE TABLE hello (id int PRIMARY KEY, msg text);
INSERT INTO hello VALUES (1,'world'),(2,'restore-me');"# RUNNABLE
pg_hardstorage repo init file:///tmp/hs-tutorial-repoThe repo is just a directory: chunks/, manifests/, wal/,
audit/, plus a top-level HSREPO magic file. Re-running against an
existing repo returns conflict.repo_exists (exit 7) — the operation
is idempotent on the URL.
# RUNNABLE
pg_hardstorage backup db1 \
--pg-connection "${PG_CONNECTION:-postgres://postgres:postgres@127.0.0.1/postgres}" \
--repo file:///tmp/hs-tutorial-repo \
--include-wal--include-wal matters here. A base backup only becomes a consistent
database after replaying the WAL written while it ran. In production
wal stream archives that WAL continuously; this tutorial does not run
it, so the backup has to carry its own. Without the flag, backup
warns that the result is not restorable yet, and restore refuses
it (preflight.backup_wal_missing) rather than producing a data
directory that would wait forever for WAL that exists nowhere.
The pipeline is BASE_BACKUP over libpq → tar parser → FastCDC
chunker → CAS PUTs → signed manifest. On the first run a signing
keypair is generated under your keyring directory (run
pg_hardstorage doctor to print the exact path).
Sample output:
✓ Backup committed
ID: db1.full.20260504T120000Z.a1b2
Deployment: db1
PostgreSQL: 17
Cluster ID: 7659398055633653799
Stop LSN / TLI: 0/2000100 / 1
Files: 967 in 1 tablespace(s)
Logical bytes: 22.2 MiB
Unique chunks: 363 (8.7 MiB after dedup)
Dedup ratio: 2.56x
Duration: 1722 ms
Encryption: none
Manifest: manifests/db1/backups/db1.full.20260504T120000Z.a1b2/manifest.jsonThe backup ID has the shape db1.full.YYYYMMDDThhmmssZ.<hash> —
UTC, no local-zone surprises at 3am, plus a 4-char hash to keep
sub-second back-to-back backups distinct.
# RUNNABLE
pg_hardstorage list db1 --repo file:///tmp/hs-tutorial-repoBackups for db1 (1):
BACKUP ID TYPE WHEN FILES SIZE DEDUP DURATION
db1.full.20260504T120000Z.a1b2 full 2026-05-04 12:00 967 22.2 MiB 2.56x 1617 msThe backup ID has the shape db1.full.YYYYMMDDThhmmssZ.<hash> —
the trailing 4-char hash disambiguates backups taken in the same
second. Capture the latest one for the next steps:
# RUNNABLE
BACKUP_ID=$(pg_hardstorage list db1 --repo file:///tmp/hs-tutorial-repo -o json \
| grep -oE 'db1\.full\.[0-9TZ]+\.[0-9a-f]+' | head -1)
echo "BACKUP_ID=$BACKUP_ID"# RUNNABLE
pg_hardstorage show db1 "$BACKUP_ID" \
--repo file:///tmp/hs-tutorial-reposhow prints the LSN range, timeline, compression, tablespaces, dedup
ratio, and the manifest's ed25519 signature fingerprint. Pipe through
-o json if you want to parse it.
# RUNNABLE
pg_hardstorage verify db1 latest \
--repo file:///tmp/hs-tutorial-repoverify validates the manifest's ed25519 signature with your local
public key, then SHA-256-round-trips every referenced chunk through
the CAS read path. Encrypted backups are decrypted in-process. No
data dir is materialised.
For a much faster pre-flight that only checks chunk presence (no
fetch, no SHA), pass --existence-only. Useful before
backup undelete to confirm chunk-GC has not yet reaped the bytes.
# RUNNABLE
pg_hardstorage restore db1 latest \
--repo file:///tmp/hs-tutorial-repo \
--target /tmp/hs-tutorial-restoredSample output:
✓ Restore complete
Backup: db1.full.20260504T120000Z.a1b2
Deployment: db1
Target: /tmp/hs-tutorial-restored
Files: 967
Bytes written: 22.2 MiB
Chunks: 919
backup_label: 230 bytes
Duration: 868 ms
Verification: passedBy default (--verify=auto) a pg_verifybackup manifest check runs
against the restored data dir when the matching postgresql-client is
on the runner's PATH, and its result shows in the Verification: line.
--verify=require makes a failed check fail the restore (exit 9);
--verify=skip turns it off (audited; do not skip in production).
docker run --rm -d --name hs-tutorial-restored \
-v /tmp/hs-tutorial-restored:/var/lib/postgresql/data \
-p 5433:5432 \
-e POSTGRES_PASSWORD=postgres \
postgres:17PGPASSWORD=postgres psql -h 127.0.0.1 -p 5433 -U postgres \
-c "SELECT * FROM hello;" id | msg
----+------------
1 | world
2 | restore-me
(2 rows)Your row survived the round-trip.
docker rm -f hs-tutorial-pg hs-tutorial-restored
rm -rf /tmp/hs-tutorial-repo /tmp/hs-tutorial-restoredYou exercised the full data plane: a base backup over the replication
protocol, content-addressed chunk storage with a signed manifest, an
independent verify step that re-hashes every chunk, and a restore that
rebuilds the data directory and confirms it with the upstream
pg_verifybackup tool.
The repo on disk is the source of truth: rerun pg_hardstorage list
or pg_hardstorage show from another machine pointing at the same
URL and you get the same answers. Local agent state is regenerable —
deleting the cache is harmless.
- PITR walkthrough — replay WAL up to a natural-language timestamp.
- Encryption walkthrough — wrap chunks with a local KEK or AWS KMS.
- Operator guide — daily operations —
what
status,list, andshowtell you in production. - R3 — Cold start from backups — what to do when the source PG is gone, only the repo remains.