Skip to content

Bring back push on Chrome and Android devices the server switched off #417

Description

@HMarzban

Summary

Web push does not reach Chrome, Edge, Android, Brave, Opera or Samsung Internet users in production. Only Firefox and Safari (macOS and iOS Home Screen apps) receive push today.

  • Root cause: commit 6d19d16 (2026-08-09) made the SSRF filter refuse fcm.googleapis.com. The sender then switched each FCM subscription off at its first push, with the same text a real 404/410 writes. The fix is Stop the SSRF filter from treating fcm.googleapis.com as a private address #398 (commit 0f6fc54), not deployed yet.
  • Why it stays broken after that fix: nothing turns a switched-off device back on. The client re-registers only every 30 days, and Settings still shows "on".
  • Severity: High (user impact). Not a security finding.
  • Source: end-to-end push review of 2026-10-06 (Notes/local-docs/push-review-2026-10-06-client.md and -server.md).

Production data (2026-10-06, read-only)

  • 12 push subscriptions. 4 FCM rows of 3 users are switched off with last_error = 'Subscription expired or invalid', all after 2026-08-09.
  • 4 more FCM rows are still active only because no push was sent to them yet.
  • Queue push_notifications: 0 waiting; 5 messages archived in 21 days.
  • Of 11 mentions in 21 days, 7 went to users whose devices were all switched off.

Scope of the fix

Server — apps/hocuspocus.server/src/lib/push/sender.ts

  • S1. An endpoint refused by the SSRF check must not switch the row off. Skip it and log a warning. Do not write the 404/410 text (sender.ts:159-162 into :70-76).
  • S2. Pass web-push options: TTL (1 day) and a request timeout (10 s). Today there is no TTL (4-week default) and no timeout (sender.ts:176-182).
  • S3. Only a permanent failure counts toward switching a device off. 404/410 switch it off now. 429, 5xx and network errors must not raise failed_count. Store the status code in last_error, for example HTTP 403, so get_push_failure_summary can group it.

Client — apps/webapp

  • C1. Re-register the existing subscription at most once a day on a signed-in load (register_push_subscription upserts). Remove the 30-day unsubscribe-then-subscribe path, which can destroy a working subscription (src/utils/push-notifications.ts:321-382). Share one in-flight promise, because the prompt card and Settings both mount the hook.
  • C2. Permission granted, signed in, no browser subscription, and the user did not opt out: subscribe again (src/components/NotificationPromptCard.tsx:107-109, src/hooks/usePushNotifications.ts).
  • C3. Sign-out calls unregister_push_subscription and subscription.unsubscribe() before signOut() (src/components/settings/hooks/useSignOut.ts:12-30). Today a shared device keeps showing the previous user's messages.
  • C4. When the existing subscription's applicationServerKey differs from NEXT_PUBLIC_VAPID_PUBLIC_KEY, subscribe again (src/utils/push-notifications.ts:233-236).

Service worker — apps/webapp/public/service-worker.js

  • W1. Handle pushsubscriptionchange: subscribe again with the same key. C1 then saves it on the next load.
  • W2. Add titles for the server types that show "New notification" today: message, content_change, system_alert, invitation (service-worker.js:87-104 against packages/supabase/scripts/01-enum.sql:71-81).
  • W3. A bad or empty payload still shows a generic notification (service-worker.js:140-152). Chrome otherwise shows its own text and may revoke the subscription.
  • W4. notificationclick opens same-origin paths only; any other origin becomes / (service-worker.js:183-211).

Database — new migration plus packages/supabase/scripts/ mirror

  • D1. update_user_online_at stamps online_at only when status changes, so the 60 s heartbeat never refreshes it. Users in an open tab are treated as offline after 2 minutes, and users stuck at ONLINE lose regular-message push. Stamp online_at on every status write.
  • D2. update_notification_preferences accepts unchecked values that the push trigger later casts (::boolean, ::time, at time zone). One bad value aborts the messages insert for the whole channel. Validate the known keys and types in the RPC.

Out of scope (deliberately, to avoid overengineering)

Per-device online suppression, a new Android badge asset, iPhone and Android device-name parsing, a deactivation metric and alert, VAPID in /health, purging pgmq.a_push_notifications, a test-push button, and deploy-time key matching in compose. File them separately if wanted.

Runbook after merge (maintainer)

  1. Deploy (includes Stop the SSRF filter from treating fcm.googleapis.com as a private address #398, commit 0f6fc54). Note the deploy time.
  2. Soon after (a cron job deletes rows switched off for 30 days), run the recovery SQL. Use Notes/local-docs/push-review-2026-10-06-server.md (queries A, B, C), with <FIX_DEPLOYED_AT> set to the deploy time:
    • A and B are read-only previews. B should list 4 rows today.
    • C updates is_active = true, failed_count = 0, last_error = null for those rows, in one transaction. Compare the count with B, then commit.
  3. Check: send a mention to a Chrome user, and confirm a row in pgmq.a_push_notifications and a notification on the device.

Acceptance criteria

  • Chrome desktop and Android receive a mention push after deploy.
  • An SSRF-refused endpoint leaves the row active and logs a warning.
  • A 5xx or network error does not raise failed_count. A 404/410 still switches the row off.
  • A device the server switched off comes back within one day of the next signed-in visit.
  • After sign-out, that browser receives no push for the old account.
  • Every server notification type shows a real title.
  • A bad payload still shows a notification.
  • A user active in a tab keeps online_at fresh, and regular-message push still reaches devices of users who left.
  • A bad preference value is refused by update_notification_preferences, and sending messages never fails because of preferences.

Related

#398 (the FCM filter fix).


Generated by Claude Code

Activity

  1. added theissue type on Oct 6, 2026
  2. HMarzban commented on Oct 6, 2026

    @HMarzban
    CollaboratorAuthor

    Fix is on branch claude/youthful-lovelace-3nvsc6:

    • da0d832 server: S1–S3
    • ea384b6 client and service worker: C1–C4, W1–W4
    • 04fa4c0 database: D1–D2, with migration 20261006120000_push_online_at_and_preference_checks.sql and the regenerated seed.sql

    Every unit passed its adviser, deslop and verify stages. The code-janitor supervisor granted the production-ready flag on round 2; round 1 asked only to regenerate seed.sql. The full pre-push check passed: lint, format, security, typecheck, webapp Jest, backend tests, and the webapp and admin build:ci.

    Deploy order:

    1. Apply the migration.
    2. Deploy the apps, including Stop the SSRF filter from treating fcm.googleapis.com as a private address #398 (0f6fc54).
    3. Run the recovery SQL from the runbook above.

    Not yet checked on real devices: Safari and Android push end to end. Use the runbook's step 3.


    Generated by Claude Code

  3. changed the title [-]PWA push: Chrome and Android devices stay switched off, and nothing brings them back[/-] [+]Bring back push on Chrome and Android devices the server switched off[/+] on Oct 6, 2026
  4. HMarzban commented on Oct 6, 2026

    @HMarzban
    CollaboratorAuthor

    Facts the comment above does not record:

    The recovery SQL is not in the repo. Runbook step 2 points at Notes/local-docs/push-review-2026-10-06-server.md, which is gitignored. Queries B and C are copied here. Replace <FIX_DEPLOYED_AT> with the deploy time of 0f6fc54. Run them soon: a cron job deletes rows that stay switched off for 30 days. Run them only after 0f6fc54 is live. First confirm the worker logs no new Refused an unsafe push endpoint line for an FCM row.

    -- B. Read-only. The rows the update will change. Compare the count with C.
    select ps.id, ps.user_id, ps.device_name, ps.platform, ps.updated_at
    from public.push_subscriptions ps
    where ps.is_active = false
      and ps.last_error = 'Subscription expired or invalid'
      and ps.push_credentials->>'endpoint' ~* '^https://fcm\.googleapis\.com/'
      and ps.updated_at >= timestamptz '2026-08-09 13:43:49+00'  -- commit 6d19d16
      and ps.updated_at <  timestamptz '<FIX_DEPLOYED_AT>'
      and not exists (
            select 1 from public.push_subscriptions a
            where a.user_id = ps.user_id and a.is_active
              and a.push_credentials->>'endpoint' = ps.push_credentials->>'endpoint')
    order by ps.updated_at;
    
    -- C. The update. Same filter as B. Check the returned count, then commit or roll back.
    begin;
    update public.push_subscriptions ps
       set is_active = true, failed_count = 0, last_error = null
     where ps.is_active = false
       and ps.last_error = 'Subscription expired or invalid'
       and ps.push_credentials->>'endpoint' ~* '^https://fcm\.googleapis\.com/'
       and ps.updated_at >= timestamptz '2026-08-09 13:43:49+00'
       and ps.updated_at <  timestamptz '<FIX_DEPLOYED_AT>'
       and not exists (
             select 1 from public.push_subscriptions a
             where a.user_id = ps.user_id and a.is_active
               and a.push_credentials->>'endpoint' = ps.push_credentials->>'endpoint')
    returning ps.id, ps.user_id;
    -- commit;  -- or: rollback;

    A dead endpoint gets 404 or 410 at its next send and is switched off again. That is the correct result.

    Migrations go up by hand. Supabase SQL has no deploy pipeline. Apply 20261006120000_push_online_at_and_preference_checks.sql before the app deploy. The branch also carries 20261006130000_close_client_write_and_grant_gaps.sql (#397, #401, #409, #410).

    D2 checks new writes only. Values stored before the migration are not checked again. A bad stored value can still abort a messages insert.

    C2 depends on a local stamp. The client repairs a missing browser subscription only when localStorage holds docsplus_push_subscription_timestamp. The stamp is the opt-in record. No subscription and no stamp counts as "opted out".


    Generated by Claude Code

  5. HMarzban commented on Oct 9, 2026

    @HMarzban
    CollaboratorAuthor

    Landed on main, not deployed yet

    Reopened: the board automation closed this issue before the push. It stays open until the prod steps are done.

    The branch merged into main as 083f37d84, with no rebase, so da0d832, ea384b6 and 04fa4c0 above still resolve. Follow-ups on top:

    • 946238c61: a row that the URL filter refuses no longer makes the push job retry and dead-letter. A bad subscription key now adds 1 to failed_count. Sign-out waits for an in-flight push sync.
    • 2a2f5c99f: the service worker no longer handles pushsubscriptionchange. The next signed-in load repairs a changed subscription instead.
    • 9e00856ec and 62c9211a1: smaller tidy-ups of the same code, with no behavior change.

    Prod steps:

    1. Apply migration 20261006120000_push_online_at_and_preference_checks.sql.
    2. Deploy the backend and the webapp.
    3. Run the recovery SQL (queries B, then C, in the comment above). Do it soon: a cron job deletes rows that stay switched off for 30 days.
    4. Check Safari and Android push end to end.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions