Skip to content

Crash reporting via Grafana Faro (opt-in) #246

Description

@srsholmes

The problem

Loadout runs in Gaming Mode on other people's handhelds, which is close to the
worst possible environment for finding out that something broke. There is no
terminal in sight, no console, and often no keyboard. When loadout falls over,
the user sees a panel that stopped working — and that is all anyone learns.

Every diagnostic path today is local and user-initiated: copy to clipboard,
"Save logs to file", journalctl. Each requires the user to notice a problem,
care enough to file an issue, find the right log, and paste it. Realistically
that happens for a small fraction of real faults, and it self-selects for the
most motivated users on the most obvious bugs.

The part that makes this worse than it looks

The backend deliberately swallows every uncaught exception except OOM in
production:

// apps/loadout/src/loader/index.ts
process.on("uncaughtException", (err) => {
  log.error(`Uncaught exception: ${err.stack ?? err.message}`);
  if (shouldRethrowUncaught(err)) throw err;   // production: false
});

This is correct and deliberate — plugins run in-process in the root daemon, so
without it one misbehaving plugin takes down the server and every other plugin
with it (documented as Audit 2026-05 A-009, pending process isolation).

But the consequence is that an entire class of fault is invisible even
locally
. Not just "the user didn't report it" — nobody finds out at all. The
plugin stops working, the log gets one line, and the process carries on. Wiring
reporting into that handler is the single highest-value thing in this issue,
and it is worth doing even if nobody ever opts in, because it forces us to
treat those exceptions as events rather than noise.

Related blind spots this should cover:

  • The overlay's Bun main process had no global error handlers at all — an
    uncaught throw killed it silently and systemd restarted it, leaving a gap in
    the journal and nothing else.
  • Plugin load and onLoad failures are caught and logged, with no user-visible
    signal — a plugin that silently fails to load looks identical to one that is
    merely disabled.
  • FFI segfaults, CEF renderer crashes, and OOM kills are invisible to any JS
    handler.

What we actually want to answer

  1. Is this crash happening to one person with an odd setup, or to everyone?
  2. Did we ship a regression? Which release started it?
  3. Which device/distro? The OS-compatibility matrix is wide (SteamOS, Bazzite,
    CachyOS; Deck, OneXPlayer, ROG Ally, Legion Go, GPD, AYANEO) and the bug
    history is full of device-specific and SteamOS-update-specific breakage.
  4. Is it loadout's fault or a plugin's?

Decision: Grafana, not Sentry

The issue title says Sentry. We're going with Grafana Cloud Frontend
Observability (Faro) instead.

Why Faro is the right shape

Loadout ships to strangers' devices, so the ingest endpoint has to be safe to
embed in a public binary. That single constraint decides most of this:

Option Verdict
Faro collector ✅ Public write-only collect/{app-key} URL. This is the RUM threat model — untrusted clients on the open internet — which is exactly ours.
Grafana Cloud OTLP ❌ Needs a bearer/basic token, and Grafana's docs explicitly advise against pushing direct from an application in production. Shipping a real credential in a root daemon is not acceptable.
Self-hosted Alloy faro.receiver ❌ Docs state it isn't designed for public exposure and doesn't work with Cloud Frontend Observability. Built for local-network clients.

It does real error triage

Worth stating plainly, because it was the main worry when weighing this against
Sentry. Frontend Observability does automatic error grouping with a
four-layer cascading fingerprint that the docs describe as "similar to
Sentry's approach"
:

  1. User-provided fingerprints (highest priority)
  2. Normalized stack traces — application frames only, dynamic values stripped
  3. File sequences + canonical message (for minified code)
  4. Error type + canonical message (no stack trace)

Layers 1 and 2 line up almost exactly with what the implementation already
does: it computes a stable fingerprint() after scrubbing, specifically so the
same bug from two different users collapses to one issue, and it already marks
frames in_app so runtime/vendor frames don't pollute grouping.

Plus, on the Grafana side we get sourcemap-based stack-trace transformation
server-side, and Grafana Assistant root-cause analysis on error groups.

Practical reasons

  • Quota headroom. Free tier is ~50,000 sessions/month against Sentry's
    5,000 errors/month. This matters more than it sounds: loadout has shipped
    crash-loop bugs before (the SteamOS lib-soname drift), and on Sentry's free
    tier a single crash-looping device can exhaust the monthly quota in an
    afternoon — after which ingest silently stops and we go blind exactly
    when something is badly broken.
  • Correlation. Errors land in Loki alongside everything else, queryable
    with LogQL rather than siloed in a separate product.
  • We already run Grafana. Dogfooding our own stack is a legitimate reason.

What this costs us

Honestly: some ergonomics. Sentry's issue-triage workflow (assign, resolve,
ignore, regression detection) is more mature as a product surface. We're
betting that error grouping + LogQL + Assistant is enough for a project this
size, and that being able to correlate frontend errors with everything else is
worth more than the polish.

Cost in code is low. The implementation deliberately isolated the
transport: scrub.ts, consent.ts, rate-limit.ts, spool.ts and every
capture point are protocol-agnostic. Only transport.ts and the event shape in
types.ts change — roughly 200 lines.


Privacy design

This is the part most likely to go wrong in public, so it's deliberate.

The original plan in this issue was "default to on". We're not doing that.
Loadout's audience is Decky-adjacent Linux homebrew — the demographic least
tolerant of telemetry — and the backend runs as root. For reference:
Audacity shipped telemetry that was opt-in and disabled by default and was
still criticised hard enough to force full removal, spawning 50+ forks;
Heroic Games Launcher, our closest peer, keeps analytics opt-in and off.

Instead:

  • Explicit prompt, no silent default. One screen, yes/no, neither
    preselected. Declining is remembered and never re-asked.
  • Crash reports only. No usage analytics, no feature tracking, no
    performance tracing, no session replay. This line has to stay literally true
    — it is what makes the feature defensible.
  • Consent is tri-state (granted / denied / unset). The third state
    matters: every existing install already has welcomeCompleted: true, so a
    step bolted onto the welcome wizard would only ever be seen by fresh
    installs. Distinguishing "not asked" from "said no" is what lets upgraders be
    prompted exactly once.
  • Fail closed everywhere. Anything that isn't literally "granted" —
    missing file, parse error, permissions problem — means nothing is sent.
  • docs/privacy.md states what is collected, what is not, and shows a real
    example payload. It must ship before any transmission is enabled.

Scrubbing

Reports are scrubbed before they leave the device, and the scrubber is treated
as security-relevant code: pure, unit-tested, with a "never-leaks gate" that
asserts on the serialized bytes rather than intermediate objects.

Removed: home directories (/home, /var/home, removable-media mounts),
usernames, hostnames, Steam account identifiers (the 32-bit account id used in
userdata/ paths, plus SteamID64/ID3/ID2 forms), API keys and bearer tokens.
Plugin install paths are normalised to <plugins>/<id> — which is needed for
cross-user grouping anyway, so it pays for itself twice.

Known limitation, stated openly in the privacy doc: a revealing folder or game
name
inside an error message cannot be pattern-matched away. Paths are
rewritten so the username is gone, but ~/Games/Some Game survives.

Abuse and quota safety

  • Rate limiting persists to disk. An in-memory limiter is useless here: a
    crash loop restarts the process and resets the counter, so the loop would
    report at full rate forever.
  • Fatal crashes are spooled to disk, not sent — a dying process can't
    await a network request. The next successful start drains the queue, with a
    retry counter so a permanently-unsendable event can't retry forever. This
    also makes offline the normal case rather than an error case, which matters
    on a handheld that is frequently suspended.
  • Third-party plugin faults are tagged and routed away from the core feed, so a
    broken plugin doesn't read as a loadout bug.

Status

Core plumbing is up as a draft PR (#249), currently inert — no ingest
endpoint configured and no consent UI, so nothing is sent. That's deliberate:
it lets the scrubber and rate limiter be exercised against real crashes on real
devices before a single byte leaves anyone's machine.

Two rounds of multi-agent review have been run against it, with every finding
adversarially verified; 21 confirmed issues fixed so far.

Remaining

  • Swap the transport from the Sentry wire format to Faro
  • Consent UI — welcome step, upgrader prompt, Settings toggle
    (⚠️ useConfigValue<T>(key, default) collapses undefined into the
    default and cannot express the tri-state; needs hasConfigValue,
    and must gate on whenUserConfigLoaded())
  • Grafana Cloud app + collector URL, wired as a build-time define rather
    than an env var (an env var is retargetable by anything that can edit a
    unit file, and this is a root daemon)
  • Source maps — the webview ships minified with no maps, so its traces are
    currently unreadable
  • Move backend egress behind the overlay (root writes a spool file, the
    user-level process does all sending) so the root daemon opens no sockets
  • Native crash proxy via ExecStopPost + $SERVICE_RESULT for FFI
    segfaults / CEF crashes / OOM kills
    (⚠️ CI diffs the unit byte-for-byte against the heredoc in
    scripts/install.sh — both must be edited identically)
  • PRIVACY.md linked from README, and a "Does loadout phone home?" FAQ
    entry — that question will appear within a day of release

Open question

The exact wire format for POSTing a Faro payload without the Web SDK isn't
fully covered in the public docs. The intent is to emit it directly, as the
current implementation does for Sentry — it keeps zero runtime dependencies,
avoids the question of whether the SDK behaves under bun build --compile, and
preserves the property that every transmitted field is spelled out in our own
source, which is what makes the privacy doc checkable. Needs confirming against
the faro-web-sdk fetch transport.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions