The problem
Loadout runs in Gaming Mode on other people's handhelds, which is close to the
worst possible environment for finding out that something broke. There is no
terminal in sight, no console, and often no keyboard. When loadout falls over,
the user sees a panel that stopped working — and that is all anyone learns.
Every diagnostic path today is local and user-initiated: copy to clipboard,
"Save logs to file", journalctl. Each requires the user to notice a problem,
care enough to file an issue, find the right log, and paste it. Realistically
that happens for a small fraction of real faults, and it self-selects for the
most motivated users on the most obvious bugs.
The part that makes this worse than it looks
The backend deliberately swallows every uncaught exception except OOM in
production:
// apps/loadout/src/loader/index.ts
process.on("uncaughtException", (err) => {
log.error(`Uncaught exception: ${err.stack ?? err.message}`);
if (shouldRethrowUncaught(err)) throw err; // production: false
});
This is correct and deliberate — plugins run in-process in the root daemon, so
without it one misbehaving plugin takes down the server and every other plugin
with it (documented as Audit 2026-05 A-009, pending process isolation).
But the consequence is that an entire class of fault is invisible even
locally. Not just "the user didn't report it" — nobody finds out at all. The
plugin stops working, the log gets one line, and the process carries on. Wiring
reporting into that handler is the single highest-value thing in this issue,
and it is worth doing even if nobody ever opts in, because it forces us to
treat those exceptions as events rather than noise.
Related blind spots this should cover:
- The overlay's Bun main process had no global error handlers at all — an
uncaught throw killed it silently and systemd restarted it, leaving a gap in
the journal and nothing else.
- Plugin load and
onLoad failures are caught and logged, with no user-visible
signal — a plugin that silently fails to load looks identical to one that is
merely disabled.
- FFI segfaults, CEF renderer crashes, and OOM kills are invisible to any JS
handler.
What we actually want to answer
- Is this crash happening to one person with an odd setup, or to everyone?
- Did we ship a regression? Which release started it?
- Which device/distro? The OS-compatibility matrix is wide (SteamOS, Bazzite,
CachyOS; Deck, OneXPlayer, ROG Ally, Legion Go, GPD, AYANEO) and the bug
history is full of device-specific and SteamOS-update-specific breakage.
- Is it loadout's fault or a plugin's?
Decision: Grafana, not Sentry
The issue title says Sentry. We're going with Grafana Cloud Frontend
Observability (Faro) instead.
Why Faro is the right shape
Loadout ships to strangers' devices, so the ingest endpoint has to be safe to
embed in a public binary. That single constraint decides most of this:
| Option |
Verdict |
| Faro collector |
✅ Public write-only collect/{app-key} URL. This is the RUM threat model — untrusted clients on the open internet — which is exactly ours. |
| Grafana Cloud OTLP |
❌ Needs a bearer/basic token, and Grafana's docs explicitly advise against pushing direct from an application in production. Shipping a real credential in a root daemon is not acceptable. |
Self-hosted Alloy faro.receiver |
❌ Docs state it isn't designed for public exposure and doesn't work with Cloud Frontend Observability. Built for local-network clients. |
It does real error triage
Worth stating plainly, because it was the main worry when weighing this against
Sentry. Frontend Observability does automatic error grouping with a
four-layer cascading fingerprint that the docs describe as "similar to
Sentry's approach":
- User-provided fingerprints (highest priority)
- Normalized stack traces — application frames only, dynamic values stripped
- File sequences + canonical message (for minified code)
- Error type + canonical message (no stack trace)
Layers 1 and 2 line up almost exactly with what the implementation already
does: it computes a stable fingerprint() after scrubbing, specifically so the
same bug from two different users collapses to one issue, and it already marks
frames in_app so runtime/vendor frames don't pollute grouping.
Plus, on the Grafana side we get sourcemap-based stack-trace transformation
server-side, and Grafana Assistant root-cause analysis on error groups.
Practical reasons
- Quota headroom. Free tier is ~50,000 sessions/month against Sentry's
5,000 errors/month. This matters more than it sounds: loadout has shipped
crash-loop bugs before (the SteamOS lib-soname drift), and on Sentry's free
tier a single crash-looping device can exhaust the monthly quota in an
afternoon — after which ingest silently stops and we go blind exactly
when something is badly broken.
- Correlation. Errors land in Loki alongside everything else, queryable
with LogQL rather than siloed in a separate product.
- We already run Grafana. Dogfooding our own stack is a legitimate reason.
What this costs us
Honestly: some ergonomics. Sentry's issue-triage workflow (assign, resolve,
ignore, regression detection) is more mature as a product surface. We're
betting that error grouping + LogQL + Assistant is enough for a project this
size, and that being able to correlate frontend errors with everything else is
worth more than the polish.
Cost in code is low. The implementation deliberately isolated the
transport: scrub.ts, consent.ts, rate-limit.ts, spool.ts and every
capture point are protocol-agnostic. Only transport.ts and the event shape in
types.ts change — roughly 200 lines.
Privacy design
This is the part most likely to go wrong in public, so it's deliberate.
The original plan in this issue was "default to on". We're not doing that.
Loadout's audience is Decky-adjacent Linux homebrew — the demographic least
tolerant of telemetry — and the backend runs as root. For reference:
Audacity shipped telemetry that was opt-in and disabled by default and was
still criticised hard enough to force full removal, spawning 50+ forks;
Heroic Games Launcher, our closest peer, keeps analytics opt-in and off.
Instead:
- Explicit prompt, no silent default. One screen, yes/no, neither
preselected. Declining is remembered and never re-asked.
- Crash reports only. No usage analytics, no feature tracking, no
performance tracing, no session replay. This line has to stay literally true
— it is what makes the feature defensible.
- Consent is tri-state (
granted / denied / unset). The third state
matters: every existing install already has welcomeCompleted: true, so a
step bolted onto the welcome wizard would only ever be seen by fresh
installs. Distinguishing "not asked" from "said no" is what lets upgraders be
prompted exactly once.
- Fail closed everywhere. Anything that isn't literally
"granted" —
missing file, parse error, permissions problem — means nothing is sent.
docs/privacy.md states what is collected, what is not, and shows a real
example payload. It must ship before any transmission is enabled.
Scrubbing
Reports are scrubbed before they leave the device, and the scrubber is treated
as security-relevant code: pure, unit-tested, with a "never-leaks gate" that
asserts on the serialized bytes rather than intermediate objects.
Removed: home directories (/home, /var/home, removable-media mounts),
usernames, hostnames, Steam account identifiers (the 32-bit account id used in
userdata/ paths, plus SteamID64/ID3/ID2 forms), API keys and bearer tokens.
Plugin install paths are normalised to <plugins>/<id> — which is needed for
cross-user grouping anyway, so it pays for itself twice.
Known limitation, stated openly in the privacy doc: a revealing folder or game
name inside an error message cannot be pattern-matched away. Paths are
rewritten so the username is gone, but ~/Games/Some Game survives.
Abuse and quota safety
- Rate limiting persists to disk. An in-memory limiter is useless here: a
crash loop restarts the process and resets the counter, so the loop would
report at full rate forever.
- Fatal crashes are spooled to disk, not sent — a dying process can't
await a network request. The next successful start drains the queue, with a
retry counter so a permanently-unsendable event can't retry forever. This
also makes offline the normal case rather than an error case, which matters
on a handheld that is frequently suspended.
- Third-party plugin faults are tagged and routed away from the core feed, so a
broken plugin doesn't read as a loadout bug.
Status
Core plumbing is up as a draft PR (#249), currently inert — no ingest
endpoint configured and no consent UI, so nothing is sent. That's deliberate:
it lets the scrubber and rate limiter be exercised against real crashes on real
devices before a single byte leaves anyone's machine.
Two rounds of multi-agent review have been run against it, with every finding
adversarially verified; 21 confirmed issues fixed so far.
Remaining
Open question
The exact wire format for POSTing a Faro payload without the Web SDK isn't
fully covered in the public docs. The intent is to emit it directly, as the
current implementation does for Sentry — it keeps zero runtime dependencies,
avoids the question of whether the SDK behaves under bun build --compile, and
preserves the property that every transmitted field is spelled out in our own
source, which is what makes the privacy doc checkable. Needs confirming against
the faro-web-sdk fetch transport.
The problem
Loadout runs in Gaming Mode on other people's handhelds, which is close to the
worst possible environment for finding out that something broke. There is no
terminal in sight, no console, and often no keyboard. When loadout falls over,
the user sees a panel that stopped working — and that is all anyone learns.
Every diagnostic path today is local and user-initiated: copy to clipboard,
"Save logs to file",
journalctl. Each requires the user to notice a problem,care enough to file an issue, find the right log, and paste it. Realistically
that happens for a small fraction of real faults, and it self-selects for the
most motivated users on the most obvious bugs.
The part that makes this worse than it looks
The backend deliberately swallows every uncaught exception except OOM in
production:
This is correct and deliberate — plugins run in-process in the root daemon, so
without it one misbehaving plugin takes down the server and every other plugin
with it (documented as Audit 2026-05 A-009, pending process isolation).
But the consequence is that an entire class of fault is invisible even
locally. Not just "the user didn't report it" — nobody finds out at all. The
plugin stops working, the log gets one line, and the process carries on. Wiring
reporting into that handler is the single highest-value thing in this issue,
and it is worth doing even if nobody ever opts in, because it forces us to
treat those exceptions as events rather than noise.
Related blind spots this should cover:
uncaught throw killed it silently and systemd restarted it, leaving a gap in
the journal and nothing else.
onLoadfailures are caught and logged, with no user-visiblesignal — a plugin that silently fails to load looks identical to one that is
merely disabled.
handler.
What we actually want to answer
CachyOS; Deck, OneXPlayer, ROG Ally, Legion Go, GPD, AYANEO) and the bug
history is full of device-specific and SteamOS-update-specific breakage.
Decision: Grafana, not Sentry
The issue title says Sentry. We're going with Grafana Cloud Frontend
Observability (Faro) instead.
Why Faro is the right shape
Loadout ships to strangers' devices, so the ingest endpoint has to be safe to
embed in a public binary. That single constraint decides most of this:
collect/{app-key}URL. This is the RUM threat model — untrusted clients on the open internet — which is exactly ours.faro.receiverIt does real error triage
Worth stating plainly, because it was the main worry when weighing this against
Sentry. Frontend Observability does automatic error grouping with a
four-layer cascading fingerprint that the docs describe as "similar to
Sentry's approach":
Layers 1 and 2 line up almost exactly with what the implementation already
does: it computes a stable
fingerprint()after scrubbing, specifically so thesame bug from two different users collapses to one issue, and it already marks
frames
in_appso runtime/vendor frames don't pollute grouping.Plus, on the Grafana side we get sourcemap-based stack-trace transformation
server-side, and Grafana Assistant root-cause analysis on error groups.
Practical reasons
5,000 errors/month. This matters more than it sounds: loadout has shipped
crash-loop bugs before (the SteamOS lib-soname drift), and on Sentry's free
tier a single crash-looping device can exhaust the monthly quota in an
afternoon — after which ingest silently stops and we go blind exactly
when something is badly broken.
with LogQL rather than siloed in a separate product.
What this costs us
Honestly: some ergonomics. Sentry's issue-triage workflow (assign, resolve,
ignore, regression detection) is more mature as a product surface. We're
betting that error grouping + LogQL + Assistant is enough for a project this
size, and that being able to correlate frontend errors with everything else is
worth more than the polish.
Cost in code is low. The implementation deliberately isolated the
transport:
scrub.ts,consent.ts,rate-limit.ts,spool.tsand everycapture point are protocol-agnostic. Only
transport.tsand the event shape intypes.tschange — roughly 200 lines.Privacy design
This is the part most likely to go wrong in public, so it's deliberate.
The original plan in this issue was "default to on". We're not doing that.
Loadout's audience is Decky-adjacent Linux homebrew — the demographic least
tolerant of telemetry — and the backend runs as root. For reference:
Audacity shipped telemetry that was opt-in and disabled by default and was
still criticised hard enough to force full removal, spawning 50+ forks;
Heroic Games Launcher, our closest peer, keeps analytics opt-in and off.
Instead:
preselected. Declining is remembered and never re-asked.
performance tracing, no session replay. This line has to stay literally true
— it is what makes the feature defensible.
granted/denied/ unset). The third statematters: every existing install already has
welcomeCompleted: true, so astep bolted onto the welcome wizard would only ever be seen by fresh
installs. Distinguishing "not asked" from "said no" is what lets upgraders be
prompted exactly once.
"granted"—missing file, parse error, permissions problem — means nothing is sent.
docs/privacy.mdstates what is collected, what is not, and shows a realexample payload. It must ship before any transmission is enabled.
Scrubbing
Reports are scrubbed before they leave the device, and the scrubber is treated
as security-relevant code: pure, unit-tested, with a "never-leaks gate" that
asserts on the serialized bytes rather than intermediate objects.
Removed: home directories (
/home,/var/home, removable-media mounts),usernames, hostnames, Steam account identifiers (the 32-bit account id used in
userdata/paths, plus SteamID64/ID3/ID2 forms), API keys and bearer tokens.Plugin install paths are normalised to
<plugins>/<id>— which is needed forcross-user grouping anyway, so it pays for itself twice.
Known limitation, stated openly in the privacy doc: a revealing folder or game
name inside an error message cannot be pattern-matched away. Paths are
rewritten so the username is gone, but
~/Games/Some Gamesurvives.Abuse and quota safety
crash loop restarts the process and resets the counter, so the loop would
report at full rate forever.
await a network request. The next successful start drains the queue, with a
retry counter so a permanently-unsendable event can't retry forever. This
also makes offline the normal case rather than an error case, which matters
on a handheld that is frequently suspended.
broken plugin doesn't read as a loadout bug.
Status
Core plumbing is up as a draft PR (#249), currently inert — no ingest
endpoint configured and no consent UI, so nothing is sent. That's deliberate:
it lets the scrubber and rate limiter be exercised against real crashes on real
devices before a single byte leaves anyone's machine.
Two rounds of multi-agent review have been run against it, with every finding
adversarially verified; 21 confirmed issues fixed so far.
Remaining
(
useConfigValue<T>(key, default)collapsesundefinedinto thedefault and cannot express the tri-state; needs
hasConfigValue,and must gate on
whenUserConfigLoaded())than an env var (an env var is retargetable by anything that can edit a
unit file, and this is a root daemon)
currently unreadable
user-level process does all sending) so the root daemon opens no sockets
ExecStopPost+$SERVICE_RESULTfor FFIsegfaults / CEF crashes / OOM kills
(
scripts/install.sh— both must be edited identically)PRIVACY.mdlinked from README, and a "Does loadout phone home?" FAQentry — that question will appear within a day of release
Open question
The exact wire format for POSTing a Faro payload without the Web SDK isn't
fully covered in the public docs. The intent is to emit it directly, as the
current implementation does for Sentry — it keeps zero runtime dependencies,
avoids the question of whether the SDK behaves under
bun build --compile, andpreserves the property that every transmitted field is spelled out in our own
source, which is what makes the privacy doc checkable. Needs confirming against
the
faro-web-sdkfetch transport.