Skip to content

Latest commit

 

History

History
427 lines (341 loc) · 19.2 KB

File metadata and controls

427 lines (341 loc) · 19.2 KB

Backlog

Future-feature notes that are scoped but not yet implemented. Promote an entry to a workstream when it's time to build; remove when shipped.


Proactive component health watchdog (v0.5.1 candidate)

Phase 5's _enter_unavailable is purely reactive: the booth only notices a hardware failure when the next operation against that hardware is attempted. On the live booth this is most visible with the camera — power it off mid-series, between shots, and nothing happens until the user presses blue, at which point the countdown plays, the capture fails, and only then does the unavailable screen appear.

Why pick it up

The current behavior is technically safe (no data loss) and the design intent ("wrap each hardware op in try/except") is sound, but the user expectation is "if the camera vanishes, tell me right away." The gap is most painful at events where the operator isn't standing at the booth to notice a tripped USB cable until a guest is mid-press.

Sketch

Two related changes:

  1. HealthMonitor.recheck_loop probes ready components too, not just ones already in unavailable. The probe is cheap (camera is ~0.5 s of gphoto2 --summary, printer is a USB enumerate, net is the existing 1 s connection check). A ready → unavailable transition publishes a StateChange to state_changes exactly the way the existing path does.
  2. booth_main subscribes to state_changes in a background task and, when a required component transitions to unavailable, interrupts the current main-loop state and routes through _enter_unavailable. The interrupt mechanism is the hard part — asyncio.CancelledError against whatever the main loop is awaiting (a button event, a screen-hold timeout, an in-flight upload task) plus a state-restoration path on recovery so the booth returns to the right screen.

Open design questions

  • Probe cadence for the camera. --summary against an idle camera is fine at 10 s intervals; cranking it lower might interfere with actual captures (gphoto2 holds an exclusive USB lock).
  • Which states are interruptable. Attract, review, and series- capture are clearly safe. Mid-print is not — interrupting an ESC/POS receipt mid-flight would leave torn paper. Probably gate the interrupt on _print_with_recovery not being on the stack.
  • Test surface. The reactive path is well-covered by test_unavailable_mode.py / test_resume_mid_series.py; the proactive path needs at least: probe-flip-during-attract, probe-flip- during-final-screen, probe-flip-during-active-capture (should be swallowed because capture itself will surface it), probe-flip-during- print (should be blocked / queued until print completes).

Smoke test for booth auto-QA

After every code change to the photobooth runtime, a smoke test should be runnable on the booth (or against a booth running in test mode) to confirm the capture → review → upload → print flow behaves end-to-end and that nothing sensitive is leaking into logs or URLs.

Behavior

  • Stage-by-stage trigger: with the booth process running, the smoke test can advance through each stage of the capture flow (idle → capture start → shot N → review → keep/redo → final → print/timeout → back to idle, including series-mode variants).
  • Baked-in sleeps: each stage pauses long enough to simulate the user loading / decision delay the booth normally sees (countdown, camera capture, compositing, upload, print) so timing-dependent code paths (e.g. event queue flushing, final-screen timeout, kiosk navigation) are exercised realistically.
  • Log and URL inspection:
    • Read /var/log/photobooth.log after each stage and verify the expected semantic events fired (see booth_main.py event mapping).
    • Capture every URL the run emits (presigned or plain) and every QR payload sent to the receipt printer.
    • Pattern-match URLs and log lines for credential material:
      • AKIA[A-Z0-9]{16} (AWS Access Key ID)
      • X-Amz-Signature=, X-Amz-Credential=, X-Amz-Security-Token= (presigned URL params)
      • Any string starting with AWS4-HMAC-SHA256
      • Long base64-ish opaque strings inside query parameters
    • Any match emits a CRITICAL log entry naming the offending line and the matched pattern. CRITICAL routes through the existing photobooth.logging_config setup (file + journald), so a leak shows up in normal operator monitoring too.

Open design questions when picking this up

  • Triggering mechanism: a BOOTH_SMOKE=1 env var that wraps booth_main.main() with a stage-driver that injects fake button events into RPi.event_queue directly? Or a separate photobooth-smoke console-script entry point that runs the smoke loop in-process and bypasses GPIO? In-process is simpler; an out-of-process IPC harness is more realistic (exercises GPIO) but requires a control socket.
  • Hardware mocks: does the smoke run hit the real camera and printer or a mock layer? Mocks are necessary to run on CI; real hardware is necessary to catch hardware-edge-case regressions. Probably both modes, selected by env var.
  • Pass/fail surface: exit code? A summary JSON written next to the log? A journalctl -u booth-smoke stream?

Why

Workstream E gave us observable events but no automated way to verify they fire correctly after a change. Today the only verification is the human "press the blue button" loop we've been doing — slow, easy to forget edge cases (redo, max-prints, upload failure paths), and impossible to run unattended after a deploy.

The credential-material check exists because v0.4.0/F caught two presigned URLs in plain-text logs that would otherwise have shipped to customer-facing receipts. A smoke test that fails LOUD on key material prevents regression on that surface.


Spurious capture — EMI on GPIO signal line (root cause identified, hardware fix in progress)

History

This entry was originally filed as "Camera sleep triggers a spurious capture" — the symptom looked like the camera waking up was queuing a phantom capture. After adding the diagnostic logging table below and catching one reproduction in the wild (2026-05-23), the actual mechanism turned out to be two stacked failures that together produced the observed symptom.

What it actually was

  1. EMI-induced falling edge on the GPIO signal line. The original button wiring left the GPIO signal line idle LOW (pulled to GND via a 10 kΩ resistor) on a long unshielded cat5e run that shared a jacket with a 12 V LED-supply pair. Capacitive coupling from nearby switching loads (receipt printer inrush, neopixel PWM, USB to the camera) injected positive-going spikes onto the signal line. Each spike briefly pulled the line HIGH; as the spike dissipated through the 10 kΩ pull-down, the line settled back to LOW — producing a falling edge that looked indistinguishable from a real press to the ISR.
  2. Camera was asleep when the phantom capture fired. gphoto2's --capture-image-and-download exited successfully (returncode 0) without producing a file: size=-1, elapsed=0.13s in the logs. The booth then crashed when downstream code tried to open the nonexistent JPEG.

Each failure on its own would have been noticeable but recoverable; together they produced a complete process exit triggered by no human action, which is why it looked so mysterious.

Hardware fix (validated 2026-05-25 on the capture button)

Rewired GPIO 25 / pin 22 (capture button) to active-low topology:

  • Pull-up: 10 kΩ from SIG → 3.3V at the Pi end (was: pull-down from SIG → GND at the Pi end)
  • Button: shorts SIG → GND when pressed (was: shorts SIG → 3.3V)
  • Series protection: 1 kΩ between SIG and GPIO 25 (unchanged)
  • Low-pass: 100 nF (0.1 µF) from SIG → GND at the Pi end (new)
  • Cable: cat5e pair carries SIG + signal-GND on one twisted pair (was: SIG + 3.3V supply, which defeated twisted-pair noise rejection)

Why this works:

  • Idle state is now HIGH (3.3V). Positive-going EMI spikes can only push the line slightly higher, clipped by the Pi input clamp diodes; the spike's dissipation back to 3.3V produces no falling edge → no phantom ISR.
  • The 100 nF + 10 kΩ low-pass adds a ~1 ms filter cutoff. Real presses (≥ 50 ms) pass; sub-millisecond EMI is rejected regardless of polarity.
  • Twisted pair common-mode rejection now works because both wires of the pair carry a related signal pair (SIG + its return).

Validation: 3.5-hour idle test on the rewired capture button produced zero phantom triggers. Red and green buttons (still on original wiring) serve as the control group; if they continue to misbehave over longer windows, the diagnosis is conclusive.

Status (2026-05-25)

  • Capture button: rewired, validated, in production on the live booth. See rpi_provisioning/HARDWARE.md for the topology and the cat5e pair assignment.
  • Red and green buttons: still on the defective v0.2.0 wiring. See "Apply rewire to red and green buttons" below.
  • Software-side resilience layer: not yet shipped — even after the rewire, gphoto2 can still return success-with-no-file when the camera is asleep. See "File-size guard in Camera.capture_async" below for the catch-and-recover plan.
  • Release-bounce as a side effect: discovered during validation, documented separately below.

Diagnostic logging (kept; still useful for future GPIO triage)

Three log lines bracket every capture. Enable the high-frequency events with BOOTH_LOG_LEVEL=DEBUG in /etc/ctp/booth.env:

Layer Level Line What to look for
GPIO ISR (rpi._on_press) DEBUG Button event: label=X pin=N queue_depth_before=Q queue_depth_before > 0 on consecutive presses of the same button ⇒ ISR firing multiple times for one physical press (bounce/EMI). Phantom event with depth=0 = a single noise-induced edge made it through.
Main loop (booth_main.run) DEBUG Main loop event: X An entry here with no matching Button event ⇒ event-bridge bug (very unlikely). Pair timestamps to measure GPIO→dispatch latency.
Camera result (camera.capture_async) INFO Camera capture: model=... file=... size=N elapsed=Ts exif_dt=... size < ~100 KB or size = -1 ⇒ junk/empty frame. exif_dt significantly older than wall-clock ⇒ stale buffered frame (note: only useful if the camera body clock has been set).

When the next suspicious capture occurs, grep /var/log/photobooth.log:

grep -B 30 -A 5 "Capture started" /var/log/photobooth.log | less

Apply rewire to red and green buttons

Apply the same active-low circuit shipped on the capture button (see "Spurious capture — EMI on GPIO signal line") to GPIO 23 (green / pin 16) and GPIO 24 (red / pin 18). Same parts per button (10 kΩ + 1 kΩ + 100 nF), same cat5e pair assignment.

Until this is done, expect occasional phantom red/green events under EMI load. The booth handles them more gracefully than a phantom capture (red is just "back to attract", green is reprint or "keep" depending on context), but cosmetic at minimum.

Promote the wiring topology in rpi_provisioning/HARDWARE.md from "in progress" to "current" once all three buttons are on the new design.


Release-bounce on long-hold-release produces a spurious event

Symptom

With the new active-low wiring, holding a button for longer than the 500 ms GPIO bouncetime window and then releasing produces an extra Button event log line on release.

Observed 2026-05-25 during wiring validation: capture press at 13:59:21, held through the capture flow, released at ~13:59:30 — a second Button event: label=capture pin=25 queue_depth_before=0 fired at 13:59:30.

Mechanism

Mechanical switch contacts physically bounce when they OPEN, not just when they close. With the active-low circuit, releasing a held button ramps the line from 0V → 3.3V via the pull-up; contact bounce briefly re-makes the closed state, pulling the line back to 0V momentarily. Each re-make is a falling edge. The bouncetime=500 parameter only suppresses events within 500 ms of the previous registered event; if the original press was longer than 500 ms ago, the first bounce-induced falling edge gets through.

This did not occur with the old (idle-LOW) wiring because the ISR fired on the natural release transition (the only falling edge available); there was no "second falling edge" to bounce.

Impact

Cosmetic if it happens at the final-screen hold (gets flushed by _flush_events()), potentially functional if it happens mid-flow where a queued event could be picked up by _review_shot and interpreted as a decision the user didn't make. Has not been observed to cause a misbehavior in normal use yet — operators typically tap rather than hold.

Fix options

  • Software flush after _take_one_shot: drain the event queue immediately after each capture, before _review_shot polls. Cheap; consistent with the existing _flush_events() pattern.
  • Asyncio-side debounce in RPi.next_event: track the last event timestamp per label and discard duplicates within a configurable window (e.g. 1 s). More general but adds state.
  • Switch to RISING-edge detection (release-detect, like the defective old wiring): works but is unintuitive and re-introduces the noise susceptibility we just fixed.

The flush-after-capture is the lightest touch and matches the existing discipline. Slot when next worth touching booth_main.run.


File-size guard in Camera.capture_async

Why

Even with the GPIO line fixed, gphoto2 can still return success without producing a file when the camera is in an unusual state (asleep, USB hiccup, battery brown-out). Today the booth treats any non-raising return as a real shot and crashes downstream when something tries to open the nonexistent JPEG.

Behavior

In photobooth/camera.py:Camera.capture_async(), after the existing log line:

  • If not os.path.exists(pic) or os.path.getsize(pic) <= 0 or < ~100 KB: raise a new CaptureFailed("gphoto2 returned success but no file on disk") exception. Do not append to _captures; reset _ready = True so the next press can try again.

In photobooth/booth_main.py:_run_single / _run_series:

  • Wrap the await self._take_one_shot() call in try / except CaptureFailed.
  • On exception: cancel the attract task, scroll a short "Camera not ready, try again" message on the neopixel, navigate back to the attract URL, drop back into the main event loop.

Why this is independent of the wiring fix

The wiring fix prevents the phantom capture from being triggered in the first place. This guard prevents the booth from crashing when a capture is triggered (legitimately or otherwise) but the camera fails to produce a file. Both are wanted; they cover different failure modes.

Optional follow-up

If the camera-sleep root cause turns out to be the dominant one even with the rewire complete, the simplest fix is one line of additional camera startup config in booth_main.py:

CAMERA_STARTUP_CONFIG = {
    "autoexposuremode": 3,
    "autopoweroff": 0,  # never sleep
}

Trade-off: sensor stays warm during long idle days; potentially shortens shutter life on a multi-day event. Worth measuring before committing.


Canon Selphy CP1500 photo printer integration + Printer → ThermalPrinter rename

Why

The booth's only "printer" today is a PBM-8350U thermal device that prints a marketing receipt with a QR code linking to the S3-hosted photo — the customer never walks away with a paper print of the actual shot. Adding a Canon Selphy CP1500 (USB dye-sub, 4×6) makes the booth a real "leave with a print in your hand" experience without removing the receipt's marketing role.

PhotoStrip.expand_for_print() already exists for exactly this — it duplicates a 2×6 strip into a 4×6 sheet — but nothing in booth_main.py calls it today.

Behavior

  • Add PhotoPrinter alongside ThermalPrinter. Both run per capture. Receipt + photo are independent run_in_executor calls; a failure on one printer never blocks the other or the return to attract.
  • Rename Printer → ThermalPrinter at the same time. The existing class is opinionated to ESC/POS / thermal hardware (escpos.printer.Usb, text / cut / qr / barcode); the bare name became misleading once a second printer joined the design. Rename scope: class, module file (printer.py → thermal_printer.py), PRINTER_MAP → THERMAL_PRINTER_MAP, Booth.add_printer → Booth.add_thermal_printer, tests/test_printer.py → tests/test_thermal_printer.py, and all call sites in booth_main.py. No behavioral change.
  • CUPS via lp, not pycups. PhotoPrinter shells out to the system lp binary via subprocess.run. No C-extension build on Buster/Py3.7 (where rawpy already broke), and CUPS is already the integration surface either way.
  • Wire expand_for_print into the flow. Currently unused; called immediately after compose() when a PhotoPrinter is online. The single-strip output stays the source of truth for the S3 upload and the kiosk review screen.

Risks

  1. Buster Gutenprint is too old for the CP1500. Buster ships Gutenprint 5.3.3 (Sept 2018); the CP1500 hardware shipped Sept 2022. Verify before starting code by SSHing to the live booth and running lpinfo --make-and-model "Canon CP1500" -m | grep -i cp1500. If empty, build Solomon Peachy's standalone selphy_print from source (https://git.shaftnet.org/cgit/selphy_print.git/) — small C build, no Python deps, registers as a CUPS backend.
  2. Pi3-class CPU on the live booth. Dye-sub data streams are ~10–20 MB per 4×6; CUPS rasterization can take 5–15 s on older ARM. Mitigation: fire _do_photo_print early, in parallel with upload.
  3. expand_for_print has never run in production. If any active template has columns > 1 and the geometry was authored without testing the 4×6 output, the first print may have unexpected margins. Mitigation: dry-run with lp -o job-hold-until=indefinite during initial bring-up.
  4. Ribbon/paper exhaustion is silent in the UI. CP1500 has a fixed-count consumable. Out of scope for the initial PR; worth a follow-up that polls lpstat -W not-completed and lights a panel LED on stall.

Reference

Full plan including file-by-file change list lives at /home/ian/.claude/plans/how-easily-could-we-zazzy-whistle.md. The two-class rationale and the CUPS-via-lp decision are captured in ARCHITECTURE.md under "Printers".

Doc alignment (2026-06-20)

ARCHITECTURE.md describes the current runtime: a single Printer class in photobooth/printer.py (PRINTER_MAP, Booth.add_printer, _do_print, tests/test_printer.py). An earlier pass had pre-staged the post-rename world in the doc while this workstream was scoped, which left it describing code that didn't exist; the doc has since been walked back to the shipping state, and this BACKLOG entry is now the canonical record of the planned design.

When this workstream lands, update ARCHITECTURE.md to match: rename Printer → ThermalPrinter (thermal_printer.py, THERMAL_PRINTER_MAP, add_thermal_printer, test_thermal_printer.py) and add the PhotoPrinter / photo_printer.py / PHOTO_PRINTER_MAP sections plus the two-class "Printers" rewrite.