See where strangers get stuck. Walkthru is the launch check for apps built with AI. AI test users try your real flows (signup, onboarding, dashboard, checkout up to payment) inside your own Chrome, think aloud, and report where they got stuck. The same report covers SEO, AI search readiness (GEO), passive security hygiene, accessibility and mobile performance, with a grounded fix prompt your coding agent can act on.
Built for solo founders, small startups and agencies shipping with Cursor, Lovable, Bolt, v0 or Claude Code.
- Start here: handoff.md, AGENTS.md, CURRENT_STATE.md
- Product spec: SPEC.md · Architecture: ARCHITECTURE.md · Design: DESIGN.md · Roadmap: ROADMAP.md
- Deep dives: system design · agent safety plan · team collaboration · auth sessions · run idempotency · provider resilience · checkpoints · deploy · billing
- What one report answers
- Product tour
- Features in depth
- Architecture
- The LangGraph agent
- Remote MCP server
- Team collaboration
- Security hardening
- Privacy and data handling
- Reliability engineering
- Design and brand
- Repository layout
- Run it locally
- Testing and checks
- Plans
- License
| Question | How Walkthru answers it |
|---|---|
| Can a stranger get in? | A persona agent (first-time visitor, phone user, buyer, skeptic, or your own custom test users on Plus) drives the page in your browser one step at a time. Logged-in pages work without sharing a password. Every finding cites a step, a screenshot or a snippet. |
| Can Google and AI search read you? | Deterministic SEO and GEO audits: crawlability, structured data, AI crawler access, content before JavaScript, llms.txt, plus a copy-paste fix pack (robots rules, JSON-LD, rendering fix for your framework). |
| Are you leaking anything? | Passive checks only: headers and CSP quality, cookies, TLS and CAA, exposed files and source maps, keys shipped to the browser, vulnerable CDN libraries, subdomain takeover, and bounded read-only Supabase and Firebase probes on owner-verified domains. Never an attack. |
On top of the report: a Launch Ready score with a live badge, automatic rerun comparison (fixed, still broken, new), competitor side by side, weekly watch with deploy hooks, AI citation tracking on free engines, team workspaces with a findings board and an in-chat assistant, and a remote MCP server so Claude Code and Cursor can scan, read fixes, open pull requests and rerun from the editor.
Screens below are real captures of the web app (headless Edge at 1.5x, reduced motion) and real public reports from the local fixture sites. The design mockups are in Design and brand.
Paste an address for a free Instant Scan, give a test user a goal, watch it drive your real tab, then fix and rerun.
Every stuck point has a step, the test user's own words and a screenshot. One Launch Ready score, weighted only by what was actually measured.
| Evidence report (real run) | Extension side panel |
|---|---|
![]() |
![]() |
| Security hygiene (real run) | SEO and AI search readiness (real run) |
![]() |
![]() |
Where a test user got stuck, with the exact step and evidence:
The signed-in workspace: runs, team, compare, watch, AI answers, MCP, plan and billing.
| Pricing | Every plan, line by line |
|---|---|
![]() |
![]() |
| Docs: connect your editor (MCP) | Docs: team workspaces |
![]() |
![]() |
| Security page: where test users may act | Sign in (Google, GitHub, magic link) |
![]() |
![]() |
Scout, the walking bird, is the mark, the in-product assistant and the test user you watch in the side panel. The Agent Lab page (/agent-lab) is the identity playground.
- Built-in personas.
first_timer(skims, impatient),phone_user(taps big obvious things, hates long forms),buyer(pricing, trust, what happens next),skeptic(reads errors, checks links, distrusts vague copy). Defined inapp/agent/persona.py. - Custom test users (Plus). Up to 10 saved test users written in your own words, and up to 6 journeys on the same page in one test set (
app/plus.py), so one report compares how different people fail. - Goal planner. One model call turns the owner's typed goal into 1 to 4 ordered checkpoints, fixes typos, reads intent instead of literal words, and refuses goals that pay, delete, spam, attack or target another site. Code ends the run as soon as the last checkpoint is reached, so "click next project" stops after one click instead of circling.
- Think aloud. Every step records the persona's thought, the action, the target element's label (never the typed value), a confusion score 0 to 3, and what the step led to (
result_url, visible errors, confirmations, "nothing visible changed"). - Test identity. Forms get a deterministic throwaway identity (
walkthru.tester+<run>@example.com), never your data. - Evidence. Up to 8 masked JPEG screenshots per run, captured on the first step, on confusion, on errors and on clicks and typing, uploaded straight to private Storage under the user's own JWT.
- Logged-in pages. The test runs in your real tab with your existing session. No password is shared and cookies never leave the browser. Logged-in testing requires a verified domain.
- Human take-over. CAPTCHAs and bot walls are never solved. You solve them yourself and press Continue, or the run stops and the report says it was the site's bot protection, not a usability problem.
| Area | Module | What it measures |
|---|---|---|
| Crawl | scans/site.py |
One bounded same-site crawl (10 pages free, 50 paid), robots.txt honoured, shared by SEO, GEO and security |
| SEO | seo.py, seo_depth.py |
Titles, descriptions, canonicals, duplicates, broken links, sitemaps, Open Graph, structured data |
| AI search (GEO) | geo.py, geo_depth.py, geo_fixes.py |
AI crawler access in robots.txt, content present before JavaScript, llms.txt, schema, citability, plus a generated fix pack per framework |
| Security hygiene | security.py, csp.py, tls.py |
Headers and CSP quality, cookies, https redirects, TLS versions from one normal handshake plus one legacy offer, CAA over DNS, CORS answer to one Origin request |
| Leaks (verified only) | secrets.py, libraries.py, takeover.py, backend.py |
Keys shipped in JS bundles (gitleaks rules), public source maps, exposed well-known files, vulnerable CDN libraries (retire.js data), dangling subdomains, bounded read-only Supabase and Firebase probes |
email.py |
SPF, DKIM selectors, DMARC policy | |
| Accessibility | accessibility.py + extension |
axe-core WCAG A/AA in the real tab, server basics for Instant Scans |
| Performance | performance.py + extension |
PageSpeed Insights (mobile) plus Core Web Vitals from the tab |
| Stack | stack.py |
Framework and backend fingerprint, so fixes match your framework |
Plain code over the stored report (app/agent/score.py), never a model:
weights: ux 30 · security 20 · geo 20 · seo 15 · speed 15
penalty: high 25 · medium 10 · low 3 (per finding, floored at 0)
journey: done, safe_stop = 100 · gave_up, stuck, budget = 50 · Walkthru's own stops = not measured
An area that was not measured is left out and its weight goes to the others; it is never scored 0. Ignored findings still count toward the score, so the public badge cannot be gamed by hiding problems. agent_ready separately answers whether an AI agent can use the site.
Each finding has a stable fingerprint kind:rule (legacy reports fall back to a normalised title key with digits folded). A rerun of the same goal on the same site marks every finding fixed, still broken or new. accept_finding records a deliberate won't-fix with a reason; it leaves the lists but still counts in the score. verify_finding re-checks just one finding on its pages, faster than a full rerun.
- Fix prompt (
fix_prompt.py): one Markdown prompt for Cursor, Claude Code, Codex, Lovable or Bolt, built in code from the stored report andrecipes.jsonwith no model call, so it can only contain what the report found. Each finding gets why it matters, the change for the detected stack (vercel.json,netlify.toml,next.config.js,public/_headers), the risk and a local check, in batches that each end with "stop and verify". Things outside the code (DNS, hosting, email, provider dashboards) are listed separately; ignored findings are left out and secret values are masked. - Fix pull requests (
app/github.py): the Walkthru GitHub App is installed on the repositories you choose. Only the installation id is stored. Connecting proves ownership with the user-to-server OAuth code from the same redirect, then revokes that token at once. Each PR uses a token minted for one action and one repository. Changes come only from deterministic recipes (security headers invercel.json,netlify.tomlorpublic/_headers, newllms.txtorai.txt); files are created or appended, never rewritten; secrets and application code are never touched. Walkthru never merges.
Compare your site with up to three competitors on the same server-side checks, as a durable background job (compare.py).
app/citations.py asks AI engines your prompts and measures in code whether the answer names your site, cites it, and your share of voice against named competitors. Engines on free quotas: Groq gpt-oss-120b with its built-in browser search (web citations matched to answer references), and Gemini without search (memory: mentions only, labelled as such). Nothing promises rankings; it reports what the engines said on the day. Data model adapted from ai-search-guru/getcito (MIT).
Watched sites are re-scanned weekly and on a deploy hook URL your CI calls (at most one check every 10 minutes). The owner is emailed only when findings change. Hook URLs are shown once and stored as SHA-256.
See Team collaboration and Remote MCP server.
Founder-approved 30-day passes through Dodo Payments with signed webhooks, payable without a credit card (docs/billing.md). The server owns entitlements: the client never says which plan it is on.
- Browser on the user's machine, brain on the server. No Chromium on our servers, so hosting stays near zero and logged-in pages work without sharing passwords.
- One HTTP call per agent step. LangGraph
interrupt()emits the next action andCommand(resume=observation)continues. A Postgres checkpointer holds state, so the API is stateless between calls and can scale out. Start, observe and stop accept anIdempotency-Key. - Deterministic code before a model call. Accessibility, performance, SEO, GEO and security are plain code. Models explain and prioritise evidence; only the persona session is an agent. Scores, comparisons and fix prompts are computed in code and cannot invent findings.
- Prompts are never a safety boundary. Every rule about where the agent may act is code that runs on the server before the run and again at every step, so neither the model nor a modified extension can talk its way past it.
- Free tier models first. Groq
gpt-oss-120b, then OpenRouter and Gemini fallbacks, with a shared circuit breaker. Paid tiers put Claude Haiku 4.5 on Vertex AI in front when configured, with the free chain behind it. Every report states which model wrote it. - The server owns the plan.
POST /runsclamps steps, sites, test users and logged-in access from the caller's entitlement. - Reads through RLS, writes through the API. The web app reads its own rows straight from Supabase under row-level security; every write and business rule goes through FastAPI.
| Part | What it does | Where |
|---|---|---|
| Web app | Landing, pricing, docs, login (Google, GitHub, magic link), dashboard, private and public reports, compare, watch, AI answers, teams, settings with API keys, billing | apps/web (React 19, Vite, TypeScript, Tailwind v4, GSAP) |
| Chrome extension | Side panel to start a run, page script that snapshots the DOM (open shadow roots included), masks PII, runs axe WCAG A/AA and Core Web Vitals, executes actions with safety gates, and captures up to eight masked screenshots | apps/extension (Chrome MV3, WXT, React) |
| API | Run lifecycle, persona graph, report graph, scans, entitlements, billing webhooks, teams, watch, citations, MCP | apps/api/app (Python 3.12, FastAPI, LangGraph, httpx) |
| Persona agent | LangGraph graph with step budgets, loop guards, safe stops on Pay and Delete, optional TypeSafe Jev provider for bounded decisions | apps/api/app/agent/ |
| Scanners | site, seo, seo_depth, geo, geo_depth, geo_fixes, security, csp, tls, secrets, libraries, takeover, backend, email, accessibility, performance, stack, agent_signals |
apps/api/app/scans/ |
| SSRF-safe fetcher | Resolves each host once, refuses private and metadata addresses (including IPv4 hidden in IPv6), pins the connection to the checked address, re-checks every redirect, and allows at most two concurrent requests per host | app/scans/fetch.py |
| Background jobs | Durable Postgres queue: claim_job with FOR UPDATE SKIP LOCKED, 15-minute leases, per-kind attempts, exponential backoff with jitter, periodic retention and watch scheduling. Survives restarts and runs across processes |
app/jobs.py |
| Rate limits and abuse | Shared Postgres limiter per address, account, route and MCP key; per-target journey caps, kill switches, run audit log and automatic suspension | app/limits.py, app/abuse.py |
| Migrations | Numbered SQL files, each applied in one transaction with a checksum; editing an applied file is refused | app/migrate.py, apps/api/migrations/ |
| MCP server | Remote MCP at /mcp (streamable HTTP, personal API key, Plus plan) with 24 tools |
app/mcp_server.py |
| Citations | AI answer tracking with evidence, on a dedicated queue | app/citations.py, app/citation_evidence.py |
| Teams | Workspaces with roles, invitations, shared reports, findings board, threads, activity and Scout in team chat | app/teams.py, app/scout.py |
| Billing | Founder-approved 30-day passes through Dodo Payments with signed webhooks | app/billing.py |
| Admin panel | Founder-only panel on 127.0.0.1:8020 with password and authenticator code. Never deployed |
apps/api/admin/ |
| Data | Supabase Postgres with RLS on every table, private run-evidence Storage bucket, LangGraph checkpoints in a private schema |
apps/api/migrations/ |
| Resend over REST (report ready, watch changes, team invites) | app/deliver.py |
|
| Evals and fixtures | Easy, hard and showcase fixture sites full of traps; LangSmith evals; a headless end-to-end driver for the built extension | evals/ |
Two graphs do the work. persona_session is the only true agent: a loop that drives your browser one step per HTTP call. The report graph is a fan-out of deterministic scans with two model calls at the edges.
The wiring, exactly as compiled in persona.py:
g = StateGraph(SessionState)
g.add_node("decide", decide) # the only node that calls a model
g.add_node("act", act) # interrupt(): hands the step to the HTTP caller, resumes with the observation
g.add_node("check", check) # pure code: stop rules
g.add_edge(START, "decide")
g.add_conditional_edges("decide", after_decide) # done / give_up -> check, anything else -> act
g.add_edge("act", "check")
g.add_conditional_edges("check", after_check) # running -> decide, else END
graph = g.compile(checkpointer=checkpointer) # thread_id = run_idState (SessionState): run id, site, goal, persona or custom persona prompt, verified, mode (owner or visitor, decided by the server), max_steps, the latest observation, the steps log, status, tokens, the goal plan and plan_done.
decide renders the persona, the goal checklist (with done and current markers), the step history (including where each step led) and a text snapshot of the page: numbered interactive elements with [filled] or [empty] field states, visible errors and confirmations, scroll position and up to 4,000 characters of page text. The model returns a PersonaStep through structured output. Then three code checks run before anything leaves the server:
- Unconfirmed done. If the model says "done" right after a click or typing that changed nothing, it is asked once more (
NOT_CONFIRMED). If it insists, the step becomesgive_upwith an honest note. A real run once declared a signup complete without submitting it; this is the fix. - Evidence-based progress.
plan_doneonly moves forward when the last action visibly worked (the URL changed or a confirmation appeared). When every checkpoint is complete, the run ends even if the model wants to keep going. _enforce(), the code safety gate, replaces unsafe or invalid choices instead of trusting the prompt:
| Model chose | _enforce() does |
|---|---|
| An element id that is not on the page | Scroll instead, confusion raised (counted toward agent_lost) |
| Typing into a non-search field on an unverified site | Stop by design (visitor_mode_limit): "everything up to here worked" |
| Like, follow, share, post, buy, add to cart on an unverified site | Stop by design (visitor_mode_limit) |
| Pay, delete, cancel, checkout, transfer | safe_stop (logged-in runs: blocked and given up) |
| Send or invite on an unverified site | safe_stop |
| A second send in the same run | Done: Walkthru never sends twice |
type with empty text |
Fills the test identity value for that field |
act calls interrupt(step). The HTTP request returns that step to the extension and the thread sleeps in the checkpointer. The next POST /runs/{id}/observe resumes it with Command(resume={observation, evidence}). Because the model runs only in decide, a resume never repeats a model call.
check is pure code. It ends the run on safe_stop, done or give_up (agent_lost when the last three actions picked missing elements), on a bot wall or CAPTCHA note, when the current URL satisfies the next checkpoint's url_contains, on the step budget (12 on Free, 30 on paid plans), on the same action three times in a row (stuck), or on arriving at the same page for the third time (looping). The report wording comes from one table of stop reasons, so Walkthru never blames the site for its own limits.
Model chain for decide (runtime.py): Claude Haiku 4.5 on Vertex AI first on Pro and Plus when configured, then Groq gpt-oss-120b, gpt-oss-20b and qwen3, then Gemini 3.5 Flash and 3.1 Flash-Lite. max_retries=0 everywhere so a 429 falls through in milliseconds instead of sleeping on retry-after; reasoning_effort=low keeps gpt-oss from long reasoning spirals. With PERSONA_DECISION_MODEL=jev, a TypeSafe Jev classifier picks from the observed elements for bounded decisions and abstains to the LLM chain when unsure.
Checkpointer. CHECKPOINTER=memory in development; production refuses to start without CHECKPOINTER=postgres. The PostgresSaver lives in a private walkthru_checkpoints schema that is never exposed through the REST API, uses a bounded pool (4 connections, statement and lock timeouts, keepalives, health checks), refuses the Supabase transaction pooler, and verifies at startup that the numbered migration matches the pinned library's schema version.
POST /runs/{id}/observe passes through, in order: Supabase JWT validation (cached by SHA-256 of the token for at most 300 seconds, never past expiry), the per-run observe rate limit (60 a minute), the idempotency claim, kill switches, the blocked-host check on the tab's current URL, the signed-in-but-unverified check, and only then graph.invoke(Command(resume=...)). No database transaction stays open while the model runs.
for n in ("first_impression", "accessibility_scan", "site_scan", "stack_scan"):
g.add_edge(START, n) # fan out in parallel
g.add_edge("site_scan", "performance_scan") # performance waits for the crawled pages
g.add_edge(["first_impression", "accessibility_scan", "performance_scan", "stack_scan"], "synthesize")
g.add_edge("synthesize", END)Only first_impression and synthesize call a model. synthesize uses the careful-writer chain (Groq, then OpenRouter Nemotron, then the rest). After it, code takes over: grounded_ux() drops any UX finding that does not point at a real step and its evidence, code findings are added verbatim, plain() strips markdown and em dashes, and the Launch Ready and Agent Ready scores are computed. A scan that errors marks its area "unavailable" and the report still ships.
app/mcp_server.py mounts a remote MCP server at /mcp on the same FastAPI app: streamable HTTP, stateless_http=True, json_response=True, one request per tool call.
Auth gate (ASGI Endpoint), before the MCP session manager sees the request:
Authorization: Bearer wt_...or 401. Keys arewt_+ 32 bytes ofsecrets.token_urlsafe, shown once, at most 5 per account.- The key's SHA-256 is looked up; the raw key is never stored. Unknown or revoked: 401.
require_plus(owner): the server reads the plan; anything but Plus is 402.limits.apply("mcp", key_id): 60 calls a minute per key, in Postgres, so it holds across processes. MCP scans have their own cap of 50 a day per Plus user and never use the free capacity.- Last use is recorded and the user is bound to the request's log context.
- The user id is set in a
contextvarfor the duration of the request; every tool reads it, and every run id is checked against it ("No run with that id in your account").
Only ToolError messages reach the agent; other exceptions stay hidden. The server instructions tell the agent that findings are grounded in evidence and must never be invented.
The 24 tools
| Group | Tools |
|---|---|
| Reports | scan_site, get_report, list_runs, rerun, compare_sites, share_report (public link + Launch Ready badge), get_plan |
| Findings | get_finding, verify_finding, accept_finding, reopen_finding, get_fix_prompt |
| Ownership | get_site_verification (meta tag for your <head> with framework-specific placement, then confirms it once deployed) |
| GitHub | list_github_repos, open_fix_pull_request (preview first, confirm=true opens it) |
| AI answers | list_ai_answer_sites, track_ai_answers, get_ai_answers, set_ai_prompts, check_ai_answers_now |
| Weekly watch | list_watched_sites, watch_site, check_watched_site_now, create_deploy_hook |
User journeys still run only from the Chrome extension; rerun repeats the server-side checks. Connect from Claude Code:
claude mcp add --transport http walkthru http://localhost:8010/mcp --header "Authorization: Bearer YOUR_KEY"Cursor and other clients take the same URL and header.
A Plus owner creates a workspace for their team or a client: shared reports (manual or auto-share), a findings board with one row per problem and site (open, in progress, fixed, won't fix, with an owner), a persistent chat with @mentions and unread counts, comment threads on every report and finding, an activity log that doubles as the audit trail, and presence. Members do not need a plan of their own. Full reference: docs/team-collaboration.md.
How it is built
- All writes and business rules go through the API. Membership and role are read from the database on every request, never cached and never taken from the client. A non-member gets 404, so a workspace id reveals nothing.
- Row-level security is the second wall. Browsers get
selectonly, on rows of workspaces they belong to, throughSECURITY DEFINERfunctionsis_team_member(team)andshared_with_me(run)that answer only forauth.uid(). Invitations are never readable from a browser. No team table grants insert, update or delete to a browser role. - Shared reports reuse the existing report view. Policies on
runsandstorage.objectslet a member read a shared run and its screenshots straight from Supabase. - Realtime is a nudge, not a data path. A
postgres_changesevent (members only, under RLS) makes the page refetch through the API. Without a socket, chat polls every 5 seconds and pages every 30. - Concurrency. Seats are counted inside one locked transaction (
public.team_join), so two people can never take the last seat. Ownership moves in one transaction (public.team_transfer).(team_id, client_id)is unique on messages, so a retried send lands once. - Invitations. Codes carry 120 random bits, are shown once and stored as SHA-256. Email invites are single use and only work for the verified address they were sent to. Invite links can be limited to an email domain, a number of uses and an end date, and never grant admin.
- Input hygiene. Control characters and bidirectional overrides are stripped from names and messages so they cannot disguise text. Invitee emails are masked in the activity feed because viewers read it.
- Plan lapse. If the owner's Plus plan ends, the workspace becomes read-only (402 on chat, shares, triage and invites) until they renew or hand it to a member on Plus. Leaving, removing people and deleting always work.
- "Found again after a fix." Findings are keyed by site origin plus the stable fingerprint, so a finding marked fixed that reappears in a later shared report is flagged.
| Action | Viewer | Member | Admin | Owner |
|---|---|---|---|---|
| Read reports, findings, chat, activity | yes | yes | yes | yes |
| Chat and comment | yes | yes | yes | yes |
| Share own reports, triage findings | no | yes | yes | yes |
| Invite, change roles, remove people, rename | no | no | yes | yes |
| Hand over, delete the workspace | no | no | no | yes |
Limits: 3 seats per workspace, 3 owned workspaces, 20 memberships, 5 open invite links, 30 invitations a day, 30 messages a minute.
Scout in team chat. @Scout <question> answers from the workspace's shared reports, findings board and recent chat (app/scout.py). Retrieval is plain code with no embeddings: the question's words and any site it names pick the findings and reports that go into one prompt (14,000 characters at most). One Gemini Flash-Lite call on its own key and quota, so chat never starves the journey models; 50 answers per workspace per UTC day, 45-second budget. Scout reads only what the members can already read, takes no actions and has no tools.
Written after a test run where the agent was told to like five posts on Instagram with the founder's logged-in account. The full threat model and launch gates are in docs/agent-safety-plan.md.
Where the agent may act (policy.py)
Decided by the server before a run starts and again at every step. The model cannot change it, and neither can a modified extension, because every step goes through the API.
| Mode | When | What the test user may do |
|---|---|---|
| Visitor | Any site you have not verified (default) | Read, scroll, click links, menus and tabs, use the site search. No other form fields, no signing in, no social or commerce actions. A page showing a signed-in account ends the run. |
| Owner | A domain you proved with <meta name="walkthru-verification">, /.well-known/walkthru.txt or a _walkthru.<host> DNS TXT record, re-checked at every run and never cached |
Everything a real visitor does; safe mode still blocks payments, deletions and cancellations, and a real send needs your confirmation in the side panel |
| Refused | Banking, payments, trading, crypto, government, tax, healthcare portals, webmail, social networks, dating, adult, gambling, password managers and sign-in providers (blocklist.json) |
Nothing, unless you verified that exact domain |
The verification token is an HMAC of the account id, so it proves one account controls the domain. Goal rules run before the planner's model call, on text that is NFKC-normalised with zero-width characters removed (so like or full-width letters read as plain words): social and commerce verbs are owner-only, and bulk goals ("like every post", "create 20 accounts", "dozens of messages") are refused everywhere. The planner's own refusal is a second layer, not the first.
Abuse controls (app/abuse.py), all counted in Postgres so they hold across processes: 20 journeys an hour against one unverified host across all users, 12 an hour per user and host, account and host kill switches re-read mid-run, an automatic pause after 5 refused goals or sites in a day, and a per-run audit log of actions (never typed values) kept for 90 days.
- Host access is requested per site when a test starts (
optional_host_permissions), never granted up front. Permissions:activeTab,sidePanel,scripting,tabs,storage. - PII is masked inside the page before upload (
lib/redact.ts): emails, any 8+ digit number (cards, phones, account numbers, IBAN tails) and key-like strings (sk_,pk_,ghp_,xox*,AKIA...). Typed values never leave the tab; screenshots hide form fields. - It never reads cookies, local or session storage, saved passwords, history, other tabs or network requests. An automated test fails the build if its code ever touches those.
externally_connectableallows only the web origin to hand it a session. The handoff is a 30-second one-use challenge bound to the user and Auth session; both Chrome's sender origin and URL origin must match; the background worker serialises credential changes so a late refresh can never restore a signed-out session (docs/auth-sessions.md).safety.tsmirrors the server's social and commerce verb lists so the side panel stops early, but the server is the authority.
app/scans/fetch.py resolves each host once, requires every resolved address to be globally routable, and pins the connection to the checked address (no DNS rebinding window):
def is_public_ip(address: str) -> bool:
ip = ipaddress.ip_address(address.split("%", 1)[0]) # drop an IPv6 zone id
if ip in METADATA or ip.is_multicast: # 169.254.169.254, fd00:ec2::254, 100.100.100.200, ...
return False
if ip.version == 6: # IPv4 hidden in IPv6 must be public too
inner = ip.ipv4_mapped or ip.sixtofour or (ip.teredo[1] if ip.teredo else None)
if ip in NAT64:
inner = ipaddress.IPv4Address(int(ip) & 0xFFFFFFFF)
if inner is not None and not is_public_ip(str(inner)):
return False
return ip.is_globalEvery redirect (at most 5) is re-checked before it is followed. Bodies are capped at 300 KB, timeouts are short, at most two concurrent requests go to one host and at most 30 scans an hour to one target. The scanner identifies itself as WalkthruBot with a page that explains how to opt out. Every request is an ordinary read: no attack payloads, no fuzzing, no writes. Exposed files, keys in JavaScript, source maps and subdomain takeover checks run only on verified domains.
- Auth. Supabase validates every bearer token; the API caches a validated token by its SHA-256 for at most 300 seconds and never past its own expiry, and refuses a known-expired token before consulting the cache.
- Idempotency (docs/run-idempotency.md).
POST /runs,/observeand/stoptake anIdempotency-Key. Claims are durable, unique per user and key, and bind the operation plus a canonical body hash: a different body for the same key is 409, an in-flight duplicate is 409 withRetry-After, and a completed claim replays its original response for 24 hours. There is no timeout-based takeover, so an operation with an unknown outcome is never silently repeated. The extension retries only the HTTP request carrying an already captured observation, never a browser action. - Rate limits (
app/limits.py): one Postgres fixed window per key. Public 120/min per address, account 240/min, auth 60/min, start run 10/min, observe 60/min per run, costly routes 10/min, MCP 60/min per key./healthis exempt. The limiter fails open if the database is unreachable so an outage does not become a lockout. - Secrets at rest. MCP keys, deploy hooks and invite codes are stored only as SHA-256 fingerprints and shown once. GitHub tokens are minted per action and never stored.
- Web app headers.
Content-Security-Policy: default-src 'self'; script-src 'self'; frame-ancestors 'none'; base-uri 'self'; form-action 'self'; object-src 'none', plusnosniff, a strict referrer policy and aPermissions-Policythat turns off camera, microphone and geolocation. CORS is locked to the web origin and the extension. - Admin. The founder panel binds to
127.0.0.1:8020only, with a password hash and a TOTP authenticator code, and is never deployed.
requirements.lock pins every tested Python dependency. CI (.github/workflows/security.yml) runs npm audit, pip-audit against the lock file and gitleaks on every push. Reused open source is licence-checked first (no AGPL, Commons Clause, Elastic or unlicensed code) and recorded in apps/api/THIRD_PARTY.md. Migrations are checksummed; editing an applied one is refused.
The in-app Privacy, Terms and Security pages are the source of truth; this is the engineering summary.
What models receive. The masked text outline of each page, the goal and earlier steps. Never your password, cookies or the values typed into forms. Runs are not used to train models; providers process requests under their API terms.
Processors. Supabase (sign-in, database, private screenshot storage), Groq and Google Gemini (test users, report writing, Scout), OpenRouter (report writing), Google PageSpeed Insights (mobile speed of the public address), Resend (email you ask for), Dodo Payments (billing), and Google Cloud Vertex AI for Claude on paid tiers when enabled.
Who can see a report. Only you, until you create a public link or share it into a team workspace. Instant Scan reports have a public link by design. Logging out does not revoke public links or workspace access; those are separate, deliberate grants.
| Data | Kept for |
|---|---|
| Screenshots | 30 days after the run (EVIDENCE_RETENTION_DAYS), then purged: Storage objects first, rows last |
| Runs, steps, reports | Until you delete them |
| Run safety audit log | 90 days |
| Idempotent replies | 24 hours, erased by the six-hourly retention pass; minimal key and hash tombstones remain until account deletion so old retries cannot recreate a deleted run |
| API keys, deploy hooks, watched sites | Until revoked or removed |
| Team chat and activity | Until the author deletes a message or the owner deletes the workspace; a deleted account's messages are signed "Former member" |
| Instant Scan report emails | Not stored in our database |
Your rights. From Settings you can download everything held about you as one file and permanently delete your account; deletion removes stored files first, then run history, then the account, so nothing is left behind unreachable. Screenshots are served only through signed URLs that expire after one hour.
Acceptable use (Terms). Test only sites you own or are authorised to test, prefer staging, and never point Walkthru at someone else's accounts. Walkthru never submits payment forms, never solves CAPTCHAs, never sends attack payloads, and never keeps repository code after a scan.
- Durable job queue (
app/jobs.py). Reports, comparisons, watch checks, Scout replies and notifications run from a Postgres table.claim_jobusesFOR UPDATE SKIP LOCKEDso two workers never take the same row; leases last 15 minutes; each kind has its own attempt budget; retries back off exponentially with jitter (15 to 30 seconds at first, capped at 30 to 60 minutes). Periodic work (retention every 6 hours, watch every 30 minutes) is queued once per window across all processes. Workers run in the API or as a separate process (python -m app.jobs). - Provider circuit breaker (docs/provider-resilience.md). Three consecutive failures open a provider's circuit for 60 seconds; during cooldown calls fail instantly and the next provider answers; one caller probes recovery; a generation counter stops a stale in-flight result from closing a newer circuit. Failed structured output counts as a failure and falls through.
| Path | Timeout | In-call retries |
|---|---|---|
| Groq structured models | 8 s | 0 |
| Gemini structured models | 20 s | 0 |
| OpenRouter writer | 90 s (5 s connect) | 0 |
| Claude on Vertex | 30 s | 0 |
| TypeSafe Jev | 10 s | 0 |
- Stateless API. Agent state lives in the Postgres checkpointer, limits and jobs in Postgres, so the API can run as several processes. Production fails fast at startup if durable persistence is missing.
- Fail soft, never fake. A failed scan marks its area unavailable instead of sinking the report; a failed audit write is logged and the journey goes on; a circuit-breaker state never manufactures a citation result.
The live product follows DESIGN.md: premium, light and minimal. Geist and Geist Mono only, thin headlines, one quiet accent (#4d7274), real photography, slow deliberate motion, and every animation behind prefers-reduced-motion. Scout, the walking bird, is the mark and the in-product assistant.
Exploratory mockups where the founder's original doodles lead the identity, one per surface of the customer journey. These are studies, not the shipped UI. Brief and per-screen notes: apps/web/design/doodle-exploration/DIRECTION.md.
Source sheets used to build them:
| Doodles | Illustrations | MCP and Claude |
|---|---|---|
![]() |
![]() |
![]() |
Ten posters in ten aspect ratios built from the same artwork and the site's Geist type. Gallery and sources: apps/web/design/doodle-posters/.
Three 60-second, 1920x1080 launch films made with Hyperframes, each with Scout as the AI test user. Click a frame to open the MP4. Sources, tools and re-render steps: apps/web/design/launch-videos/README.md.
The 30-second cut (launch-video-30s/), frame by frame:
The hand-drawn diagrams are generated from code: specs in docs/diagrams/diagrams.js, drawn by rough.js (MIT) in Excalidraw's Virgil font (OFL), both loaded from unpkg at render time. Re-render after editing a spec:
node docs/diagrams/render.cjsThe script needs Playwright with Microsoft Edge (set PLAYWRIGHT_PATH if playwright is not resolvable).
apps/web/ React 19 + Vite + TypeScript + Tailwind v4 + GSAP
src/pages/ One file per page (Landing, Login, Dashboard, Report, Compare, Watch, Visibility, Teams, Mcp, ...)
src/components/ Shared UI (Nav, Pricing, Faq, ScanForm, Footer, Logo, Loading orbs, team/*)
src/lib/ supabase.ts, auth.ts, runs.ts, teams.ts, motion.ts, session handoff to the extension
src/content.ts All marketing copy
public/assets/ Images and product screens
design/ mocks/, doodle-exploration/, doodle-posters/, launch-videos/, launch-video-30s/, source-images/
apps/api/ FastAPI + LangGraph
app/ main.py (routes), agent/, scans/, jobs.py, limits.py, abuse.py, idempotency.py, migrate.py, mcp_server.py, teams.py, scout.py, ...
migrations/ 0001_initial.sql, 0002_agent_checkpoints.sql, 0003_run_idempotency.sql
admin/ Local founder admin panel
scripts/ test_user.py, grant_plan.py, billing.py
tests/ pytest suite
requirements.lock Pinned, tested Python dependencies
apps/extension/ Chrome MV3 extension (WXT + React)
entrypoints/ background, inject (page script), sidepanel
lib/ snapshot, redact, execute, safety, evidence, diagnostics, api, session
tests/ vitest
evals/ fixtures/easy, fixtures/hard, showcase, traps.json, serve.py, e2e_extension.py
docs/ Deploy, billing, system design, agent safety, auth sessions, idempotency, provider resilience, ...
diagrams/ Hand-drawn README diagrams (specs, renderer, PNGs)
readme/ README screenshots
dev.py, dev.cmd One-command dev stack (development only)
- Python 3.12+
- Node.js 20.19+ (Vite 8 and WXT)
- Google Chrome (for the extension)
- A Supabase project (free tier is enough): URL, publishable key, secret key and a Postgres connection string
- A free Groq API key and a free Google AI Studio (Gemini) key
- Optional: PageSpeed API key, Resend key, LangSmith key, OpenRouter key, Dodo test-mode keys
git clone https://github.com/watermelon588/Walkthru.gitEvery variable is listed with comments in .env.example. Copy the API block into apps/api/.env and the VITE_* block into apps/web/.env. For the extension, copy apps/extension/.env.beta.example to apps/extension/.env. Never commit any .env file.
Minimum for a local run:
| File | Variables |
|---|---|
apps/api/.env |
GROQ_API_KEY, GOOGLE_API_KEY, SUPABASE_URL, SUPABASE_PUBLISHABLE_KEY, SUPABASE_SECRET_KEY, DATABASE_URL, WEB_URL=http://localhost:5173, ALLOW_LOCAL_SCANS=1 (to scan the local fixtures) |
apps/web/.env |
VITE_SUPABASE_URL, VITE_SUPABASE_ANON_KEY, VITE_API_URL=http://localhost:8010, VITE_EXTENSION_ID (after loading the extension) |
apps/extension/.env |
VITE_API_URL, VITE_WEB_URL, VITE_SUPABASE_URL, VITE_SUPABASE_ANON_KEY |
If your network cannot reach Supabase's direct IPv6 host, use the session pooler string (port 5432) for DATABASE_URL. Never use the transaction pooler (6543).
cd apps/api
python -m venv .venv
.venv\Scripts\python -m pip install -r requirements.lock
.venv\Scripts\python -m pip install -e . --no-deps
.venv\Scripts\python -m app.migrate # apply pending migrations (status, down also available)
.venv\Scripts\python scripts/test_user.py # optional: throwaway account and session for local testingOn macOS or Linux use .venv/bin/python instead of .venv\Scripts\python.
cd apps/web && npm installcd apps/extension && npm install && npm run buildOpen chrome://extensions, turn on Developer mode, click Load unpacked and pick apps/extension/.output/chrome-mv3. Copy the extension id into VITE_EXTENSION_ID in apps/web/.env. After every rebuild, click Reload on the Walkthru card.
From the repository root:
.\devThis builds the extension and starts the API (http://localhost:8010), the web app (http://localhost:5173) and the fixture sites (easy :8101, hard :8102, showcase :8103) with labelled output, and restarts the API when its code changes. Ctrl+C stops all of them. dev.py and dev.cmd are development conveniences and are removed before production.
cd apps/api && .venv/Scripts/python -m uvicorn app.main:app --host 127.0.0.1 --port 8010cd apps/web && npm run dev -- --host 127.0.0.1 --port 5173apps/api/.venv/Scripts/python evals/serve.py- Check http://127.0.0.1:8010/health.
- Sign in at http://localhost:5173/login (the dashboard hands your session to the extension).
- Open http://127.0.0.1:8101 (easy) or http://127.0.0.1:8102 (hard).
- Open the Walkthru side panel from the Chrome toolbar.
- Enter a goal such as
Sign up for an account, pick a test user and click Start test. - When it finishes, open the report from the side panel or dashboard and check the steps, screenshot evidence, Share, Email me, CSV export and Save PDF.
The showcase fixture (http://127.0.0.1:8103) is a well-built product site whose journeys hold deliberate traps: circular links, a cookie banner, a newsletter pop-up, a dead button, new-tab and external links, content hidden until scrolled, and Pay and Delete buttons Walkthru must never press.
# Drive the built extension's page script in headless Chrome through the real API and model, then print the report
apps\api\.venv\Scripts\python evals\e2e_extension.py https://your-site.vercel.app "Find the projects and a way to get in touch"
# Free accounts get 3 runs a month. Give a test account a dev pass to try paid plans (no payment involved)
apps\api\.venv\Scripts\python apps\api\scripts\grant_plan.py you@example.com pro # --revoke ends it
# Run background jobs in their own process (set JOB_WORKERS=0 in the API first)
cd apps/api; .venv\Scripts\python -m app.jobs
# Founder admin panel on http://127.0.0.1:8020 (this computer only)
cd apps/api; .venv\Scripts\python -m admin setup; .venv\Scripts\python -m adminTo test a real form send on the easy fixture as a verified owner, write your verification token to evals/.walkthru-token (git-ignored; see GET /verification) and run with AUTO_CONFIRM=1. Without it the run stops at the send button.
All of these must stay green before a push.
cd apps/api && .venv/Scripts/python -m pytest -q && .venv/Scripts/ruff check .cd apps/web && npm run build && npm run lintcd apps/extension && npm test && npx tsc --noEmit && npm run lintAt the last full run (2026-10-02): 711 API tests passed (plus real-Postgres RLS and PostgREST suites that run when a live stack is configured), 90 extension tests passed, and the web build and lint were clean. Notable suites: the persona graph against scripted models, every _enforce and policy rule, idempotency races, SSRF address forms, circuit-breaker concurrency, team permissions, and an SQL test proving private, public and team report reads under RLS. CI (.github/workflows/security.yml) runs npm audit, pip-audit against requirements.lock, and gitleaks on every push.
| Free | Launch Pack | Pro | Plus | |
|---|---|---|---|---|
| Price | $0 | $9 once | $19/mo ($15 founding) | $49/mo ($39 founding) |
| Test runs | 3 a month | 20 within 30 days | 40 a month | 150 a month |
| Steps per run | 12 | 30 | 30 | 30 |
| Pages the test user may enter | Public | Public and logged-in | Public and logged-in | Public and logged-in |
| Rerun compare, fix prompt, competitors | No | Yes | Yes | Yes |
| MCP server, weekly watch, team workspaces, custom test users | No | No | No | Yes |
Full table and what is still coming: SPEC.md.
Walkthru is released under the MIT License. Copyright (c) 2026 Rohit Maity.
Third-party code and assets keep their own licences: reused open source is listed in apps/api/THIRD_PARTY.md, the Geist fonts ship with their OFL licence files, and rough.js (MIT) and the Virgil font (OFL) are only loaded at diagram render time.













































