An agent that investigates open GitHub issues with tools, then rules on them.
An autonomous agent bot triggered by GitHub Actions' own event system (genuinely webhook-driven under the hood — GitHub Actions is built on GitHub's internal webhook/event delivery — but there is no custom HTTP webhook receiver endpoint in this repo to audit; see .github/workflows/triage.yml for the actual trigger), not a human-clicked UI demo — see Production usage for the run, issue, and comment it produced on its own, plus the measured accuracy benchmark in ACCURACY.md.
Live demo | Quick start | How it works | Docs
Open https://triage-desk-iota.vercel.app, enter any public owner/name repository (or click one of the suggested examples) and open the docket. Ruling needs your own Anthropic or OpenAI API key, entered in the app (see Privacy).
- Enter
owner/nameand press Open docket. - Select an issue and press
R, or pressBto triage the next five. - Copy the suggested reply with
C, or export all rulings as JSON.
- Investigates before ruling. The agent calls tools (get issue, find similar issues, list labels, read a repo file) before it decides.
- Structured ruling. Kind, priority, labels, duplicate, missing information, a reply you can paste, and confidence.
- Auditable. A timeline shows every tool call the agent made.
- Fast to drive. Keyboard shortcuts, batch triage of five issues, JSON export.
- Read-only in the browser. The browser app never writes to GitHub; a human applies the ruling. The CI bot below is the one exception, scoped to its own workflow token.
- Bring your own key. Anthropic or OpenAI, entered in the browser. No server, no environment variables.
- Runs unattended too. A GitHub Actions bot (
scripts/triage-bot.ts) triages every new issue automatically — see Production usage.
flowchart LR
I[Issue] --> A[Agent loop]
A -->|tool call| T[Tools]
T --> A
T --> G[get issue]
T --> D[find similar]
T --> L[list labels]
T --> F[read repo file]
A -->|submit| V[Validate ruling]
V --> R[Ruling view]
The page builds an adapter for the chosen provider and runs the loop with the open issues as context. Each tool call is reported as a step. The loop ends when submit_triage passes validation. Full diagrams and the module map are in docs/architecture.md.
Requires Node 22 or newer.
git clone https://github.com/edgeorgie/triage-desk.git
cd triage-desk
npm ci
npm run devOpen http://localhost:3000. There are no environment variables: click Add API key, choose Anthropic or OpenAI and paste your key. The key stays in your browser.
Sandboxed/CI environments with
NODE_ENV=productionset:npm ci/npm installwill silently skip devDependencies (includingtypescriptand@tailwindcss/postcss), causingnpm run typecheck/npm run buildto fail with missing-module errors that look like real bugs but aren't. Fix:unset NODE_ENV && npm install --include=devbefore running either command.
| Data | Where it goes | Stored |
|---|---|---|
| Repository issues | Fetched from GitHub in the browser | Memory only |
| Issue text and tool results | Sent to the chosen model provider | Not stored |
| Provider key | Sent only to the provider | sessionStorage by default; localStorage only if you choose to remember it on this device |
| Rulings | Kept in memory, exported on request | Not stored |
Issue text is untrusted input to an agent with tools. Blast radius is limited: tools are read-only, file paths are guarded, the loop is capped at 8 steps and the ruling is schema-validated. A hostile issue could still mislead the ruling, so a human applies it.
Beyond the browser demo, this repo ships a headless agent bot that runs autonomously — no human click, no browser, no manually-entered key for the GitHub actions it takes.
flowchart LR
W[GitHub webhook: issues.opened] --> GA[.github/workflows/triage.yml]
GA --> S[scripts/triage-bot.ts]
S -->|reuses| L[lib/agent.ts, lib/tools.ts, lib/triage.ts]
S -->|GITHUB_TOKEN| C[Issue comment]
S -->|GITHUB_TOKEN| LB[Issue labels]
- Trigger: GitHub Actions' own
issues: [opened]/pull_request_target: [opened]events — these are genuinely delivered by GitHub's internal webhook/event system, not a manual run, but note this repo has no custom HTTP webhook receiver; the trigger is entirely GitHub Actions' built-inon:config in .github/workflows/triage.yml, so there's no signature-verification code path to review here. - Logic reuse:
scripts/triage-bot.tsimports the exact samelib/agent.tsagent loop,lib/tools.tstools andlib/triage.tsschema the browser UI uses; it just swaps the browser'sfetch-based "paste your key" flow for a Node/CI environment. - Writes to GitHub autonomously: the workflow's built-in
GITHUB_TOKENis enough to post a triage comment and apply labels — no new secret required for that part. - LLM reasoning (optional): if
ANTHROPIC_API_KEYorOPENAI_API_KEYis present as a repository secret (Settings → Secrets and variables → Actions), the bot calls that provider for the kind/priority/duplicate/reply reasoning, same as the browser app. Onpull_request_targetevents (PRs can come from any anonymous fork) the workflow only forwards these paid keys whenauthor_associationisOWNER/MEMBER/COLLABORATOR— everyone else's PR still gets triaged, just via the free heuristic fallback below, so an anonymous account can't run up the repo owner's LLM bill. See the cost-gate comment in .github/workflows/triage.yml for the full reasoning. - Heuristic fallback (documented, no fabricated LLM usage): with neither secret configured,
scripts/triage-bot.tsruns a deterministic, keyword/overlap-based triage (seeheuristicTriagein the script) so the full pipeline — trigger → investigate → comment → label — still produces real output end-to-end, clearly labeled_Automated heuristic triage (no LLM key configured)_in the posted comment. - Evidence it runs:
- Example run (triggered by the real
issues.openedwebhook, 22s, green): https://github.com/edgeorgie/triage-desk/actions/runs/37983913473 - Example issue the bot triaged on its own: #26
- Example comment it posted: #26 (comment)
- Labels it applied autonomously:
feature,p3 - Each run also writes
triage-logs/runs.jsonlandtriage-logs/last-run.json, uploaded as a workflow artifact (triage-log-<run id>) on every run — a machine-readable record alongside the human-readable Actions log.
- Example run (triggered by the real
- To enable LLM-grade triage: add an
ANTHROPIC_API_KEY(orOPENAI_API_KEY) secret under Settings → Secrets and variables → Actions. No code changes needed; the bot detects it automatically and switches modes. - I built the webhook-triggered bot, see PR #25. I wired eval-lab into this repo's own CI to test the bot's real triage logic, see PR #27.
The accuracy/duplicate-detection numbers below are self-labeled internal consistency checks (n=18), not an external benchmark — I wrote the ground truth labels using the same mental model the heuristic implements, so they measure "does the code match my own intuition" more than "does it match independent human judgment." Treat them as a documented starting point with a known root-cause bug (see ACCURACY.md), not a validated accuracy claim. The webhook-latency and CI numbers below them are real production/operational measurements, not self-graded.
- 22s — webhook-to-comment latency, real run: triage-desk/actions/runs/37983913473
- 83.3% (15/18) — kind classification accuracy, self-labeled internal-consistency check, n=18 (ACCURACY.md)
- 83.3% (15/18) — priority classification accuracy, same self-labeled check
- 1/2 (50%) — duplicate detection rate on the same self-labeled check (near-verbatim caught, paraphrase missed; n=2 is too small to generalize)
- 0.274ms → ~0.03ms — heuristic triage latency per case, cold vs. JIT-warmed
- 6/6 — eval-lab eval cases passing against this bot's real heuristic logic in CI (run 37987251361)
- 22/22 — local unit tests passing (
npm test)
- Public repositories only; 60 unauthenticated GitHub calls per hour.
- Token overlap misses paraphrased duplicates.
- The agent never writes to GitHub.
Next.js 16 (App Router), React 19, TypeScript, Tailwind CSS 4. Client-side only; no backend. Models: Anthropic claude-haiku-4-5-20251001 or OpenAI gpt-4o-mini.
| Script | Purpose |
|---|---|
npm run dev |
Development server |
npm run build |
Production build |
npm run typecheck |
TypeScript check |
npm run lint |
ESLint |
npm test |
Unit tests |
npm run spec:check |
Traceability gate |
npm run verify |
All of the above |
npm run deploy:pages |
Static export to the gh-pages branch |
npm run triage-bot |
Headless triage for one issue (used by .github/workflows/triage.yml) |
The app is fully client-side, so it can be hosted as static files.
- GitHub Pages:
npm run deploy:pagesbuilds a static export and publishes it to thegh-pagesbranch. Enable Pages from that branch; on a free plan the repository must be public. - Vercel or any Node host: use the Deploy button above. No configuration is needed.
| Term | Meaning |
|---|---|
| Agent loop | Repeat: ask the model, run the tools it requests, feed back results, until it submits. |
| Tool use | The model returns structured calls to named functions instead of free text. |
| Terminal tool | submit_triage: the call that ends the loop and carries the structured ruling. |
| Ruling | The validated decision: kind, priority, labels, duplicate, missing information, reply, confidence. |
| Duplicate detection | Ranking similar open issues by token overlap computed locally. |
| Step budget | The loop stops after 8 steps if no ruling is submitted. |
Typography: Display, Syne; Text, Onest; Code, JetBrains Mono.
| Token | Value | Use |
|---|---|---|
cream |
#fff9e8 |
Page background |
ink |
#17130a |
Text and ruling card |
tangerine |
#ff6a2b |
Primary action |
lemon |
#ffe45e |
Highlights |
mint |
#19b47a |
Positive |
rose |
#ef4565 |
Errors and bugs |
- Make the agent's work visible and auditable.
- Chunky borders and offset shadows for a tactile feel.
Motion, components and rationale: docs/design-system.md.
| Document | What it answers |
|---|---|
| docs/index.md | Map of all documentation |
| docs/architecture.md | Diagrams and modules |
| docs/spec/spec.md | Requirements and acceptance criteria |
| docs/spec/traceability.md | Requirement to code, test and evidence |
| docs/design-system.md | Tokens, motion, components |
| docs/glossary.md | Definitions |
| docs/evaluation.md | Self-assessment against a review rubric |
| docs/adr | Decision records |
- AGENTS.md defines the workflow and quality gates for agents and people.
- llms.txt is served at
/llms.txtwhen deployed and points to the key documents. - docs/spec/requirements.json is the machine-readable requirement list with status, files and tests.
npm run verifyis the single deterministic gate: typecheck, lint, traceability check, tests and build.
See CONTRIBUTING.md. Security reports: SECURITY.md.