Skip to content
getjumpPublic

About

An Agent Skill that reviews a pull request by running it: three builds, one subtraction, then probes for what no test covers.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

pr-review

An Agent Skill that reviews a pull request by running it, not by reading it.

getjump.github.io/pr-review — the method, and a real report it produced.

A line-by-line diff answers a question nobody asked. This skill answers the one people do ask: what changes for whoever uses this, and what of it is noise — and it proves the answer instead of asserting it.

The method

Build three copies of the repository and run the test suite in each:

build code tests what it tells you
base before before what is already red — flakes and broken environments
mix after before anything failing here changed behavior
head after after whether the PR is even green

Then it is subtraction, not opinion:

  • failed in mix, not in base → a behavior change
  • did the PR edit that test or its fixtures? edited → the author knew. untouched but the result moved → an unstated change, and that is the finding worth leading with

Execution only covers scenarios somebody already wrote down, so the skill adds a second pass: probes. A probe is a throwaway test case written for measurement, run on both sides, aimed at whatever the diff narrowed — a value added to a blacklist, a condition replaced by a new lookup, a newly required input. Probes are where the regressions nobody tested for turn up.

The output is a single self-contained HTML page with four levels behind a detail slider: consequences → rules → examples → code, cross-filtered by change.

Install

A skill is just a directory with a SKILL.md. Put this one wherever your agent looks for skills — for Claude Code that is ~/.claude/skills/ (personal) or .claude/skills/ (per project):

git clone https://github.com/getjump/pr-review ~/.claude/skills/pr-review

Other agents that read the same format (Cursor, Codex, OpenCode, Gemini CLI, Goose, Copilot, Amp, Junie, and others) use their own directory — see the client list for each one's setup page.

Nothing here is tied to a particular agent or vendor: the skill body is plain Markdown, and the bundled scripts are bash and dependency-free Python 3.8+.

Use

Ask your agent for a PR review, a diff summary, or "what does this branch actually change" — the skill triggers on that. Or invoke it by name.

It needs git and the ability to run the project's test suite. It does not need network access beyond fetching the branch under review.

What is in the box

SKILL.md                        the method: five steps, and the rules that keep it honest
references/execution.md         worktrees, priming a build that will not compile, runner formats
references/probes.md            where to place a probe, how to keep the pair honest, pitfalls
references/report.md            report structure and the numbers the verdict must carry
scripts/setup_worktrees.sh      builds base/head/mix, including the revert that is easy to get wrong
scripts/collect_results.py      normalizes go test -json / JUnit XML / Jest JSON → {test: status}
scripts/compare_runs.py         three runs → behavior changes, flakes filtered out
assets/report-template.html     working skeleton of the report: levels, filtering, themes
examples/crush-pr-3456/         a full report produced by this skill on a public repo

Example

examples/crush-pr-3456/ is a real report on charmbracelet/crush#3456 — open it in a browser. It is worth reading because the test suite had nothing to say about that PR — zero of 2541 scenarios moved, since the file it changes had no test file before the branch. Every finding in it came from probes, including one behavior change the PR description does not mention.

Language support

Any runner that can emit machine-readable results works. collect_results.py auto-detects go test -json, JUnit/xUnit XML (pytest, Gradle, Maven, RSpec, nextest, dotnet, …), and the Jest/Vitest JSON reporter, and accepts a plain {"test name": "pass|fail|skip"} object from anything else.

What it is not

Not a code-quality review. It has no opinions about style, naming, or structure, and it will not hand you a list of suggested refactors. It tells you what changed, for whom, and how that was established.

License

MIT — see LICENSE.

About

An Agent Skill that reviews a pull request by running it: three builds, one subtraction, then probes for what no test covers.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages