Reproducible measurement harness for local AI systems, by AGmind Systems Lab.
We qualify exact configurations — hardware × runtime build × model artifact × workload — and publish evidence, not impressions. This repository is the instrument: everything needed to run our measurements on your own machine and check our numbers, or produce your own.
Measured on AMD Ryzen AI Max+ 395 / Radeon 8060S (gfx1151, Strix Halo), llama.cpp pinned by image digest, Qwen3.6-35B-A3B pinned by artifact hash — every number re-derived from raw run records in CI:
| Question | Measured answer |
|---|---|
| llama.cpp Vulkan vs ROCm decode | 15.9 vs 18.7 ms/token (claim) |
| Prompt cache: 2nd question over a 32k doc | 860 ms vs 33,728 ms cache off (claim) |
| Reasoning on vs off, time to first answer token | 20,191 ms vs 210 ms at equal task success (claim) |
| Concurrency ladder c1/c4/c8, first answer token | 210 / 338 / 865 ms, 100% completion (report) |
Full table of all 40 claims: RESULTS.md · machine-readable: claims.json
Most benchmarks report server-side tokens-per-second. This harness measures from the user's seat and treats several things as first-class that conventional tooling ignores:
- Client-side timing. The clock starts when the request leaves the client and stops on tokens as they arrive over the wire. Server-reported timings can hide what the user actually experiences.
- TTFT vs TTFA. Time to first token (any streamed token, including a reasoning model's thinking) is recorded separately from time to first answer token (the first token of the visible reply). On reasoning models these differ by orders of magnitude.
- An empty answer is a failure. A response that returns HTTP 200 with no answer content (for example, the reasoning pass consumed the whole token budget) is counted as a failed request and stays in the denominator of every statistic.
- Quality gates. Each response is checked against per-prompt format, language and repetition gates; gate outcomes ship with the run.
- Evidence bundles. Every run emits a manifest (full identity of the
cell), per-request records including failures, per-request gate outcomes,
the runtime log, and a SHA-256 checksum file. Published numbers are
re-derived from these records by the SQL in
claims/sql/— never typed by hand.
bench/run_cell.py the harness: one measured cell per invocation
workloads/ workload definitions with frozen prompt corpora
schemas/ JSON Schema (2020-12) for the seven catalog entities
claims/sql/ DuckDB queries that derive published values from raw runs
tools/ artifact fetcher (hash-verified), fingerprint capture,
runtime smoke test
docs/METHODOLOGY.md the rules a run must satisfy to count
Requirements: Python 3.10+ (stdlib only), Docker, and a quiesced machine. The harness launches the runtime container itself, records the exact command in the manifest, and removes the container when the run ends.
# 1. fetch the model artifact at a pinned revision, verify SHA-256
DEST=./models/upstream ./tools/fetch-artifacts.sh
# 2. run one cell (one container per run, identity fully pinned)
python3 bench/run_cell.py \
--run-id my-run-01 \
--system my-box \
--runtime-id llamacpp-vulkan-b9049 \
--image ghcr.io/ggml-org/llama.cpp:server-vulkan-b9049 \
--model-id qwen36-35b-a3b-q4km-upstream \
--model /models/upstream/Qwen3.6-35B-A3B-Q4_K_M.gguf \
--models-dir ./models \
--workload interactive-assistant-v2 \
--workload-revision 2026-08-02 \
--max-tokens 1024 \
--out ./runs
# 3. inspect the bundle
cat ./runs/my-run-01/metrics.jsonFlags of note: --no-think disables the reasoning block via the chat
template; --max-tokens sets the completion budget; --backend-args
passes device flags to docker run (default --device /dev/dri for
Vulkan); --duration-min N switches to endurance mode — the corpus is
cycled with bounded closed-loop submission until N minutes elapse, and
every record carries a t_start_s offset so drift can be derived over
5-minute windows; --cache-prompt enables the server prompt cache for
session/cache studies (pair it with --warmup 1: only corpus[0] is
sacrificial there, and the default two warmup requests would prefill a
document that must be measured cold); --allow-shared-host records, but
does not invalidate, a run made while other GPU-capable containers were
present — without it a run is invalid if the host inventory shows a
GPU-capable neighbor container at the start or the end of the run.
The prompt cache is disabled by default and every request carries a unique salt so nothing is answered from a previous request's work; the cache setting is part of the recorded cell identity in the manifest.
A run only counts if it satisfies docs/METHODOLOGY.md — including host quiescing, pinned identities, and failure accounting.
Each file in claims/sql/ is a DuckDB query with a $runs placeholder.
A claim names its run IDs; the derivation replaces $runs with a DuckDB
list of those runs' requests.jsonl files (queries over gate outcomes
declare quality.jsonl as their source instead) and executes the query.
For claims/sql/ttft-ctx32k-en.sql over two runs, that is literally:
duckdb -c "
SELECT median(ttft_ms) AS value
FROM read_json_auto(['runs/my-run-01/requests.jsonl', 'runs/my-run-02/requests.jsonl'])
WHERE ok AND band = 'ctx32k' AND lang = 'en';
"The query must return a single value; re-running it against the named
runs must reproduce the published number exactly.
Good — that is the point of publishing the instrument. Run the same cell on the same configuration and open an issue with your bundle. Negative results and non-reproductions are treated as first-class outcomes.
Apache-2.0. Prompt corpora in workloads/ are covered by the same license.