Skip to content

Repository files navigation

agmind-bench

Reproducible measurement harness for local AI systems, by AGmind Systems Lab.

We qualify exact configurations — hardware × runtime build × model artifact × workload — and publish evidence, not impressions. This repository is the instrument: everything needed to run our measurements on your own machine and check our numbers, or produce your own.

Headline results (from the sealed evidence)

Measured on AMD Ryzen AI Max+ 395 / Radeon 8060S (gfx1151, Strix Halo), llama.cpp pinned by image digest, Qwen3.6-35B-A3B pinned by artifact hash — every number re-derived from raw run records in CI:

Question Measured answer
llama.cpp Vulkan vs ROCm decode 15.9 vs 18.7 ms/token (claim)
Prompt cache: 2nd question over a 32k doc 860 ms vs 33,728 ms cache off (claim)
Reasoning on vs off, time to first answer token 20,191 ms vs 210 ms at equal task success (claim)
Concurrency ladder c1/c4/c8, first answer token 210 / 338 / 865 ms, 100% completion (report)

Full table of all 40 claims: RESULTS.md · machine-readable: claims.json

What it measures differently

Most benchmarks report server-side tokens-per-second. This harness measures from the user's seat and treats several things as first-class that conventional tooling ignores:

  • Client-side timing. The clock starts when the request leaves the client and stops on tokens as they arrive over the wire. Server-reported timings can hide what the user actually experiences.
  • TTFT vs TTFA. Time to first token (any streamed token, including a reasoning model's thinking) is recorded separately from time to first answer token (the first token of the visible reply). On reasoning models these differ by orders of magnitude.
  • An empty answer is a failure. A response that returns HTTP 200 with no answer content (for example, the reasoning pass consumed the whole token budget) is counted as a failed request and stays in the denominator of every statistic.
  • Quality gates. Each response is checked against per-prompt format, language and repetition gates; gate outcomes ship with the run.
  • Evidence bundles. Every run emits a manifest (full identity of the cell), per-request records including failures, per-request gate outcomes, the runtime log, and a SHA-256 checksum file. Published numbers are re-derived from these records by the SQL in claims/sql/ — never typed by hand.

Repository layout

bench/run_cell.py     the harness: one measured cell per invocation
workloads/            workload definitions with frozen prompt corpora
schemas/              JSON Schema (2020-12) for the seven catalog entities
claims/sql/           DuckDB queries that derive published values from raw runs
tools/                artifact fetcher (hash-verified), fingerprint capture,
                      runtime smoke test
docs/METHODOLOGY.md   the rules a run must satisfy to count

Reproduce a run

Requirements: Python 3.10+ (stdlib only), Docker, and a quiesced machine. The harness launches the runtime container itself, records the exact command in the manifest, and removes the container when the run ends.

# 1. fetch the model artifact at a pinned revision, verify SHA-256
DEST=./models/upstream ./tools/fetch-artifacts.sh

# 2. run one cell (one container per run, identity fully pinned)
python3 bench/run_cell.py \
  --run-id my-run-01 \
  --system my-box \
  --runtime-id llamacpp-vulkan-b9049 \
  --image ghcr.io/ggml-org/llama.cpp:server-vulkan-b9049 \
  --model-id qwen36-35b-a3b-q4km-upstream \
  --model /models/upstream/Qwen3.6-35B-A3B-Q4_K_M.gguf \
  --models-dir ./models \
  --workload interactive-assistant-v2 \
  --workload-revision 2026-08-02 \
  --max-tokens 1024 \
  --out ./runs

# 3. inspect the bundle
cat ./runs/my-run-01/metrics.json

Flags of note: --no-think disables the reasoning block via the chat template; --max-tokens sets the completion budget; --backend-args passes device flags to docker run (default --device /dev/dri for Vulkan); --duration-min N switches to endurance mode — the corpus is cycled with bounded closed-loop submission until N minutes elapse, and every record carries a t_start_s offset so drift can be derived over 5-minute windows; --cache-prompt enables the server prompt cache for session/cache studies (pair it with --warmup 1: only corpus[0] is sacrificial there, and the default two warmup requests would prefill a document that must be measured cold); --allow-shared-host records, but does not invalidate, a run made while other GPU-capable containers were present — without it a run is invalid if the host inventory shows a GPU-capable neighbor container at the start or the end of the run.

The prompt cache is disabled by default and every request carries a unique salt so nothing is answered from a previous request's work; the cache setting is part of the recorded cell identity in the manifest.

A run only counts if it satisfies docs/METHODOLOGY.md — including host quiescing, pinned identities, and failure accounting.

Check a published number

Each file in claims/sql/ is a DuckDB query with a $runs placeholder. A claim names its run IDs; the derivation replaces $runs with a DuckDB list of those runs' requests.jsonl files (queries over gate outcomes declare quality.jsonl as their source instead) and executes the query. For claims/sql/ttft-ctx32k-en.sql over two runs, that is literally:

duckdb -c "
SELECT median(ttft_ms) AS value
FROM read_json_auto(['runs/my-run-01/requests.jsonl', 'runs/my-run-02/requests.jsonl'])
WHERE ok AND band = 'ctx32k' AND lang = 'en';
"

The query must return a single value; re-running it against the named runs must reproduce the published number exactly.

Disagree with our numbers?

Good — that is the point of publishing the instrument. Run the same cell on the same configuration and open an issue with your bundle. Negative results and non-reproductions are treated as first-class outcomes.

License

Apache-2.0. Prompt corpora in workloads/ are covered by the same license.

About

Reproducible local-LLM benchmark harness: llama.cpp on AMD Strix Halo (gfx1151, Ryzen AI Max+ 395) and NVIDIA DGX Spark — frozen corpora, quality gates with unit tests, sealed run bundles. Apache-2.0

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages