AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results
-
Updated
May 25, 2026 - Python
AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results
Methodology and deterministic evaluation tools for iterative improvement of skills, repositories and agent workflows, with independent review and regression checks.
A Simple Way to Eliminate Reward Hacking in GRPO Diffusion Alignment
Real-time reward debugging and hacking detection for reinforcement learning
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
An agent for auditing repositories of traces for violations of safety properties. Automatically finds cheating (task-level gaming and harness-level cheating) on top benchmarks.
CHERRL: A Controllable Hacking Environment for Rubric-Based Reinforcement Learning
Runtime detector for reward hacking and misalignment in LLM agents (89.7% F1 on 5,391 trajectories).
Exit-code checks for AI coding agents: catches test tampering in diffs, tests that cannot fail, fixes that only satisfy the tested inputs, and weakened assertions. Benchmarked on 630 real PRs and QuixBugs.
Stop trusting 'done'. Make AI coding agents prove their work — frozen specs, tamper-detected tests, an independent blind verifier, hard-blocking gates. Fails loudly instead of pretending. For Claude Code, Codex, OpenCode & any Agent Skills runtime. No unverified work passes.
Reverse-engineer the RL reward an LLM was trained on — from black-box agentic coding behavior alone. Tested on Opus 4.8, Fable 5, GPT-5.5.
Forward and reverse KL divergence monitor detecting policy drift and reward hacking across distributions
Forward and reverse KL divergence monitor detecting policy drift and reward hacking across distributions
A GitHub App that catches AI coding agents cheating to make CI green. Your Coderabbit open source alternative at lightning fast speed.
Experiment code for 'Are we really tilting? The mechanics of reward guidance in flow and diffusion models' — plug-in Doob h-transform sampling, reward damping, best-of-n, and flow map reward guidance for Gaussian mixtures, a 2D checkerboard, and FLUX.1 text-to-image generation.
Plug-and-play reward monitoring for RL training loops. Catch reward hacking, component imbalance, and starvation before they tank your run. Drop in one .step() call — get balance reports, auto weight correction, alignment scores, and WandB/TensorBoard/SB3 integrations out of the box. → rewardguard.dev
🐀 Fuzz your verifier before an RL agent does. Static + dynamic LLM security auditor to detect reward-hacking in RL post-training environments (OpenEnv, verifiers-spec, Gymnasium).
Sealed-lab research harness for emergent multi-agent (LLM) coordination and containment — replicates the dynamics of the 2026 OpenAI–Hugging Face incident against self-hosted, intentionally-vulnerable targets. Air-gapped by design. Defensive / AI-safety only.
End-to-end RLHF pipeline: reward modeling, PPO/DPO/GRPO, reward signal design, FSDP scaling analysis, and agent evaluation on GPT-2
Code for the paper "Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement"
To associate your repository with the reward-hacking topic, visit your repo's landing page and select "manage topics."