Pinned Loading
-
policy-to-eval-harness
policy-to-eval-harness PublicTurn a written AI usage policy into a running evaluation. 12-category taxonomy, 400 borderline prompts, 5 open-weight models, judge calibrated against blind human labels.
Python
-
multilingual-enforcement-consistency
multilingual-enforcement-consistency PublicDoes a model's safety survive translation? Takes a policy-derived taxonomy, translates 30 borderline prompts into English, Chinese, French, Singlish and Singapore Mandarin, and measures whether ref…
Python
-
apac-regulatory-readiness
apac-regulatory-readiness PublicA maintained tracker of platform regulation across 10 APAC jurisdictions, plus a regulator response workflow, an AI toolkit with a measured evaluation harness, and a readiness assessment framework.
HTML
-
distress-conversation-safety-eval
distress-conversation-safety-eval PublicRubric-based safety evaluation for multi-turn conversations with escalating user distress. Measures the failure single-turn evals cannot see: position hold rate, drift slope, turns to first failure.
Python
-
enforcement-ops-simulator
enforcement-ops-simulator PublicDiscrete-event simulation of a Trust & Safety enforcement queue, with the operating model it is built to test: severity matrix, decision tree, escalation paths, calibration cadence, and AAR.
Python
-
crisis-comms-wargame
crisis-comms-wargame PublicAn adversarial multi-agent war-game for AI-safety crisis communications, and the playbook it was built to test. Four incidents, four adversary agents, three response postures, twelve deterministic …
Python
If the problem persists, check the GitHub status page or contact support.