Frontline screening & triage for people with disabilities, caregivers, and families — based on the project proposal "แพลตฟอร์มคัดกรองและประเมินสุขภาพจิต เบื้องต้นด้วย Privacy-Preserving GenAI".
This is a screening tool, not a diagnosis. It never replaces a clinician. Crisis path always surfaces hotline 1323 (free, 24/7).
We do not claim "AI guesses mental health with 90% accuracy" — that would be a diagnostic claim and is indefensible. Instead, screening performance is inherited from clinically validated instruments, scored deterministically:
| Instrument | Cutoff | Published sensitivity / specificity |
|---|---|---|
| PHQ-9 (depression) | ≥10 | ~88% / ~88% |
| GAD-7 (anxiety) | ≥10 | ~89% / ~82% |
The LLM is only a conversational front-end (empathetic interviewing, psychological first aid). It never scores risk and never decides crisis status.
Report sensitivity / specificity / recall-at-crisis — never a single "accuracy" number (a do-nothing model is ~90% accurate on imbalanced data).
user input (questions / free text / voice)
│
▼
crisis.py ← deterministic self-harm scan (TH+EN), high recall
│ crisis_flag
▼
screening.py ← PHQ-9 + GAD-7 scoring → Green / Yellow / Orange / Red
│
▼
LLM front-end ← Qwen, guardrailed: psychological first aid only (TODO)
│
▼
triage / escalation (hotline 1323 on RED)
| File | Role |
|---|---|
screening.py |
Deterministic PHQ-9/GAD-7 scoring + risk banding (accuracy core) |
crisis.py |
Crisis keyword/pattern layer (Thai + English), forces RED |
llm_interviewer.py |
Local Qwen 2.5 front-end (Ollama), crisis-guarded, psychological first aid only |
evaluate.py |
Sensitivity/specificity/recall report over a labeled dataset |
voice_features.py |
Vocal biomarker extraction (in-memory, features-only logging) |
data/sample_dataset.jsonl |
Tiny example dataset (replace with real labeled data) |
data/runtime/ |
All runtime-generated files (user config, logs, journals, datasets) — gitignored, private |
tests/ |
pytest suite (pytest tests/ -v) |
python3 screening.py # smoke test the scorer
python3 crisis.py # smoke test crisis detection
python3 evaluate.py data/sample_dataset.jsonl # metrics reportThe conversational layer runs entirely on your Mac through Ollama — no text leaves the machine.
ollama serve # start the local model server (if not running)
ollama pull qwen2.5:3b # one-time, ~1.9 GB (already present here)
python3 llm_interviewer.py # empathetic chat front-end (crisis-guarded)
python3 llm_interviewer.py --screen # structured PHQ-9/GAD-7 flow + reflectionThe crisis layer runs before the model on every turn; on a hit it returns the deterministic hotline message and the model is never consulted for the verdict. The model is psychological-first-aid scope only — it never scores or diagnoses.
A phone-style chat UI over the same local model. The Electron window is pure UI;
it spawns server.py (stdlib HTTP bridge) which reuses the exact crisis.py /
screening.py / llm_interviewer.py modules — one source of truth for safety logic.
cd app
npm install # one-time (Electron + React + Vite)
npm start # vite build -> opens Electron; auto-starts the Python bridge
npm run dev # UI-only hot reload in a browser (run `python3 server.py` first;
# /api is proxied to the bridge on :8765)app/
main.js # Electron main: spawns the Python bridge, loads renderer-dist/
preload.js # secure contextBridge -> window.mh (fetch to localhost only)
vite.config.mjs # React renderer build config (+ /api dev proxy)
src/ # React renderer
components/ # App, TopBar, Messages, Composer, sheets, ModePicker, VoiceFirst
lib/ # bridge (window.mh fallback), recorder hook, TTS, icons
styles.css # aurora-glass theme (plain CSS, class names = design contract)
renderer-dist/ # vite build output (gitignored; served to phones by server.py)
server.py # localhost HTTP bridge: /api/health /api/chat /api/screen
The npm start script strips ELECTRON_RUN_AS_NODE (set inside some IDE terminals)
so Electron always boots as a GUI app. Requires Python 3 and Ollama on PATH.
- Speech-to-text: whisper.cpp. The mic
is captured as 16 kHz WAV in the app and transcribed by a local
whisper-cli. - Text-to-speech: the OS voices via the Web Speech API (macOS
th-TH/en-US).
One-time setup:
brew install whisper-cpp
mkdir -p models
# base (~142MB, fast). For better Thai use ggml-small.bin (~466MB) or ggml-medium.bin.
curl -L -o models/ggml-base.bin \
https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.binOn first launch the app asks how the user prefers to communicate — instead of assuming one interface fits every disability:
| Mode | For | What changes |
|---|---|---|
| เสียงเป็นหลัก (voice-first) | blind / low-vision users | whole screen becomes the talk button; replies always spoken; zero reading required |
| ตัวหนังสือใหญ่ (large text) | low vision / older users | max font scale + high-contrast theme, on by default |
| แชทปกติ (standard) | everyone else | regular chat UI |
Always available in the top bar: font scaling (ก / ก+ / ก++) and a
high-contrast toggle (the pastel theme's muted text is ~2.4:1 contrast,
below WCAG AA 4.5:1 — the toggle lifts every text token above AA). Icon
buttons carry Thai aria-labels, keyboard focus is visible, and animations
respect prefers-reduced-motion. Preferences persist locally; the mode can
be changed anytime from settings.
Every voice clip is additionally analyzed in memory for vocal biomarkers,
then the audio is discarded. Raw audio is never stored — only derived
numbers persist (data/runtime/voice_sessions.jsonl), from which speech
cannot be reconstructed. Features per clip:
| Feature | What it captures | Depression-linked direction |
|---|---|---|
f0_sd_semitones |
pitch variability | ↓ (monotone voice) |
speech_rate_sps |
syllables/sec over speech time | ↓ (slower speech) |
pause_ratio, long_pauses |
silence behavior | ↑ (longer/more pauses) |
rms_mean |
vocal energy | ↓ (quieter voice) |
After ≥3 sessions a personal baseline forms; each new clip is compared to the user's own average (absolute values differ per person, trends don't). Deviations surface as neutral notes in the chat UI ("น้ำเสียงราบเรียบกว่าช่วงก่อนหน้า") — a supplementary signal alongside PHQ-9/GAD-7, never a score, never a diagnosis, never consulted by the crisis layer. Basis: Cummins et al. 2015 (Speech Communication), Mundt et al. 2012.
pip install librosa numpy # one-time; server degrades gracefully without it
python3 voice_features.py # smoke test on a synthetic clipIn the app: 🎤 records (click to start, click to stop → auto-transcribed and sent);
🔊 toggles spoken replies. Override the model with MH_WHISPER_MODEL=models/ggml-small.bin.
Chromium's built-in SpeechRecognition is deliberately not used — it streams
audio to Google, which would break the privacy guarantee.
| Band | Meaning | Action |
|---|---|---|
| 🟢 Green | minimal / low | self-care content |
| 🟡 Yellow | early signs | preventive tips, mood tracking |
| 🟠 Orange | moderate–high | recommend professional / hotline 1323 |
| 🔴 Red | crisis / self-harm risk | immediate safety mode + 1323 |
-
llm_interviewer.py— Qwen 2.5 wired as the empathetic front-end with strict guardrails (psychological first aid scope only; never diagnoses), local via Ollama. - Vocal-biomarker feature extraction (supplementary signal only) —
voice_features.py: in-memory analysis, features-only logging, personal baseline trends. Audio never touches the database. - Replace
data/sample_dataset.jsonlwith a real labeled validation set and re-runevaluate.pyto report defensible metrics. - PDPA consent flow + edge/on-device processing.