Skip to content

About

Privacy-preserving, local-LLM mental-health screening companion (PHQ-9/GAD-7 triage) with an Electron/React clay UI

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Privacy-Preserving Mental-Health Screening Platform

Frontline screening & triage for people with disabilities, caregivers, and families — based on the project proposal "แพลตฟอร์มคัดกรองและประเมินสุขภาพจิต เบื้องต้นด้วย Privacy-Preserving GenAI".

This is a screening tool, not a diagnosis. It never replaces a clinician. Crisis path always surfaces hotline 1323 (free, 24/7).

The accuracy story (read this before claiming a number)

We do not claim "AI guesses mental health with 90% accuracy" — that would be a diagnostic claim and is indefensible. Instead, screening performance is inherited from clinically validated instruments, scored deterministically:

Instrument Cutoff Published sensitivity / specificity
PHQ-9 (depression) ≥10 ~88% / ~88%
GAD-7 (anxiety) ≥10 ~89% / ~82%

The LLM is only a conversational front-end (empathetic interviewing, psychological first aid). It never scores risk and never decides crisis status.

Report sensitivity / specificity / recall-at-crisis — never a single "accuracy" number (a do-nothing model is ~90% accurate on imbalanced data).

Architecture

user input (questions / free text / voice)
        │
        ▼
  crisis.py        ← deterministic self-harm scan (TH+EN), high recall
        │ crisis_flag
        ▼
  screening.py     ← PHQ-9 + GAD-7 scoring → Green / Yellow / Orange / Red
        │
        ▼
  LLM front-end    ← Qwen, guardrailed: psychological first aid only (TODO)
        │
        ▼
  triage / escalation (hotline 1323 on RED)

Files

File Role
screening.py Deterministic PHQ-9/GAD-7 scoring + risk banding (accuracy core)
crisis.py Crisis keyword/pattern layer (Thai + English), forces RED
llm_interviewer.py Local Qwen 2.5 front-end (Ollama), crisis-guarded, psychological first aid only
evaluate.py Sensitivity/specificity/recall report over a labeled dataset
voice_features.py Vocal biomarker extraction (in-memory, features-only logging)
data/sample_dataset.jsonl Tiny example dataset (replace with real labeled data)
data/runtime/ All runtime-generated files (user config, logs, journals, datasets) — gitignored, private
tests/ pytest suite (pytest tests/ -v)

Quick start

python3 screening.py                       # smoke test the scorer
python3 crisis.py                          # smoke test crisis detection
python3 evaluate.py data/sample_dataset.jsonl   # metrics report

Local AI front-end (private, on-device via Ollama)

The conversational layer runs entirely on your Mac through Ollama — no text leaves the machine.

ollama serve                               # start the local model server (if not running)
ollama pull qwen2.5:3b                      # one-time, ~1.9 GB (already present here)
python3 llm_interviewer.py                  # empathetic chat front-end (crisis-guarded)
python3 llm_interviewer.py --screen         # structured PHQ-9/GAD-7 flow + reflection

The crisis layer runs before the model on every turn; on a hit it returns the deterministic hotline message and the model is never consulted for the verdict. The model is psychological-first-aid scope only — it never scores or diagnoses.

Desktop chat app (Electron + React + Vite)

A phone-style chat UI over the same local model. The Electron window is pure UI; it spawns server.py (stdlib HTTP bridge) which reuses the exact crisis.py / screening.py / llm_interviewer.py modules — one source of truth for safety logic.

cd app
npm install            # one-time (Electron + React + Vite)
npm start              # vite build -> opens Electron; auto-starts the Python bridge
npm run dev            # UI-only hot reload in a browser (run `python3 server.py` first;
                       #   /api is proxied to the bridge on :8765)
app/
  main.js              # Electron main: spawns the Python bridge, loads renderer-dist/
  preload.js           # secure contextBridge -> window.mh (fetch to localhost only)
  vite.config.mjs      # React renderer build config (+ /api dev proxy)
  src/                 # React renderer
    components/        #   App, TopBar, Messages, Composer, sheets, ModePicker, VoiceFirst
    lib/               #   bridge (window.mh fallback), recorder hook, TTS, icons
    styles.css         #   aurora-glass theme (plain CSS, class names = design contract)
  renderer-dist/       # vite build output (gitignored; served to phones by server.py)
server.py              # localhost HTTP bridge: /api/health /api/chat /api/screen

The npm start script strips ELECTRON_RUN_AS_NODE (set inside some IDE terminals) so Electron always boots as a GUI app. Requires Python 3 and Ollama on PATH.

Voice mode (fully local — no audio leaves the machine)

  • Speech-to-text: whisper.cpp. The mic is captured as 16 kHz WAV in the app and transcribed by a local whisper-cli.
  • Text-to-speech: the OS voices via the Web Speech API (macOS th-TH / en-US).

One-time setup:

brew install whisper-cpp
mkdir -p models
# base (~142MB, fast). For better Thai use ggml-small.bin (~466MB) or ggml-medium.bin.
curl -L -o models/ggml-base.bin \
  https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.bin

Accessibility & communication modes

On first launch the app asks how the user prefers to communicate — instead of assuming one interface fits every disability:

Mode For What changes
เสียงเป็นหลัก (voice-first) blind / low-vision users whole screen becomes the talk button; replies always spoken; zero reading required
ตัวหนังสือใหญ่ (large text) low vision / older users max font scale + high-contrast theme, on by default
แชทปกติ (standard) everyone else regular chat UI

Always available in the top bar: font scaling (ก / ก+ / ก++) and a high-contrast toggle (the pastel theme's muted text is ~2.4:1 contrast, below WCAG AA 4.5:1 — the toggle lifts every text token above AA). Icon buttons carry Thai aria-labels, keyboard focus is visible, and animations respect prefers-reduced-motion. Preferences persist locally; the mode can be changed anytime from settings.

Vocal biomarker analysis (supplementary signal — never a diagnosis)

Every voice clip is additionally analyzed in memory for vocal biomarkers, then the audio is discarded. Raw audio is never stored — only derived numbers persist (data/runtime/voice_sessions.jsonl), from which speech cannot be reconstructed. Features per clip:

Feature What it captures Depression-linked direction
f0_sd_semitones pitch variability ↓ (monotone voice)
speech_rate_sps syllables/sec over speech time ↓ (slower speech)
pause_ratio, long_pauses silence behavior ↑ (longer/more pauses)
rms_mean vocal energy ↓ (quieter voice)

After ≥3 sessions a personal baseline forms; each new clip is compared to the user's own average (absolute values differ per person, trends don't). Deviations surface as neutral notes in the chat UI ("น้ำเสียงราบเรียบกว่าช่วงก่อนหน้า") — a supplementary signal alongside PHQ-9/GAD-7, never a score, never a diagnosis, never consulted by the crisis layer. Basis: Cummins et al. 2015 (Speech Communication), Mundt et al. 2012.

pip install librosa numpy    # one-time; server degrades gracefully without it
python3 voice_features.py    # smoke test on a synthetic clip

In the app: 🎤 records (click to start, click to stop → auto-transcribed and sent); 🔊 toggles spoken replies. Override the model with MH_WHISPER_MODEL=models/ggml-small.bin. Chromium's built-in SpeechRecognition is deliberately not used — it streams audio to Google, which would break the privacy guarantee.

Risk bands

Band Meaning Action
🟢 Green minimal / low self-care content
🟡 Yellow early signs preventive tips, mood tracking
🟠 Orange moderate–high recommend professional / hotline 1323
🔴 Red crisis / self-harm risk immediate safety mode + 1323

TODO (next)

  • llm_interviewer.py — Qwen 2.5 wired as the empathetic front-end with strict guardrails (psychological first aid scope only; never diagnoses), local via Ollama.
  • Vocal-biomarker feature extraction (supplementary signal only) — voice_features.py: in-memory analysis, features-only logging, personal baseline trends. Audio never touches the database.
  • Replace data/sample_dataset.jsonl with a real labeled validation set and re-run evaluate.py to report defensible metrics.
  • PDPA consent flow + edge/on-device processing.

About

Privacy-preserving, local-LLM mental-health screening companion (PHQ-9/GAD-7 triage) with an Electron/React clay UI

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages