Skip to content

About

Diagnostic Assistant Demo for ESWC 2026

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

125 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

A Prototyping Environment for Comparing Neuro-symbolic Diagnostic Assistants

Codebase associated with the ESWC 2026 demo paper "A Prototyping Environment for Comparing Neuro-symbolic Diagnostic Assistants".

The environment runs diagnostic scenarios: a fault is injected into a small simulated electrical system by a Saboteur agent, and a Diagnostic Assistant guides a Service Agent toward identifying the root cause. The goal is to compare different assistant strategies — LLM-based, neuro-symbolic, and baselines — under controlled conditions.

Environment architecture


Repository structure

ESWC_2026_Demo/
├── environment_classes.py          # Abstract agent roles and scenario orchestrator
├── configuration.py                # Configuration dataclass and defaults
├── run_diagnostic_scenario.py      # Entry point: run a single diagnostic scenario
├── run_evaluation_protocol.py      # Entry point: run the Adaptive Evaluation Protocol
├── run_many_scenarios.py           # Batch runner for multiple scenarios
├── make_tables.py                  # Aggregate evaluation results into tables
├── SCENARIOS_MASTER.csv            # Master list of all 135 scenarios with metadata
│
├── Implementations/                # Concrete agent implementations
│   ├── diagnosticAssistantLLM.py           # Monolithic LLM assistant
│   ├── diagnosticAssistantEvidenceKGOptimal.py  # Neuro-symbolic assistant
│   ├── diagnosticAssistantRandomSearch.py  # Random baseline (components one-by-one)
│   ├── diagnosticAssistantFixedRandomTrajectories.py  # Fixed random baseline
│   ├── diagnosticAssistantUnhelpful.py     # Always-wrong baseline
│   ├── fault_injections.py         # SCENARIOS list (135 injectable faults)
│   ├── saboteur*.py                # Saboteur implementations
│   └── serviceAgent*.py            # Service agent implementations
│
├── Evaluation/                     # Adaptive Evaluation Protocol
│   ├── protocol.py                 # Main protocol loop (HDBSCAN + ARI convergence)
│   ├── clustering.py               # Intent & execution clustering
│   ├── qualitative.py              # Qualitative rubric evaluation (LLM-based)
│   ├── report.py                   # Per-scenario report formatting
│   ├── metrics.py                  # Numerical metrics (success rate, cost, actions)
│   └── checkpoint.py               # Checkpoint save/load
│
├── Knowledge_sources/
│   ├── Structured_knowledge_sources/   # OWL ontology (TBox) + per-system ABox KGs
│   └── Unstructured_knowledge_sources/ # Textual descriptions + circuit schematics
│
├── Utilities/                      # Shared helpers
│   ├── OWL_reasoning.py            # HermiT reasoner wrapper (owlready2)
│   ├── retrieval.py                # RAG: chunking, embedding, top-k retrieval
│   ├── topology.py                 # Graph algorithms over the KG
│   ├── caching.py                  # SQLite-backed LLM call cache
│   ├── chat_log.py                 # HTML chat transcript writer
│   ├── agents_boilerplate.py       # OpenAI Agents SDK setup helpers
│   └── ...
│
├── Tests/                          # pytest test suite
├── static/client.html              # Browser voice interface
├── voice_server.py                 # FastAPI voice server (Whisper STT + TTS)
└── voice_client.py                 # Voice client helpers

Systems

Five simulated electrical systems are supported, all implemented as SPICE circuits via the diagnosable-systems-simulation package:

System --system flag Scenarios
3-module cube circuit 3CubesSystem 1–25
10-module cube circuit 10CubesSystem 26–50
Asymmetric chains AsymmetricChainsSystem 51–75
Ambient light sensor AmbientLightSensorSystem 76–105
Current sensor CurrentSensorSystem 106–135

Assistant implementations

--assistant flag Class Description
LLM DiagnosticAssistantLLM Monolithic GPT agent with RAG and vision
EvidenceKGOptimal DiagnosticAssistantEvidenceKGOptimal Neuro-symbolic: OWL reasoning + SPARQL + information-gain heuristic
RandomSearch DiagnosticAssistantRandomSearch Random baseline: hypothesises components one-by-one without replacement
FixedRandomTrajectories DiagnosticAssistantFixedRandomTrajectories Fixed random action sequences (reproducible baseline)
Unhelpful DiagnosticAssistantUnhelpful Always-wrong baseline

The neuro-symbolic assistant pipeline is summarised below:

Neuro-symbolic assistant architecture


Installation

1. Install Python dependencies

pip install -r requirements.txt

Note: torch is a heavy dependency pulled in by sentence-transformers. If you only want to run scenarios without the evaluation protocol, you can skip it.

2. Install the simulation backend

The SPICE simulation backend is a separate package:

git clone https://github.com/kataph/diagnosable-systems-simulation
pip install -e diagnosable-systems-simulation/

3. Set API credentials

export OPENAI_API_KEY="sk-..."          # OpenAI, or
export OPENAI_BASE_URL="https://openrouter.ai/api/v1"  # any OpenAI-compatible endpoint
export OPENAI_API_KEY="sk-or-..."       # OpenRouter key

Running a scenario

python -m run_diagnostic_scenario --help

3-cubes system, LLM assistant, CLI:

python -m run_diagnostic_scenario \
  --text-input-file "Knowledge_sources/Unstructured_knowledge_sources/3_cubes/3_cubes_description.txt" \
  --diagram "Knowledge_sources/Unstructured_knowledge_sources/3_cubes/3_cubes_schematics.png" \
  --kg "Knowledge_sources/Structured_knowledge_sources/3_cubes/zorro-ontology-3-cubes-abox.ttl" \
  --ontology "Knowledge_sources/Structured_knowledge_sources/zorro-ontology-tbox.ttl" \
  --retrieval-folder "Knowledge_sources/Unstructured_knowledge_sources/3_cubes" \
  --system 3CubesSystem --saboteur Human --service Human --assistant LLM --interface cli

10-cubes system, neuro-symbolic assistant:

python -m run_diagnostic_scenario \
  --text-input-file "Knowledge_sources/Unstructured_knowledge_sources/10_cubes/10_cubes_description.txt" \
  --diagram "Knowledge_sources/Unstructured_knowledge_sources/10_cubes/10_cubes_schematics.png" \
  --kg "Knowledge_sources/Structured_knowledge_sources/10_cubes/zorro-ontology-10-cubes-abox.ttl" \
  --ontology "Knowledge_sources/Structured_knowledge_sources/zorro-ontology-tbox.ttl" \
  --retrieval-folder "Knowledge_sources/Unstructured_knowledge_sources/10_cubes" \
  --system 10CubesSystem --saboteur Human --service Human --assistant EvidenceKGOptimal --interface cli

Voice interface (requires running uvicorn separately):

uvicorn voice_server:app --host 0.0.0.0 --port 8000
# Then browse to http://127.0.0.1:8000/client.html

Running the evaluation protocol

The Adaptive Evaluation Protocol collects batches of trajectories per scenario and checks for convergence of the trajectory clustering (via inter-batch ARI).

python -m run_evaluation_protocol --help

Example (automated, all scenarios):

python -m run_evaluation_protocol \
  --all-scenarios \
  --assistant EvidenceKGOptimal --service SpiceSim \
  --service-config '{"model":"openai/gpt-4.1-mini"}' \
  --assistant-config '{"model":"openai/gpt-4.1","embed_model":"openai/text-embedding-3-small"}' \
  --labeling-model openai/gpt-4.1-mini \
  --batch-size 10 --rounds 10 --max-concurrent 5

Incremental runs (manual convergence control):

# First batch
python -m run_evaluation_protocol --scenarios 1,2,3,4,5 \
  --assistant EvidenceKGOptimal --service SpiceSim \
  [flags] --batch-size 10 --max-batches 1
# note the run_id printed at startup

# Add another batch
python -m run_evaluation_protocol --scenarios 1,2,3,4,5 \
  --resume <run_id> [flags] --batch-size 10 --max-batches 2

# Add more scenarios to the same run
python -m run_evaluation_protocol --scenarios 6,7,8,9,10 \
  --resume <run_id> [flags] --batch-size 10 --max-batches 1

Monitor progress:

tail -f Logs/Evaluation/<AssistantType>/<run_id>/protocol.log

Generating evaluation tables

# Auto-detect latest run per assistant:
python make_tables.py

# Specify runs explicitly:
python make_tables.py --run EvidenceKGOptimal:<run_id> --run LLM:<run_id>

Produces tables broken down by scenario, system, and fault type, including qualitative analysis aggregates. LaTeX is auto-copied to clipboard if pyperclip is installed.


Running tests

pytest Tests/
pytest Tests/test_scenario1_spice.py -s   # single file with output

Cost model

Each diagnostic action and hypothesis verification incurs a cost recorded in the trajectory. The cost model has two tiers depending on whether SpiceSim is the service agent.

Action cost

Service agent Cost source Values
SpiceSim precise_action_cost — actual simulation time in minutes Varies per action type (e.g. observe_component=10, measure_voltage=20, replace_component=120)
Human, LLM, Mock ACTION_COST_MAP fallback Observe=1, Test=2, Adjust=4, Replace=12 (qualitative ordering only)

Both values are logged side-by-side as kg_cost (fallback) and sim_time (precise) in the debugging log.

Hypothesis verification cost

Hypothesis verification cost is also computed differently per service agent:

Service agent Correct/partial hypothesis Wrong hypothesis
SpiceSim apply_repairs() — sum of actual per-component repair costs: cable reconnection = 10 per port, component replacement = 120 _estimate_repair_cost() — same per-type rates applied to the attempted components, without mutating state
Human, LLM, Mock FALLBACK_HYPOTHESIS_VERIFICATION_COST = 120 flat Same flat cost

The SpiceSim costs scale naturally with the hypothesis: naming two crossed cables costs 20 (two reconnections at 10 each), while naming one degraded component costs 120. This means MODE 2 hypothesis testing is cost-equivalent to performing the same repairs via MODE 1 actions, with no arbitrage opportunity.


How repair correctness is determined (two paths)

A session ends successfully either because the assistant submitted a correct hypothesis (the hypothesis path) or because a direct repair action restored the system without any hypothesis (the action path). The correctness check differs substantially between the two.

Hypothesis path (verify_hypothesis)

Called when DiagnosticAssistant.suggest_action() returns a DiagnosticFaultHypothesis. The service agent calls system.test_repair(component_ids), which:

  1. Resets the circuit to the fault snapshot (via restore_snapshot).
  2. Applies the hypothesised repairs to the live circuit (reconnects cables, clears fault overlays, sets is_rotated=True for PhysicalEnclosure targets).
  3. Re-simulates and checks is_system_nominal().
  4. Restores the circuit back to the fault state before returning, so the test is non-destructive.

If test_repair returns True, the session ends with end=success.

Action path (execute_action + decide_finish)

Called for every ordinary diagnostic action. The service agent uses nl_interface to execute the action on the live circuit, then calls decide_finish to check whether the system has been restored.

The complication: after executing each batch of actions, ServiceAgentSpiceSim calls restore_snapshot to undo any transient side-effects (e.g. a switch left open, a measurement that perturbed state) before decide_finish re-simulates. This reset must not undo actions that were genuine repairs — otherwise decide_finish would see the broken circuit again and fail to detect success.

To distinguish repair actions from transient diagnostic actions, the system uses two layers:

  1. Always-permanent actions (replace_component, reconnect_cable): these remove a fault overlay or reconnect a disconnected cable. They are always tracked in _repaired_comp_ids and excluded from the snapshot reset, regardless of the scenario.

  2. Scenario-specific permanent actions: some actions are only repairs in specific scenarios. For example, rotate_enclosure on cube_ctrl is a repair in scenarios 101–105 (optical feedback loop fault), but merely a diagnostic action in any other scenario. reverse_battery is a repair only when the fault is a reversed battery. These are declared per-scenario in _PERMANENT_REPAIRS inside fault_injections.py as a frozenset[tuple[action_id, subject_component_id]] and stored on each Scenario as permanent_repair_actions.

When execute_action completes, it checks each executed action against both layers and adds qualifying subject component IDs to _repaired_comp_ids. restore_snapshot(exclude_ids=_repaired_comp_ids) then preserves those components in their repaired state. decide_finish re-simulates and calls is_system_nominal() — if True, the session ends with end=success_no_hypothesis.

Adding a new scenario with a novel repair action type: add a _PERMANENT_REPAIRS[scenario_number] entry in fault_injections.py with the (action_id, subject_component_id) pairs for that scenario's solution. Do not add the action to the always-permanent set unless it unconditionally removes a fault overlay regardless of scenario.

Test coverage

Tests/test_all_scenarios.py mirrors this two-path logic via conftest.run_sequence, which applies the same restore_snapshot + exclude_ids logic using the scenario's permanent_repair_actions. Tests that use RotateEnclosure, ReverseBattery, or SwapCablePolarities as their repair will fail if the corresponding entry is missing from _PERMANENT_REPAIRS — this is intentional: the tests catch B3-type bugs where a repair action is reset by restore_snapshot before decide_finish can detect success.


Trajectory end values and statistical treatment

Each trajectory is closed with one of five end values written to the JSON file:

Value Meaning Counted in statistics?
success Assistant submitted a correct hypothesis Yes — success
success_no_hypothesis System restored via a direct repair action (no hypothesis submitted) Yes — success
timeout Action budget exhausted without restoring the system Yes — failure
surrender Assistant returned None (no further suggestions) Yes — failure
llm_truncation Session terminated by a truncated/unparseable LLM response or an API network error Excluded from all statistics

Rationale for excluding llm_truncation: the cause is ambiguous. It can be a genuine model schema-compliance failure, a token-budget limit hit during a long reasoning chain, or a transient network/server error — all three produce the same observable outcome. Including these as failures would conflate infrastructure reliability with diagnostic competence and penalise models that happened to hit more API errors. Reported success rates are therefore a lower bound on true diagnostic ability given a working API call.


Logging

Each run produces two outputs under Logs/:

  • DebuggingLogs/DIAGNOSTIC_SCENARIO_RUN_<timestamp> — structured log
  • Chats/DIAGNOSTIC_SCENARIO_RUN_<timestamp>_CHAT.html — human-readable HTML transcript

Evaluation protocol outputs go to Logs/Evaluation/<AssistantType>/<run_id>/.


Reproducibility and variance

EvidenceKGOptimal: near-zero variance by design

The neuro-symbolic assistant is deterministic end-to-end: SPARQL queries, OWL reasoning, and the information-gain heuristic are pure functions of the knowledge graph and the observed evidence. The only stochastic element is the NER step, which maps a symptom description to a set of anomalous/nominal components. Its input is fixed per scenario (system description + symptom + component list), so it is cached — same input, same output, always.

One might ask whether the LLM assistant should be cached the same way. It should not: its LLM calls take a growing conversation history as input, which is unique per trajectory because each action outcome differs. Cache hits would be negligible. More fundamentally, the LLM assistant's variance is not localised to one upstream step — it is distributed across every call in the session. The two architectures are qualitatively different in this regard.

A further point: the NER step is not inherently tied to an LLM. Classifying a symptom description against a closed, known set of components (provided by the KG) is a well-defined task that could be solved with rule-based matching, dictionary lookup, or a fine-tuned deterministic model. The LLM is used here for convenience and generality, not necessity — which makes caching a natural and principled choice rather than a workaround.

Why the nl_interface is not cached

The natural language interface has two LLM calls per action: _parse (free text → action list) and _verbalize (action results → narrative). Caching them would require stable inputs across runs. For the LLM assistant, both inputs depend on the assistant's own free-text suggestions, which are non-deterministic — so cache hits would be negligible. For the neuro-symbolic assistant, _parse inputs are more stable, but _verbalize inputs include the raw simulation results (voltages, resistances, etc.) which vary. Since the nl_interface must work correctly for all assistants, and at least one assistant (LLM) makes caching ineffective, we do not cache the nl_interface.

What we tried before settling on caching

We attempted to make the NER step deterministic through API parameters: temperature=0.0, top_p=1.0, seed=42, switching from the Responses API to Chat Completions (which actually accepts seed), and pinning a specific upstream provider through OpenRouter. None of this fully worked. OpenAI's documentation states that seed offers only "best effort" reproducibility. In practice we observed different NER outputs for identical inputs across all combinations tested. Caching is the only solution that actually works.


Video

A demo video is available in the releases and on YouTube.

About

Diagnostic Assistant Demo for ESWC 2026

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages