Codebase associated with the ESWC 2026 demo paper "A Prototyping Environment for Comparing Neuro-symbolic Diagnostic Assistants".
The environment runs diagnostic scenarios: a fault is injected into a small simulated electrical system by a Saboteur agent, and a Diagnostic Assistant guides a Service Agent toward identifying the root cause. The goal is to compare different assistant strategies — LLM-based, neuro-symbolic, and baselines — under controlled conditions.
ESWC_2026_Demo/
├── environment_classes.py # Abstract agent roles and scenario orchestrator
├── configuration.py # Configuration dataclass and defaults
├── run_diagnostic_scenario.py # Entry point: run a single diagnostic scenario
├── run_evaluation_protocol.py # Entry point: run the Adaptive Evaluation Protocol
├── run_many_scenarios.py # Batch runner for multiple scenarios
├── make_tables.py # Aggregate evaluation results into tables
├── SCENARIOS_MASTER.csv # Master list of all 135 scenarios with metadata
│
├── Implementations/ # Concrete agent implementations
│ ├── diagnosticAssistantLLM.py # Monolithic LLM assistant
│ ├── diagnosticAssistantEvidenceKGOptimal.py # Neuro-symbolic assistant
│ ├── diagnosticAssistantRandomSearch.py # Random baseline (components one-by-one)
│ ├── diagnosticAssistantFixedRandomTrajectories.py # Fixed random baseline
│ ├── diagnosticAssistantUnhelpful.py # Always-wrong baseline
│ ├── fault_injections.py # SCENARIOS list (135 injectable faults)
│ ├── saboteur*.py # Saboteur implementations
│ └── serviceAgent*.py # Service agent implementations
│
├── Evaluation/ # Adaptive Evaluation Protocol
│ ├── protocol.py # Main protocol loop (HDBSCAN + ARI convergence)
│ ├── clustering.py # Intent & execution clustering
│ ├── qualitative.py # Qualitative rubric evaluation (LLM-based)
│ ├── report.py # Per-scenario report formatting
│ ├── metrics.py # Numerical metrics (success rate, cost, actions)
│ └── checkpoint.py # Checkpoint save/load
│
├── Knowledge_sources/
│ ├── Structured_knowledge_sources/ # OWL ontology (TBox) + per-system ABox KGs
│ └── Unstructured_knowledge_sources/ # Textual descriptions + circuit schematics
│
├── Utilities/ # Shared helpers
│ ├── OWL_reasoning.py # HermiT reasoner wrapper (owlready2)
│ ├── retrieval.py # RAG: chunking, embedding, top-k retrieval
│ ├── topology.py # Graph algorithms over the KG
│ ├── caching.py # SQLite-backed LLM call cache
│ ├── chat_log.py # HTML chat transcript writer
│ ├── agents_boilerplate.py # OpenAI Agents SDK setup helpers
│ └── ...
│
├── Tests/ # pytest test suite
├── static/client.html # Browser voice interface
├── voice_server.py # FastAPI voice server (Whisper STT + TTS)
└── voice_client.py # Voice client helpers
Five simulated electrical systems are supported, all implemented as SPICE circuits via the diagnosable-systems-simulation package:
| System | --system flag |
Scenarios |
|---|---|---|
| 3-module cube circuit | 3CubesSystem |
1–25 |
| 10-module cube circuit | 10CubesSystem |
26–50 |
| Asymmetric chains | AsymmetricChainsSystem |
51–75 |
| Ambient light sensor | AmbientLightSensorSystem |
76–105 |
| Current sensor | CurrentSensorSystem |
106–135 |
--assistant flag |
Class | Description |
|---|---|---|
LLM |
DiagnosticAssistantLLM |
Monolithic GPT agent with RAG and vision |
EvidenceKGOptimal |
DiagnosticAssistantEvidenceKGOptimal |
Neuro-symbolic: OWL reasoning + SPARQL + information-gain heuristic |
RandomSearch |
DiagnosticAssistantRandomSearch |
Random baseline: hypothesises components one-by-one without replacement |
FixedRandomTrajectories |
DiagnosticAssistantFixedRandomTrajectories |
Fixed random action sequences (reproducible baseline) |
Unhelpful |
DiagnosticAssistantUnhelpful |
Always-wrong baseline |
The neuro-symbolic assistant pipeline is summarised below:
pip install -r requirements.txtNote:
torchis a heavy dependency pulled in bysentence-transformers. If you only want to run scenarios without the evaluation protocol, you can skip it.
The SPICE simulation backend is a separate package:
git clone https://github.com/kataph/diagnosable-systems-simulation
pip install -e diagnosable-systems-simulation/export OPENAI_API_KEY="sk-..." # OpenAI, or
export OPENAI_BASE_URL="https://openrouter.ai/api/v1" # any OpenAI-compatible endpoint
export OPENAI_API_KEY="sk-or-..." # OpenRouter keypython -m run_diagnostic_scenario --help3-cubes system, LLM assistant, CLI:
python -m run_diagnostic_scenario \
--text-input-file "Knowledge_sources/Unstructured_knowledge_sources/3_cubes/3_cubes_description.txt" \
--diagram "Knowledge_sources/Unstructured_knowledge_sources/3_cubes/3_cubes_schematics.png" \
--kg "Knowledge_sources/Structured_knowledge_sources/3_cubes/zorro-ontology-3-cubes-abox.ttl" \
--ontology "Knowledge_sources/Structured_knowledge_sources/zorro-ontology-tbox.ttl" \
--retrieval-folder "Knowledge_sources/Unstructured_knowledge_sources/3_cubes" \
--system 3CubesSystem --saboteur Human --service Human --assistant LLM --interface cli10-cubes system, neuro-symbolic assistant:
python -m run_diagnostic_scenario \
--text-input-file "Knowledge_sources/Unstructured_knowledge_sources/10_cubes/10_cubes_description.txt" \
--diagram "Knowledge_sources/Unstructured_knowledge_sources/10_cubes/10_cubes_schematics.png" \
--kg "Knowledge_sources/Structured_knowledge_sources/10_cubes/zorro-ontology-10-cubes-abox.ttl" \
--ontology "Knowledge_sources/Structured_knowledge_sources/zorro-ontology-tbox.ttl" \
--retrieval-folder "Knowledge_sources/Unstructured_knowledge_sources/10_cubes" \
--system 10CubesSystem --saboteur Human --service Human --assistant EvidenceKGOptimal --interface cliVoice interface (requires running uvicorn separately):
uvicorn voice_server:app --host 0.0.0.0 --port 8000
# Then browse to http://127.0.0.1:8000/client.htmlThe Adaptive Evaluation Protocol collects batches of trajectories per scenario and checks for convergence of the trajectory clustering (via inter-batch ARI).
python -m run_evaluation_protocol --helpExample (automated, all scenarios):
python -m run_evaluation_protocol \
--all-scenarios \
--assistant EvidenceKGOptimal --service SpiceSim \
--service-config '{"model":"openai/gpt-4.1-mini"}' \
--assistant-config '{"model":"openai/gpt-4.1","embed_model":"openai/text-embedding-3-small"}' \
--labeling-model openai/gpt-4.1-mini \
--batch-size 10 --rounds 10 --max-concurrent 5Incremental runs (manual convergence control):
# First batch
python -m run_evaluation_protocol --scenarios 1,2,3,4,5 \
--assistant EvidenceKGOptimal --service SpiceSim \
[flags] --batch-size 10 --max-batches 1
# note the run_id printed at startup
# Add another batch
python -m run_evaluation_protocol --scenarios 1,2,3,4,5 \
--resume <run_id> [flags] --batch-size 10 --max-batches 2
# Add more scenarios to the same run
python -m run_evaluation_protocol --scenarios 6,7,8,9,10 \
--resume <run_id> [flags] --batch-size 10 --max-batches 1Monitor progress:
tail -f Logs/Evaluation/<AssistantType>/<run_id>/protocol.log# Auto-detect latest run per assistant:
python make_tables.py
# Specify runs explicitly:
python make_tables.py --run EvidenceKGOptimal:<run_id> --run LLM:<run_id>Produces tables broken down by scenario, system, and fault type, including qualitative analysis aggregates. LaTeX is auto-copied to clipboard if pyperclip is installed.
pytest Tests/
pytest Tests/test_scenario1_spice.py -s # single file with outputEach diagnostic action and hypothesis verification incurs a cost recorded in the trajectory. The cost model has two tiers depending on whether SpiceSim is the service agent.
| Service agent | Cost source | Values |
|---|---|---|
SpiceSim |
precise_action_cost — actual simulation time in minutes |
Varies per action type (e.g. observe_component=10, measure_voltage=20, replace_component=120) |
Human, LLM, Mock |
ACTION_COST_MAP fallback |
Observe=1, Test=2, Adjust=4, Replace=12 (qualitative ordering only) |
Both values are logged side-by-side as kg_cost (fallback) and sim_time (precise) in the debugging log.
Hypothesis verification cost is also computed differently per service agent:
| Service agent | Correct/partial hypothesis | Wrong hypothesis |
|---|---|---|
SpiceSim |
apply_repairs() — sum of actual per-component repair costs: cable reconnection = 10 per port, component replacement = 120 |
_estimate_repair_cost() — same per-type rates applied to the attempted components, without mutating state |
Human, LLM, Mock |
FALLBACK_HYPOTHESIS_VERIFICATION_COST = 120 flat |
Same flat cost |
The SpiceSim costs scale naturally with the hypothesis: naming two crossed cables costs 20 (two reconnections at 10 each), while naming one degraded component costs 120. This means MODE 2 hypothesis testing is cost-equivalent to performing the same repairs via MODE 1 actions, with no arbitrage opportunity.
A session ends successfully either because the assistant submitted a correct hypothesis (the hypothesis path) or because a direct repair action restored the system without any hypothesis (the action path). The correctness check differs substantially between the two.
Called when DiagnosticAssistant.suggest_action() returns a DiagnosticFaultHypothesis. The service agent calls system.test_repair(component_ids), which:
- Resets the circuit to the fault snapshot (via
restore_snapshot). - Applies the hypothesised repairs to the live circuit (reconnects cables, clears fault overlays, sets
is_rotated=TrueforPhysicalEnclosuretargets). - Re-simulates and checks
is_system_nominal(). - Restores the circuit back to the fault state before returning, so the test is non-destructive.
If test_repair returns True, the session ends with end=success.
Called for every ordinary diagnostic action. The service agent uses nl_interface to execute the action on the live circuit, then calls decide_finish to check whether the system has been restored.
The complication: after executing each batch of actions, ServiceAgentSpiceSim calls restore_snapshot to undo any transient side-effects (e.g. a switch left open, a measurement that perturbed state) before decide_finish re-simulates. This reset must not undo actions that were genuine repairs — otherwise decide_finish would see the broken circuit again and fail to detect success.
To distinguish repair actions from transient diagnostic actions, the system uses two layers:
-
Always-permanent actions (
replace_component,reconnect_cable): these remove a fault overlay or reconnect a disconnected cable. They are always tracked in_repaired_comp_idsand excluded from the snapshot reset, regardless of the scenario. -
Scenario-specific permanent actions: some actions are only repairs in specific scenarios. For example,
rotate_enclosureoncube_ctrlis a repair in scenarios 101–105 (optical feedback loop fault), but merely a diagnostic action in any other scenario.reverse_batteryis a repair only when the fault is a reversed battery. These are declared per-scenario in_PERMANENT_REPAIRSinsidefault_injections.pyas afrozenset[tuple[action_id, subject_component_id]]and stored on eachScenarioaspermanent_repair_actions.
When execute_action completes, it checks each executed action against both layers and adds qualifying subject component IDs to _repaired_comp_ids. restore_snapshot(exclude_ids=_repaired_comp_ids) then preserves those components in their repaired state. decide_finish re-simulates and calls is_system_nominal() — if True, the session ends with end=success_no_hypothesis.
Adding a new scenario with a novel repair action type: add a _PERMANENT_REPAIRS[scenario_number] entry in fault_injections.py with the (action_id, subject_component_id) pairs for that scenario's solution. Do not add the action to the always-permanent set unless it unconditionally removes a fault overlay regardless of scenario.
Tests/test_all_scenarios.py mirrors this two-path logic via conftest.run_sequence, which applies the same restore_snapshot + exclude_ids logic using the scenario's permanent_repair_actions. Tests that use RotateEnclosure, ReverseBattery, or SwapCablePolarities as their repair will fail if the corresponding entry is missing from _PERMANENT_REPAIRS — this is intentional: the tests catch B3-type bugs where a repair action is reset by restore_snapshot before decide_finish can detect success.
Each trajectory is closed with one of five end values written to the JSON file:
| Value | Meaning | Counted in statistics? |
|---|---|---|
success |
Assistant submitted a correct hypothesis | Yes — success |
success_no_hypothesis |
System restored via a direct repair action (no hypothesis submitted) | Yes — success |
timeout |
Action budget exhausted without restoring the system | Yes — failure |
surrender |
Assistant returned None (no further suggestions) |
Yes — failure |
llm_truncation |
Session terminated by a truncated/unparseable LLM response or an API network error | Excluded from all statistics |
Rationale for excluding llm_truncation: the cause is ambiguous. It can be a genuine model schema-compliance failure, a token-budget limit hit during a long reasoning chain, or a transient network/server error — all three produce the same observable outcome. Including these as failures would conflate infrastructure reliability with diagnostic competence and penalise models that happened to hit more API errors. Reported success rates are therefore a lower bound on true diagnostic ability given a working API call.
Each run produces two outputs under Logs/:
DebuggingLogs/DIAGNOSTIC_SCENARIO_RUN_<timestamp>— structured logChats/DIAGNOSTIC_SCENARIO_RUN_<timestamp>_CHAT.html— human-readable HTML transcript
Evaluation protocol outputs go to Logs/Evaluation/<AssistantType>/<run_id>/.
The neuro-symbolic assistant is deterministic end-to-end: SPARQL queries, OWL reasoning, and the information-gain heuristic are pure functions of the knowledge graph and the observed evidence. The only stochastic element is the NER step, which maps a symptom description to a set of anomalous/nominal components. Its input is fixed per scenario (system description + symptom + component list), so it is cached — same input, same output, always.
One might ask whether the LLM assistant should be cached the same way. It should not: its LLM calls take a growing conversation history as input, which is unique per trajectory because each action outcome differs. Cache hits would be negligible. More fundamentally, the LLM assistant's variance is not localised to one upstream step — it is distributed across every call in the session. The two architectures are qualitatively different in this regard.
A further point: the NER step is not inherently tied to an LLM. Classifying a symptom description against a closed, known set of components (provided by the KG) is a well-defined task that could be solved with rule-based matching, dictionary lookup, or a fine-tuned deterministic model. The LLM is used here for convenience and generality, not necessity — which makes caching a natural and principled choice rather than a workaround.
The natural language interface has two LLM calls per action: _parse (free text → action list) and _verbalize (action results → narrative). Caching them would require stable inputs across runs. For the LLM assistant, both inputs depend on the assistant's own free-text suggestions, which are non-deterministic — so cache hits would be negligible. For the neuro-symbolic assistant, _parse inputs are more stable, but _verbalize inputs include the raw simulation results (voltages, resistances, etc.) which vary. Since the nl_interface must work correctly for all assistants, and at least one assistant (LLM) makes caching ineffective, we do not cache the nl_interface.
We attempted to make the NER step deterministic through API parameters: temperature=0.0, top_p=1.0, seed=42, switching from the Responses API to Chat Completions (which actually accepts seed), and pinning a specific upstream provider through OpenRouter. None of this fully worked. OpenAI's documentation states that seed offers only "best effort" reproducibility. In practice we observed different NER outputs for identical inputs across all combinations tested. Caching is the only solution that actually works.

