diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..964d0b4 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,103 @@ +# Project Configuration + +## Behavioral Rules + +- Do what has been asked; nothing more, nothing less +- NEVER create files unless absolutely necessary for the goal +- ALWAYS prefer editing an existing file to creating a new one +- NEVER proactively create documentation files unless explicitly requested +- NEVER save working files, tests, or docs to the root folder +- ALWAYS read a file before editing it +- Keep files under 500 lines +- NEVER commit secrets, credentials, or .env files + +## File Organization + +- Use `src/` for source code +- Use `tests/` for test files +- Use `docs/` for documentation +- Use `scripts/` for utility scripts and orchestration scripts +- Use `config/` for configuration files + +## Parallelism — ALWAYS parallel by default + +- EVERY task must be analyzed for parallelism BEFORE execution +- Batch ALL related file reads in ONE message +- Batch ALL file edits in ONE message +- Batch ALL independent Bash commands in ONE message +- Spawn ALL independent Agent calls in ONE message with `run_in_background: true` +- After spawning background agents, STOP and wait for results — do NOT poll +- When a task has multiple independent fix targets, spawn one Agent per target in a single message +- When reviewing results from parallel agents, read ALL results before deciding next action +- Sequential steps run only when there is a true data dependency on a prior step + +## Closed-Loop Execution + +When a workflow specifies a loop (repeat-until-pass), follow this protocol: + +1. **Read the workflow** to identify: steps, pass condition, max iterations, and what to capture per iteration +2. **Run the check/capture step** to establish baseline metrics +3. **Analyze results** — categorize issues, group by fix type +4. **Spawn parallel fix agents** — one Agent per independent issue category, ALL in one message +5. **Wait for all agents** — review ALL results together +6. **Re-run the check** — compare metrics to previous iteration +7. **Log iteration** — append to `context/MEMORY.md`: iteration number, pass/fail counts, key fixes, regressions +8. **Decide**: + - All pass → exit loop, run final confirmation + - Regression detected → revert, log what failed, try different approach + - Issues remain and under max iterations → go to step 3 + - Max iterations reached → stop, report remaining issues +9. **On exit** — write final summary to `context/MEMORY.md` + +### Regression handling +- If an iteration produces MORE issues than the previous one, it is a regression +- Revert the changes from that iteration immediately +- Log what was attempted and why it regressed +- Try a different fix approach in the next iteration +- Never repeat the same fix that caused a regression + +## Security + +- NEVER hardcode API keys, secrets, or credentials in source files +- NEVER commit .env files or any file containing secrets +- Always validate user input at system boundaries +- Always sanitize file paths to prevent directory traversal + +## Memory Protocol + +- At session start, read `context/MEMORY.md` for ongoing context +- Before session ends, update `context/MEMORY.md` with progress and findings +- Log important architectural decisions in `context/DECISIONS.md` +- Check `context/CONVENTIONS.md` for project-specific patterns before writing code +- During loops, append iteration results to `context/MEMORY.md` after each iteration + +## Orchestration Scripts + +- Orchestration scripts live in `scripts/` and automate multi-step pipelines +- Scripts should be idempotent — safe to re-run from any iteration +- Scripts must accept `--iteration N` to resume from a specific point +- Scripts must write machine-readable output (JSON) for Claude to parse +- Scripts must exit with code 0 on success, non-zero on failure +- Use `scripts/orchestrate.sh` as the template for new orchestration scripts + +## Workflows + +When the task matches a common pattern, follow the corresponding workflow: + +- Bug fixes: follow `workflows/bug-fix.md` +- New features: follow `workflows/new-feature.md` +- Refactoring: follow `workflows/refactor.md` +- Code reviews: follow `workflows/code-review.md` +- Closed-loop QA: follow `workflows/closed-loop.md` + +## Skills + +Use these agent profiles when the task calls for a specialized role: + +- `/planner` — Task decomposition, parallel execution planning +- `/implementer` — Writing production code +- `/reviewer` — Code review with severity ratings +- `/debugger` — Systematic bug investigation +- `/refactorer` — Safe code restructuring +- `/architect` — System design and architecture decisions +- `/orchestrator` — Design and run closed-loop pipelines diff --git a/JavaDuckerMcpServer.java b/JavaDuckerMcpServer.java index eb95254..7ac6d0d 100644 --- a/JavaDuckerMcpServer.java +++ b/JavaDuckerMcpServer.java @@ -111,6 +111,58 @@ public static void main(String[] args) throws Exception { "to monitor bulk ingestion progress.", "{}"), (ex, a) -> call(JavaDuckerMcpServer::stats)) + .tool( + tool("javaducker_summarize", + "Get a structural summary of an indexed file: class names, method names, imports, " + + "line count. One-call overview without reading the full text.", + schema(props( + "artifact_id", str("Artifact ID to summarize")), + "artifact_id")), + (ex, a) -> call(() -> summarize((String) a.get("artifact_id")))) + .tool( + tool("javaducker_map", + "Get a project map showing directory structure, file counts, largest files, and " + + "recently indexed files. Use for codebase orientation.", + "{}"), + (ex, a) -> call(JavaDuckerMcpServer::projectMap)) + .tool( + tool("javaducker_stale", + "Check which indexed files are stale (modified on disk since last indexing). " + + "Accepts file_paths (list of absolute paths) or git_diff_ref (e.g. HEAD~3) to auto-detect changed files.", + schema(props( + "file_paths", str("JSON array of absolute file paths to check (optional if git_diff_ref given)"), + "git_diff_ref", str("Git ref for diff, e.g. HEAD~3 or main (optional if file_paths given)")))), + (ex, a) -> call(() -> checkStale( + (String) a.getOrDefault("file_paths", ""), + (String) a.getOrDefault("git_diff_ref", "")))) + .tool( + tool("javaducker_dependencies", + "Get the import/dependency list for an indexed file. Shows what this file imports " + + "and which indexed artifacts those imports resolve to.", + schema(props( + "artifact_id", str("Artifact ID to get dependencies for")), + "artifact_id")), + (ex, a) -> call(() -> dependencies((String) a.get("artifact_id")))) + .tool( + tool("javaducker_dependents", + "Find which indexed files import/depend on this file. Useful for impact analysis.", + schema(props( + "artifact_id", str("Artifact ID to find dependents of")), + "artifact_id")), + (ex, a) -> call(() -> dependents((String) a.get("artifact_id")))) + .tool( + tool("javaducker_watch", + "Start or stop auto-indexing a directory. When watching, file changes are " + + "automatically detected and re-indexed. Use action=start with a directory, or action=stop.", + schema(props( + "action", str("start or stop"), + "directory", str("Absolute path to watch (required for start)"), + "extensions", str("Comma-separated extensions, e.g. .java,.xml,.md (optional)")), + "action")), + (ex, a) -> call(() -> watch( + (String) a.get("action"), + (String) a.getOrDefault("directory", ""), + (String) a.getOrDefault("extensions", "")))) .build(); } @@ -204,6 +256,84 @@ static Map stats() throws Exception { return httpGet("/stats"); } + static Map summarize(String artifactId) throws Exception { + Map r = httpGet("/summary/" + artifactId); + if (r == null) throw new RuntimeException("Artifact not found or no summary available: " + artifactId); + return r; + } + + static Map projectMap() throws Exception { + return httpGet("/map"); + } + + @SuppressWarnings("unchecked") + static Map checkStale(String filePathsJson, String gitDiffRef) throws Exception { + List paths = new ArrayList<>(); + + // If git_diff_ref is given, run git diff to get file paths + if (gitDiffRef != null && !gitDiffRef.isBlank()) { + ProcessBuilder pb = new ProcessBuilder("git", "diff", "--name-only", gitDiffRef); + pb.directory(Path.of(PROJECT_ROOT).toFile()); + pb.redirectErrorStream(true); + Process proc = pb.start(); + String output = new String(proc.getInputStream().readAllBytes()).trim(); + proc.waitFor(); + if (!output.isEmpty()) { + Path root = Path.of(PROJECT_ROOT).toAbsolutePath(); + for (String line : output.split("\n")) { + paths.add(root.resolve(line.trim()).toString()); + } + } + } + + // If file_paths is given, parse it + if (filePathsJson != null && !filePathsJson.isBlank()) { + try { + List parsed = MAPPER.readValue(filePathsJson, List.class); + paths.addAll(parsed); + } catch (Exception e) { + // Try as comma-separated + for (String p : filePathsJson.split(",")) { + if (!p.isBlank()) paths.add(p.trim()); + } + } + } + + if (paths.isEmpty()) { + throw new RuntimeException("Provide file_paths or git_diff_ref"); + } + + return httpPost("/stale", Map.of("file_paths", paths)); + } + + static Map watch(String action, String directory, String extensions) throws Exception { + if ("stop".equalsIgnoreCase(action)) { + return httpPost("/watch/stop", Map.of()); + } + if ("start".equalsIgnoreCase(action)) { + Map body = new LinkedHashMap<>(); + body.put("directory", directory); + if (extensions != null && !extensions.isBlank()) body.put("extensions", extensions); + return httpPost("/watch/start", body); + } + if ("status".equalsIgnoreCase(action)) { + return httpGet("/watch/status"); + } + throw new RuntimeException("Unknown action: " + action + ". Use start, stop, or status."); + } + + static Map dependencies(String artifactId) throws Exception { + Map r = httpGet("/dependencies/" + artifactId); + if (r == null) throw new RuntimeException("Artifact not found: " + artifactId); + return r; + } + + static Map dependents(String artifactId) throws Exception { + Map r = httpGet("/dependents/" + artifactId); + if (r == null) throw new RuntimeException("Artifact not found: " + artifactId); + return r; + } + // ── HTTP helpers ────────────────────────────────────────────────────────── static Map httpGet(String path) throws Exception { diff --git a/context/CONVENTIONS.md b/context/CONVENTIONS.md new file mode 100644 index 0000000..5ca5eba --- /dev/null +++ b/context/CONVENTIONS.md @@ -0,0 +1,14 @@ +# Project Conventions + + + +## Naming + + +## Imports + + +## Error Handling + + +## Testing diff --git a/context/DECISIONS.md b/context/DECISIONS.md new file mode 100644 index 0000000..8566392 --- /dev/null +++ b/context/DECISIONS.md @@ -0,0 +1,8 @@ +# Architecture Decisions + + diff --git a/context/MEMORY.md b/context/MEMORY.md new file mode 100644 index 0000000..052ad50 --- /dev/null +++ b/context/MEMORY.md @@ -0,0 +1,27 @@ +# Session Memory + +## Current Focus +All 8 Claude Companion features implemented and tests passing (65/65). + +## Recent Decisions +- DuckDB UPDATE with PK can fail (ART index bug) — use DELETE+INSERT pattern for artifact reindex +- ALTER TABLE in SchemaBootstrap needs separate Statement objects (error closes shared stmt in DuckDB) +- chunk_embeddings cleanup requires subquery (keyed by chunk_id, not artifact_id) + +## Key Findings +- All 8 features from plans/claude-companion-features.md implemented in one session +- 7 new Java classes created, 10+ existing classes modified +- 6 new MCP tools added (summarize, map, stale, dependencies, dependents, watch) +- 2 existing MCP tools enhanced (search with line numbers, index with incremental re-indexing) + +## Open Questions +- HNSW index is not auto-built on startup — needs explicit buildHnswIndex() call +- Watch mode FileWatcher uses polling WatchService which may miss rapid changes on some OS + +## Session Log +- 2026-03-28: Implemented all 8 features from claude-companion-features.md plan + - Phase 1 (parallel): F1 Line Numbers, F3 File Summaries, F4 Project Map, F6 Diff-Aware Search + - Phase 2 (parallel): F2 Incremental Re-indexing, F5 Dependency Graph + - Phase 3+4 (parallel): F7 Watch Mode, F8 HNSW Index + - Fixed test compilation (3 test files), DuckDB PK constraint bug, Statement closure bug + - All 65 tests green diff --git a/plans/claude-companion-features.md b/plans/claude-companion-features.md new file mode 100644 index 0000000..2d0142d --- /dev/null +++ b/plans/claude-companion-features.md @@ -0,0 +1,333 @@ +# JavaDucker: Claude Companion Features — Multi-Agent Implementation Plan + +## Goal +Make JavaDucker a better companion for Claude Code when working with large codebases. All features must be **pure Java**, no new non-Java dependencies, everything local. + +## Constraint +- Java only — no Python, no Node.js, no external embedding APIs +- No new Maven dependencies — use what's already in pom.xml (Spring Boot, DuckDB, POI, etc.) +- All processing local — no network calls to AI services +- Existing interfaces must not break + +--- + +## Features (priority order) + +### Feature 1: Line Numbers in Search Results +### Feature 2: Incremental Re-indexing (replace stale artifacts) +### Feature 3: File Summaries (auto-generated digest per file) +### Feature 4: Project Map (high-level codebase overview) +### Feature 5: Dependency/Import Graph +### Feature 6: Diff-Aware Search (git-changed file detection) +### Feature 7: Watch Mode (auto-index on file change) +### Feature 8: ANN Indexing (HNSW for scale) + +--- + +## Feature 1: Line Numbers in Search Results + +**Why**: Claude needs `file:line` to do `Read(file, offset=N)`. Without line numbers, Claude has to search again after finding a chunk. + +**Changes**: + +| File | Change | +|------|--------| +| `Chunker.java` | Track line_start and line_end per chunk (count `\n` in text before charStart/charEnd) | +| `SchemaBootstrap.java` | Add `line_start INTEGER, line_end INTEGER` columns to `artifact_chunks` | +| `IngestionWorker.java` | Pass line numbers when inserting chunks | +| `SearchService.java` | Include `line_start`, `line_end` in result maps | +| `JavaDuckerMcpServer.java` | Include line numbers in search result formatting | + +**Parallel group**: All file changes are independent — one agent per file. + +**Agent plan**: +``` +Agent 1: Chunker.java — add line counting to chunk() method +Agent 2: SchemaBootstrap.java — add columns (with ALTER TABLE migration for existing DBs) +Agent 3: SearchService.java — add line_start/line_end to all search result maps +Agent 4: IngestionWorker.java — pass line numbers during chunk insert +Agent 5: JavaDuckerMcpServer.java — format line numbers in search tool output +``` + +**Verification**: Index a file, search for a known function, confirm line numbers match. + +--- + +## Feature 2: Incremental Re-indexing + +**Why**: Currently re-indexing a changed file creates a duplicate. Need to detect "same file, new content" and replace. + +**Changes**: + +| File | Change | +|------|--------| +| `UploadService.java` | On duplicate path detection: delete old chunks/embeddings/text, reset status to RECEIVED, update sha256 and size | +| `SchemaBootstrap.java` | Add `original_client_path` index if not exists (already has one — verify) | +| `ArtifactService.java` | Add `deleteArtifactData(artifactId)` — deletes from artifact_chunks, chunk_embeddings, artifact_text | + +**Sequential**: UploadService depends on ArtifactService having the delete method. + +**Agent plan**: +``` +Group 1 (parallel): + Agent A: ArtifactService.java — add deleteArtifactData() method + Agent B: SchemaBootstrap.java — verify indexes exist + +Group 2 (after Group 1): + Agent C: UploadService.java — modify upload() to replace existing artifact on same path +``` + +**Verification**: Index a file, modify it, re-index same path, confirm old chunks gone and new chunks present. + +--- + +## Feature 3: File Summaries + +**Why**: Claude can understand the codebase at a glance without reading every file. One-call overview per file. + +**Approach**: Generate a summary during ingestion using the extracted text — no LLM needed. Extract: +- File type and language +- Class/interface/function names (regex-based extraction for Java, JS, Python, etc.) +- Import statements (first 10) +- Line count +- Top terms (from TF-IDF embedding — highest-weight tokens) + +**Changes**: + +| File | Change | +|------|--------| +| `FileSummarizer.java` (NEW) | Extract structural summary from text: class names, method names, imports, line count, top terms | +| `SchemaBootstrap.java` | Add `artifact_summaries` table: artifact_id, summary_text, class_names, method_names, import_count, line_count | +| `IngestionWorker.java` | After CHUNKED stage, generate summary and insert | +| `ArtifactService.java` | Add `getSummary(artifactId)` method | +| `JavaDuckerRestController.java` | Add `GET /api/summary/{artifactId}` endpoint | +| `JavaDuckerMcpServer.java` | Add `javaducker_summarize` tool | + +**Agent plan**: +``` +Group 1 (parallel): + Agent A: FileSummarizer.java — new class, regex-based extraction for Java/JS/Python/Go/Rust + Agent B: SchemaBootstrap.java — add artifact_summaries table + Agent C: JavaDuckerMcpServer.java — add javaducker_summarize tool (calls /api/summary/{id}) + +Group 2 (after Group 1): + Agent D: IngestionWorker.java — call FileSummarizer after chunking + Agent E: ArtifactService.java + REST controller — add getSummary() and endpoint +``` + +**Verification**: Index a Java file, call summarize, confirm class names and method names appear. + +--- + +## Feature 4: Project Map + +**Why**: Claude gets a mental map of the codebase in one call — directory structure, file counts, most-connected files, recently changed files. + +**Approach**: Query DuckDB for all indexed artifacts, group by directory path, count files per directory, identify largest files, most recently indexed. + +**Changes**: + +| File | Change | +|------|--------| +| `ProjectMapService.java` (NEW) | Query artifacts grouped by directory prefix, return tree structure with counts | +| `JavaDuckerRestController.java` | Add `GET /api/map` endpoint | +| `JavaDuckerMcpServer.java` | Add `javaducker_map` tool | + +**Agent plan**: +``` +All parallel (independent): + Agent A: ProjectMapService.java — new service, DuckDB queries + Agent B: REST controller — add /api/map endpoint + Agent C: MCP server — add javaducker_map tool +``` + +**Verification**: Index a directory, call map, confirm directory tree with file counts. + +--- + +## Feature 5: Dependency/Import Graph + +**Why**: Claude can trace call chains and understand what depends on what — the highest-value feature. + +**Approach**: Parse import/require/include statements from source code during ingestion. Store as edges in a graph table. Query for callers/callees. + +**Changes**: + +| File | Change | +|------|--------| +| `ImportParser.java` (NEW) | Regex-based import extraction for Java (`import`), JS/TS (`import`/`require`), Python (`import`/`from`), Go (`import`), Rust (`use`) | +| `SchemaBootstrap.java` | Add `artifact_imports` table: artifact_id, import_statement, resolved_artifact_id (nullable) | +| `IngestionWorker.java` | After PARSING, extract imports and insert | +| `DependencyService.java` (NEW) | Query import graph: `getDependencies(artifactId)`, `getDependents(artifactId)`, `getImportChain(from, to)` | +| `JavaDuckerRestController.java` | Add `GET /api/dependencies/{artifactId}` and `GET /api/dependents/{artifactId}` | +| `JavaDuckerMcpServer.java` | Add `javaducker_dependencies` and `javaducker_dependents` tools | + +**Agent plan**: +``` +Group 1 (parallel): + Agent A: ImportParser.java — new class, regex patterns per language + Agent B: SchemaBootstrap.java — add artifact_imports table + Agent C: DependencyService.java — new service, graph queries + +Group 2 (after Group 1): + Agent D: IngestionWorker.java — call ImportParser after text extraction + Agent E: REST controller + MCP server — add dependency endpoints and tools +``` + +**Import resolution**: Match import paths to indexed artifacts by `original_client_path`. E.g., `import com.javaducker.server.service.SearchService` → find artifact where path ends with `com/javaducker/server/service/SearchService.java`. Store `resolved_artifact_id` when found, NULL when external. + +**Verification**: Index the code-helper project itself, query dependencies of SearchService, confirm it lists EmbeddingService and DuckDBDataSource. + +--- + +## Feature 6: Diff-Aware Search + +**Why**: After code changes, Claude can ask "what indexed content is stale?" + +**Approach**: Run `git diff --name-only HEAD~N` or accept a file list, cross-reference against indexed artifacts by `original_client_path`. + +**Changes**: + +| File | Change | +|------|--------| +| `StalenessService.java` (NEW) | Accept file paths, query artifacts table, return which indexed artifacts are stale (file modified after `indexed_at`) | +| `JavaDuckerRestController.java` | Add `POST /api/stale` endpoint (accepts list of file paths) | +| `JavaDuckerMcpServer.java` | Add `javaducker_stale` tool (accepts `git_diff_ref` or `file_paths`) | + +**No git dependency in Java** — the MCP server runs `git diff --name-only` via ProcessBuilder and passes the file list to the REST endpoint. + +**Agent plan**: +``` +All parallel: + Agent A: StalenessService.java — query by paths, compare timestamps + Agent B: REST controller — add /api/stale endpoint + Agent C: MCP server — add javaducker_stale tool with git diff integration +``` + +**Verification**: Index project, modify a file (don't re-index), call stale, confirm that file appears. + +--- + +## Feature 7: Watch Mode + +**Why**: Auto-index changed files so the search index stays current without manual re-indexing. + +**Approach**: Use Java's `WatchService` (java.nio.file) to monitor directories. On file change, trigger re-index via UploadService. Requires Feature 2 (incremental re-indexing). + +**Changes**: + +| File | Change | +|------|--------| +| `FileWatcher.java` (NEW) | `WatchService`-based directory monitor, filters by extension, debounces (500ms), calls UploadService on change | +| `AppConfig.java` | Add `watchDirs` (list of paths), `watchEnabled` (boolean), `watchExtensions` (string) | +| `JavaDuckerServerApp.java` | Start FileWatcher as Spring bean if enabled | +| `JavaDuckerMcpServer.java` | Add `javaducker_watch` tool to start/stop watching a directory | +| `JavaDuckerRestController.java` | Add `POST /api/watch/start` and `POST /api/watch/stop` | + +**Depends on**: Feature 2 (incremental re-indexing must work first). + +**Agent plan**: +``` +Group 1 (parallel): + Agent A: FileWatcher.java — WatchService implementation with debounce + Agent B: AppConfig.java — add watch config properties + +Group 2 (after Group 1): + Agent C: ServerApp + REST + MCP — wire up start/stop watch endpoints +``` + +**Verification**: Start watch on a directory, modify a file, wait 5s, confirm it's automatically re-indexed. + +--- + +## Feature 8: HNSW Indexing (ANN) + +**Why**: Brute-force cosine works for ~10k chunks. HNSW gives O(log n) search for 100k+ chunks. + +**Approach**: Implement HNSW in pure Java. This is the most complex feature. The algorithm is well-documented and doesn't require external libraries. + +**Changes**: + +| File | Change | +|------|--------| +| `HnswIndex.java` (NEW) | Pure Java HNSW implementation: insert, search, serialize/deserialize. Parameters: M=16, efConstruction=200, efSearch=50 | +| `EmbeddingService.java` | Add methods to build and query HNSW index | +| `SearchService.java` | Replace brute-force loop with HNSW query when index is available, fallback to brute-force if not | +| `IngestionWorker.java` | After EMBEDDED, insert into HNSW index | +| `SchemaBootstrap.java` | Add `hnsw_state` table for serialized index persistence | + +**Agent plan**: This is a single-agent task — HNSW is tightly coupled and shouldn't be split. + +``` +Agent 1: HnswIndex.java — full implementation (insert, knn-search, serialization) +Agent 2 (after Agent 1): Wire into SearchService + IngestionWorker +``` + +**Verification**: Index 1000+ chunks, compare search results and latency between brute-force and HNSW. + +--- + +## Execution Order (respecting dependencies) + +``` +Phase 1 — Independent features (ALL PARALLEL): + ├── Feature 1: Line Numbers (5 agents) + ├── Feature 3: File Summaries (5 agents) + ├── Feature 4: Project Map (3 agents) + └── Feature 6: Diff-Aware Search (3 agents) + +Phase 2 — Depends on nothing but benefits from Phase 1: + ├── Feature 2: Incremental Re-indexing (3 agents, 2 groups) + └── Feature 5: Dependency Graph (5 agents, 2 groups) + +Phase 3 — Depends on Feature 2: + └── Feature 7: Watch Mode (3 agents, 2 groups) + +Phase 4 — Independent but complex, do last: + └── Feature 8: HNSW Index (2 agents, sequential) +``` + +## New MCP Tools Summary + +| Tool | Feature | Purpose | +|------|---------|---------| +| `javaducker_summarize` | 3 | One-paragraph file digest with class/method names | +| `javaducker_map` | 4 | Project directory tree with file counts | +| `javaducker_dependencies` | 5 | What does this file import? | +| `javaducker_dependents` | 5 | What files import this one? | +| `javaducker_stale` | 6 | Which indexed files have changed since last index? | +| `javaducker_watch` | 7 | Start/stop auto-indexing a directory | + +Existing tools enhanced: +| Tool | Feature | Enhancement | +|------|---------|-------------| +| `javaducker_search` | 1 | Results include `line_start`, `line_end` | +| `javaducker_index_file` | 2 | Replaces stale artifact instead of duplicating | +| `javaducker_index_directory` | 2 | Replaces stale artifacts instead of duplicating | +| `javaducker_search` | 8 | Uses HNSW for faster semantic search when available | + +## New Java Classes + +| Class | Feature | Lines (est.) | +|-------|---------|-------------| +| `FileSummarizer.java` | 3 | ~150 | +| `ProjectMapService.java` | 4 | ~80 | +| `ImportParser.java` | 5 | ~120 | +| `DependencyService.java` | 5 | ~100 | +| `StalenessService.java` | 6 | ~60 | +| `FileWatcher.java` | 7 | ~120 | +| `HnswIndex.java` | 8 | ~300 | + +## Modified Java Classes + +| Class | Features | Changes | +|-------|----------|---------| +| `Chunker.java` | 1 | Add line counting | +| `SchemaBootstrap.java` | 1, 2, 3, 5, 8 | New tables and columns | +| `IngestionWorker.java` | 1, 3, 5, 8 | Call summarizer, import parser, HNSW insert | +| `SearchService.java` | 1, 8 | Line numbers in results, HNSW search path | +| `ArtifactService.java` | 2, 3 | deleteArtifactData(), getSummary() | +| `UploadService.java` | 2 | Replace stale artifacts | +| `AppConfig.java` | 7 | Watch config properties | +| `JavaDuckerRestController.java` | 3, 4, 5, 6, 7 | New endpoints | +| `JavaDuckerMcpServer.java` | 1, 3, 4, 5, 6, 7 | New tools, enhanced search output | diff --git a/scripts/orchestrate.sh b/scripts/orchestrate.sh new file mode 100644 index 0000000..22a510b --- /dev/null +++ b/scripts/orchestrate.sh @@ -0,0 +1,116 @@ +#!/bin/bash +# drom-flow orchestration script template +# Copy and customize this for your project's pipeline. +# +# Usage: +# ./scripts/orchestrate.sh [--iteration N] [--max N] [--check-only] +# +# Output: +# Writes JSON report to ./reports/iteration-N.json +# Exit 0 = all pass, Exit 1 = issues remain, Exit 2 = error + +set -euo pipefail + +# --- Configuration (customize these) --- +CHECK_CMD="echo 'Override CHECK_CMD with your test/check command'" +REPORT_DIR="./reports" +MAX_ITERATIONS=10 +# ---------------------------------------- + +# Parse arguments +ITERATION=1 +CHECK_ONLY=false +while [[ $# -gt 0 ]]; do + case $1 in + --iteration) ITERATION="$2"; shift 2 ;; + --max) MAX_ITERATIONS="$2"; shift 2 ;; + --check-only) CHECK_ONLY=true; shift ;; + *) echo "Unknown arg: $1"; exit 2 ;; + esac +done + +mkdir -p "$REPORT_DIR" + +run_check() { + local iter=$1 + local report="$REPORT_DIR/iteration-${iter}.json" + local start_time=$(date +%s) + + echo "[orchestrate] Iteration $iter — running check..." + + # Run the check command, capture output + local exit_code=0 + local output + output=$(eval "$CHECK_CMD" 2>&1) || exit_code=$? + + local end_time=$(date +%s) + local duration=$((end_time - start_time)) + + # Write report + cat > "$report" </dev/null || echo "\"$output\"") +} +EOF + + echo "[orchestrate] Report written to $report (exit code: $exit_code, ${duration}s)" + return $exit_code +} + +compare_iterations() { + local prev="$REPORT_DIR/iteration-$(($1 - 1)).json" + local curr="$REPORT_DIR/iteration-$1.json" + + if [ ! -f "$prev" ]; then + echo "[orchestrate] No previous iteration to compare" + return 0 + fi + + local prev_exit=$(python3 -c "import json; print(json.load(open('$prev'))['exitCode'])" 2>/dev/null || echo "1") + local curr_exit=$(python3 -c "import json; print(json.load(open('$curr'))['exitCode'])" 2>/dev/null || echo "1") + + echo "[orchestrate] Previous exit: $prev_exit → Current exit: $curr_exit" + + if [ "$curr_exit" -gt "$prev_exit" ]; then + echo "[orchestrate] WARNING: Possible regression detected" + return 1 + fi + return 0 +} + +# --- Main --- + +if [ "$CHECK_ONLY" = true ]; then + run_check "$ITERATION" + exit $? +fi + +echo "[orchestrate] Starting closed loop: iteration $ITERATION, max $MAX_ITERATIONS" + +while [ "$ITERATION" -le "$MAX_ITERATIONS" ]; do + if run_check "$ITERATION"; then + echo "[orchestrate] ALL CHECKS PASSED at iteration $ITERATION" + exit 0 + fi + + if [ "$ITERATION" -gt 1 ]; then + if ! compare_iterations "$ITERATION"; then + echo "[orchestrate] Regression at iteration $ITERATION — stopping for review" + exit 1 + fi + fi + + echo "[orchestrate] Issues remain. Report: $REPORT_DIR/iteration-${ITERATION}.json" + echo "[orchestrate] Waiting for fixes before next iteration..." + # Script exits here — Claude reads the report, spawns fix agents, + # then re-runs: ./scripts/orchestrate.sh --iteration $((ITERATION+1)) + exit 1 + +done + +echo "[orchestrate] Max iterations ($MAX_ITERATIONS) reached" +exit 1 diff --git a/src/main/java/com/javaducker/server/config/AppConfig.java b/src/main/java/com/javaducker/server/config/AppConfig.java index 3d386df..a18d9fc 100644 --- a/src/main/java/com/javaducker/server/config/AppConfig.java +++ b/src/main/java/com/javaducker/server/config/AppConfig.java @@ -15,6 +15,8 @@ public class AppConfig { private int ingestionPollSeconds = 5; private int ingestionWorkerThreads = 4; private int maxSearchResults = 20; + private boolean watchEnabled = false; + private String watchExtensions = ".java,.xml,.md,.yml,.json,.txt"; public String getDbPath() { return dbPath; } public void setDbPath(String dbPath) { this.dbPath = dbPath; } @@ -39,4 +41,10 @@ public class AppConfig { public int getMaxSearchResults() { return maxSearchResults; } public void setMaxSearchResults(int maxSearchResults) { this.maxSearchResults = maxSearchResults; } + + public boolean isWatchEnabled() { return watchEnabled; } + public void setWatchEnabled(boolean watchEnabled) { this.watchEnabled = watchEnabled; } + + public String getWatchExtensions() { return watchExtensions; } + public void setWatchExtensions(String watchExtensions) { this.watchExtensions = watchExtensions; } } diff --git a/src/main/java/com/javaducker/server/db/SchemaBootstrap.java b/src/main/java/com/javaducker/server/db/SchemaBootstrap.java index ce193ba..48335ad 100644 --- a/src/main/java/com/javaducker/server/db/SchemaBootstrap.java +++ b/src/main/java/com/javaducker/server/db/SchemaBootstrap.java @@ -81,7 +81,9 @@ CREATE TABLE IF NOT EXISTS artifact_chunks ( chunk_index INTEGER NOT NULL, chunk_text VARCHAR NOT NULL, char_start BIGINT, - char_end BIGINT + char_end BIGINT, + line_start INTEGER, + line_end INTEGER ) """); @@ -114,6 +116,45 @@ ON artifacts (file_name, size_bytes) ON artifacts (original_client_path) """); + // Feature 1: Add line_start/line_end columns to existing DBs + try (Statement alter = conn.createStatement()) { + alter.execute("ALTER TABLE artifact_chunks ADD COLUMN line_start INTEGER"); + } catch (Exception ignored) {} + try (Statement alter = conn.createStatement()) { + alter.execute("ALTER TABLE artifact_chunks ADD COLUMN line_end INTEGER"); + } catch (Exception ignored) {} + + // Feature 3: File summaries table + stmt.execute(""" + CREATE TABLE IF NOT EXISTS artifact_summaries ( + artifact_id VARCHAR PRIMARY KEY, + summary_text VARCHAR, + class_names VARCHAR, + method_names VARCHAR, + import_count INTEGER, + line_count INTEGER + ) + """); + + // Feature 5: Dependency/import graph table + stmt.execute(""" + CREATE TABLE IF NOT EXISTS artifact_imports ( + artifact_id VARCHAR NOT NULL, + import_statement VARCHAR NOT NULL, + resolved_artifact_id VARCHAR + ) + """); + + stmt.execute(""" + CREATE INDEX IF NOT EXISTS idx_artifact_imports_artifact + ON artifact_imports (artifact_id) + """); + + stmt.execute(""" + CREATE INDEX IF NOT EXISTS idx_artifact_imports_resolved + ON artifact_imports (resolved_artifact_id) + """); + log.info("Database schema created/verified"); } } diff --git a/src/main/java/com/javaducker/server/ingestion/Chunker.java b/src/main/java/com/javaducker/server/ingestion/Chunker.java index 72691fa..12fc147 100644 --- a/src/main/java/com/javaducker/server/ingestion/Chunker.java +++ b/src/main/java/com/javaducker/server/ingestion/Chunker.java @@ -8,7 +8,8 @@ @Component public class Chunker { - public record Chunk(int index, String text, long charStart, long charEnd) {} + public record Chunk(int index, String text, long charStart, long charEnd, + int lineStart, int lineEnd) {} public List chunk(String text, int chunkSize, int overlap) { List chunks = new ArrayList<>(); @@ -21,11 +22,31 @@ public List chunk(String text, int chunkSize, int overlap) { int index = 0; int pos = 0; + int currentLine = 1; + int scannedUpTo = 0; while (pos < text.length()) { int end = Math.min(pos + effectiveSize, text.length()); + + // Advance line count up to pos (for lineStart) + while (scannedUpTo < pos) { + if (text.charAt(scannedUpTo) == '\n') { + currentLine++; + } + scannedUpTo++; + } + int lineStart = currentLine; + + // Count additional newlines from pos to end (for lineEnd) + int lineEnd = lineStart; + for (int i = pos; i < end; i++) { + if (text.charAt(i) == '\n') { + lineEnd++; + } + } + String chunkText = text.substring(pos, end); - chunks.add(new Chunk(index, chunkText, pos, end)); + chunks.add(new Chunk(index, chunkText, pos, end, lineStart, lineEnd)); index++; pos += step; if (end == text.length()) break; diff --git a/src/main/java/com/javaducker/server/ingestion/FileSummarizer.java b/src/main/java/com/javaducker/server/ingestion/FileSummarizer.java new file mode 100644 index 0000000..8ca1564 --- /dev/null +++ b/src/main/java/com/javaducker/server/ingestion/FileSummarizer.java @@ -0,0 +1,179 @@ +package com.javaducker.server.ingestion; + +import org.springframework.stereotype.Component; + +import java.util.*; +import java.util.regex.Matcher; +import java.util.regex.Pattern; + +@Component +public class FileSummarizer { + + private record LangDef(String language, List classPatterns, + List methodPatterns, Pattern importPattern) {} + + private final Map langDefs = new HashMap<>(); + + public FileSummarizer() { + langDefs.put("java", new LangDef("Java", + List.of(Pattern.compile("class\\s+(\\w+)"), Pattern.compile("interface\\s+(\\w+)")), + List.of(Pattern.compile("(?:public|private|protected|static|\\s)+[\\w<>\\[\\]]+\\s+(\\w+)\\s*\\(")), + Pattern.compile("import\\s+[\\w.]+;"))); + + langDefs.put("js", new LangDef("JavaScript", + List.of(Pattern.compile("class\\s+(\\w+)")), + List.of(Pattern.compile("function\\s+(\\w+)"), + Pattern.compile("(?:const|let|var)\\s+(\\w+)\\s*=\\s*(?:async\\s*)?\\(")), + Pattern.compile("(?:import|require)\\s*\\("))); + + langDefs.put("ts", new LangDef("TypeScript", + List.of(Pattern.compile("class\\s+(\\w+)")), + List.of(Pattern.compile("function\\s+(\\w+)"), + Pattern.compile("(?:const|let|var)\\s+(\\w+)\\s*=\\s*(?:async\\s*)?\\(")), + Pattern.compile("(?:import|require)\\s*\\("))); + + langDefs.put("py", new LangDef("Python", + List.of(Pattern.compile("class\\s+(\\w+)")), + List.of(Pattern.compile("def\\s+(\\w+)")), + Pattern.compile("(?:import|from)\\s+\\w+"))); + + langDefs.put("go", new LangDef("Go", + List.of(Pattern.compile("type\\s+(\\w+)\\s+struct")), + List.of(Pattern.compile("func\\s+(?:\\([^)]+\\)\\s+)?(\\w+)")), + Pattern.compile("import\\s+"))); + + langDefs.put("rs", new LangDef("Rust", + List.of(Pattern.compile("(?:pub\\s+)?struct\\s+(\\w+)")), + List.of(Pattern.compile("(?:pub\\s+)?fn\\s+(\\w+)")), + Pattern.compile("use\\s+[\\w:]+"))); + + // Aliases + langDefs.put("jsx", langDefs.get("js")); + langDefs.put("tsx", langDefs.get("ts")); + langDefs.put("mjs", langDefs.get("js")); + } + + public Map summarize(String text, String fileName) { + Map result = new LinkedHashMap<>(); + + if (text == null) text = ""; + if (fileName == null) fileName = ""; + + String ext = extractExtension(fileName); + LangDef lang = langDefs.get(ext); + int lineCount = text.isEmpty() ? 0 : text.split("\n", -1).length; + + result.put("file_type", ext.isEmpty() ? "unknown" : ext); + result.put("language", lang != null ? lang.language() : humanLanguage(ext)); + result.put("line_count", lineCount); + + List classNames = new ArrayList<>(); + List methodNames = new ArrayList<>(); + List imports = new ArrayList<>(); + + if (lang != null) { + for (Pattern p : lang.classPatterns()) { + extractAll(p, text, classNames); + } + for (Pattern p : lang.methodPatterns()) { + extractAll(p, text, methodNames); + } + extractImports(lang.importPattern(), text, imports, 10); + } + + result.put("class_names", classNames); + result.put("method_names", methodNames); + result.put("imports", imports); + result.put("summary_text", buildSummary(fileName, lang, lineCount, classNames, methodNames, imports)); + + return result; + } + + private String extractExtension(String fileName) { + int dot = fileName.lastIndexOf('.'); + if (dot < 0 || dot == fileName.length() - 1) return ""; + return fileName.substring(dot + 1).toLowerCase(); + } + + private String humanLanguage(String ext) { + return switch (ext) { + case "java" -> "Java"; + case "js", "jsx", "mjs" -> "JavaScript"; + case "ts", "tsx" -> "TypeScript"; + case "py" -> "Python"; + case "go" -> "Go"; + case "rs" -> "Rust"; + case "rb" -> "Ruby"; + case "cpp", "cc", "cxx" -> "C++"; + case "c", "h" -> "C"; + case "cs" -> "C#"; + case "kt" -> "Kotlin"; + case "scala" -> "Scala"; + case "swift" -> "Swift"; + case "php" -> "PHP"; + case "sh", "bash" -> "Shell"; + case "yml", "yaml" -> "YAML"; + case "json" -> "JSON"; + case "xml" -> "XML"; + case "md" -> "Markdown"; + case "sql" -> "SQL"; + case "html", "htm" -> "HTML"; + case "css" -> "CSS"; + default -> "Unknown"; + }; + } + + private void extractAll(Pattern pattern, String text, List dest) { + Matcher m = pattern.matcher(text); + while (m.find()) { + String name = m.group(1); + if (name != null && !dest.contains(name)) { + dest.add(name); + } + } + } + + private void extractImports(Pattern pattern, String text, List dest, int limit) { + Matcher m = pattern.matcher(text); + while (m.find() && dest.size() < limit) { + dest.add(m.group().strip()); + } + } + + private String buildSummary(String fileName, LangDef lang, int lineCount, + List classNames, List methodNames, + List imports) { + StringBuilder sb = new StringBuilder(); + sb.append(fileName); + + if (lang != null) { + sb.append(" is a ").append(lang.language()).append(" file"); + } else { + sb.append(" is a file"); + } + + sb.append(" with ").append(lineCount).append(" lines."); + + if (!classNames.isEmpty()) { + sb.append(" It defines "); + sb.append(classNames.size() == 1 ? "class " : "classes "); + sb.append(String.join(", ", classNames)).append("."); + } + + if (!methodNames.isEmpty()) { + sb.append(" It contains ").append(methodNames.size()); + sb.append(methodNames.size() == 1 ? " method: " : " methods: "); + List shown = methodNames.size() > 8 + ? methodNames.subList(0, 8) : methodNames; + sb.append(String.join(", ", shown)); + if (methodNames.size() > 8) sb.append(" and others"); + sb.append("."); + } + + if (!imports.isEmpty()) { + sb.append(" It has ").append(imports.size()).append(" import(s)."); + } + + return sb.toString(); + } +} diff --git a/src/main/java/com/javaducker/server/ingestion/FileWatcher.java b/src/main/java/com/javaducker/server/ingestion/FileWatcher.java new file mode 100644 index 0000000..49a2d5b --- /dev/null +++ b/src/main/java/com/javaducker/server/ingestion/FileWatcher.java @@ -0,0 +1,142 @@ +package com.javaducker.server.ingestion; + +import com.javaducker.server.service.UploadService; +import org.slf4j.Logger; +import org.slf4j.LoggerFactory; +import org.springframework.stereotype.Component; + +import java.io.IOException; +import java.nio.file.*; +import java.nio.file.attribute.BasicFileAttributes; +import java.util.Map; +import java.util.Set; +import java.util.concurrent.ConcurrentHashMap; + +@Component +public class FileWatcher { + + private static final Logger log = LoggerFactory.getLogger(FileWatcher.class); + private static final long DEBOUNCE_MS = 500; + private static final Set EXCLUDED_DIRS = Set.of( + "node_modules", ".git", "target", "build", "__pycache__", ".idea", "temp"); + + private final UploadService uploadService; + private volatile WatchService watchService; + private volatile Path watchedDirectory; + private volatile boolean watching; + private Thread watchThread; + private Set allowedExtensions; + private final Map lastModified = new ConcurrentHashMap<>(); + + public FileWatcher(UploadService uploadService) { + this.uploadService = uploadService; + } + + public void startWatching(Path directory, Set extensions) throws IOException { + stopWatching(); + this.allowedExtensions = extensions; + this.watchService = FileSystems.getDefault().newWatchService(); + registerRecursive(directory); + this.watchedDirectory = directory; + this.watching = true; + + watchThread = new Thread(this::watchLoop, "file-watcher"); + watchThread.setDaemon(true); + watchThread.start(); + log.info("Started watching directory: {} for extensions: {}", directory, extensions); + } + + public void stopWatching() { + watching = false; + if (watchService != null) { + try { watchService.close(); } catch (IOException e) { log.warn("Error closing WatchService", e); } + watchService = null; + } + if (watchThread != null) { + watchThread.interrupt(); + watchThread = null; + } + watchedDirectory = null; + lastModified.clear(); + log.info("Stopped watching"); + } + + public boolean isWatching() { return watching; } + + public Path getWatchedDirectory() { return watchedDirectory; } + + private void registerRecursive(Path root) throws IOException { + Files.walkFileTree(root, new SimpleFileVisitor<>() { + @Override + public FileVisitResult preVisitDirectory(Path dir, BasicFileAttributes attrs) { + if (EXCLUDED_DIRS.contains(dir.getFileName().toString())) { + return FileVisitResult.SKIP_SUBTREE; + } + try { + dir.register(watchService, + StandardWatchEventKinds.ENTRY_CREATE, + StandardWatchEventKinds.ENTRY_MODIFY); + } catch (IOException e) { + log.warn("Failed to register: {}", dir, e); + } + return FileVisitResult.CONTINUE; + } + }); + } + + private void watchLoop() { + while (watching) { + try { + WatchKey key = watchService.take(); + Path dir = (Path) key.watchable(); + for (WatchEvent event : key.pollEvents()) { + if (event.kind() == StandardWatchEventKinds.OVERFLOW) continue; + Path changed = dir.resolve((Path) event.context()); + handleEvent(changed); + } + key.reset(); + } catch (ClosedWatchServiceException e) { + break; + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + break; + } + } + watching = false; + } + + private void handleEvent(Path path) { + try { + if (Files.isDirectory(path)) { + if (!EXCLUDED_DIRS.contains(path.getFileName().toString())) { + registerRecursive(path); + } + return; + } + String fileName = path.getFileName().toString(); + if (!hasAllowedExtension(fileName)) return; + for (Path part : path) { + if (EXCLUDED_DIRS.contains(part.toString())) return; + } + + long now = System.currentTimeMillis(); + Long prev = lastModified.put(path, now); + if (prev != null && (now - prev) < DEBOUNCE_MS) return; + + byte[] content = Files.readAllBytes(path); + String mediaType = Files.probeContentType(path); + if (mediaType == null) mediaType = "application/octet-stream"; + + String artifactId = uploadService.upload( + fileName, path.toAbsolutePath().toString(), mediaType, content.length, content); + log.info("Auto-indexed file: {} -> {}", path, artifactId); + } catch (Exception e) { + log.error("Error processing file event for {}: {}", path, e.getMessage()); + } + } + + private boolean hasAllowedExtension(String fileName) { + if (allowedExtensions == null || allowedExtensions.isEmpty()) return true; + return allowedExtensions.stream().anyMatch(fileName::endsWith); + } +} diff --git a/src/main/java/com/javaducker/server/ingestion/HnswIndex.java b/src/main/java/com/javaducker/server/ingestion/HnswIndex.java new file mode 100644 index 0000000..4467a44 --- /dev/null +++ b/src/main/java/com/javaducker/server/ingestion/HnswIndex.java @@ -0,0 +1,292 @@ +package com.javaducker.server.ingestion; + +import java.util.*; +import java.util.concurrent.ConcurrentHashMap; +import java.util.concurrent.ThreadLocalRandom; + +/** + * Pure Java HNSW (Hierarchical Navigable Small World) index for approximate nearest neighbor search. + * Uses cosine distance (1 - cosine_similarity) as the distance metric. + */ +public class HnswIndex { + + private final int dimension; + private final int m; + private final int efConstruction; + private final int efSearch; + private final double mL; // level multiplier: 1 / ln(M) + + private final ConcurrentHashMap nodes = new ConcurrentHashMap<>(); + private volatile String entryPointId; + private volatile int maxLevel; + + public HnswIndex(int dimension, int m, int efConstruction, int efSearch) { + this.dimension = dimension; + this.m = m; + this.efConstruction = efConstruction; + this.efSearch = efSearch; + this.mL = 1.0 / Math.log(m); + this.maxLevel = -1; + } + + public record Result(String id, double distance) {} + + private static class Node { + final String id; + final double[] vector; + final int level; + final List> connections; // connections[layer] = neighbor IDs + + Node(String id, double[] vector, int level) { + this.id = id; + this.vector = vector; + this.level = level; + this.connections = new ArrayList<>(level + 1); + for (int i = 0; i <= level; i++) { + this.connections.add(Collections.synchronizedList(new ArrayList<>())); + } + } + } + + public synchronized void insert(String id, double[] vector) { + if (vector.length != dimension) { + throw new IllegalArgumentException("Expected dimension " + dimension + ", got " + vector.length); + } + if (nodes.containsKey(id)) return; + + int nodeLevel = randomLevel(); + Node newNode = new Node(id, vector, nodeLevel); + nodes.put(id, newNode); + + if (entryPointId == null) { + entryPointId = id; + maxLevel = nodeLevel; + return; + } + + String currentId = entryPointId; + int currentMaxLevel = maxLevel; + + // Greedy descent from top to nodeLevel + 1 + for (int layer = currentMaxLevel; layer > nodeLevel; layer--) { + currentId = greedyClosest(currentId, vector, layer); + } + + // For each layer from min(nodeLevel, currentMaxLevel) down to 0, search and connect + for (int layer = Math.min(nodeLevel, currentMaxLevel); layer >= 0; layer--) { + List neighbors = searchLayer(currentId, vector, efConstruction, layer); + // Select M closest + List selected = selectNeighbors(vector, neighbors, m); + + // Set connections for new node at this layer + newNode.connections.get(layer).addAll(selected); + + // Add bidirectional connections and prune if needed + for (String neighborId : selected) { + Node neighbor = nodes.get(neighborId); + if (neighbor == null || layer > neighbor.level) continue; + List nConns = neighbor.connections.get(layer); + nConns.add(id); + if (nConns.size() > m) { + pruneConnections(neighbor, layer); + } + } + + if (!neighbors.isEmpty()) { + // Update currentId to closest found for next layer descent + currentId = neighbors.get(0); + } + } + + // Update entry point if new node has higher level + if (nodeLevel > maxLevel) { + entryPointId = id; + maxLevel = nodeLevel; + } + } + + public List search(double[] query, int k) { + if (query.length != dimension) { + throw new IllegalArgumentException("Expected dimension " + dimension + ", got " + query.length); + } + if (entryPointId == null) return Collections.emptyList(); + + String currentId = entryPointId; + + // Greedy descent from top layer to layer 1 + for (int layer = maxLevel; layer > 0; layer--) { + currentId = greedyClosest(currentId, query, layer); + } + + // Beam search at layer 0 + List candidates = searchLayer(currentId, query, Math.max(efSearch, k), 0); + + // Return top-k + List results = new ArrayList<>(); + for (String candidateId : candidates) { + Node node = nodes.get(candidateId); + if (node != null) { + results.add(new Result(candidateId, cosineDistance(query, node.vector))); + } + } + results.sort(Comparator.comparingDouble(Result::distance)); + return results.size() > k ? results.subList(0, k) : results; + } + + public int size() { + return nodes.size(); + } + + public boolean isEmpty() { + return nodes.isEmpty(); + } + + /** + * Greedily traverse the layer to find the single closest node to the query vector. + */ + private String greedyClosest(String startId, double[] query, int layer) { + String bestId = startId; + double bestDist = cosineDistance(query, nodes.get(startId).vector); + + boolean improved = true; + while (improved) { + improved = false; + Node bestNode = nodes.get(bestId); + if (bestNode == null || layer > bestNode.level) break; + List conns = bestNode.connections.get(layer); + for (String neighborId : conns) { + Node neighbor = nodes.get(neighborId); + if (neighbor == null) continue; + double dist = cosineDistance(query, neighbor.vector); + if (dist < bestDist) { + bestDist = dist; + bestId = neighborId; + improved = true; + } + } + } + return bestId; + } + + /** + * Beam search at a given layer. Returns candidate IDs sorted by distance (closest first). + */ + private List searchLayer(String startId, double[] query, int ef, int layer) { + Set visited = new HashSet<>(); + double startDist = cosineDistance(query, nodes.get(startId).vector); + visited.add(startId); + + TreeMap> candidateMap = new TreeMap<>(); + TreeMap> resultMap = new TreeMap<>(); + addToMap(candidateMap, startDist, startId); + addToMap(resultMap, startDist, startId); + + while (!candidateMap.isEmpty()) { + // Get closest candidate + Map.Entry> closest = candidateMap.firstEntry(); + double cDist = closest.getKey(); + String cId = closest.getValue().remove(0); + if (closest.getValue().isEmpty()) candidateMap.pollFirstEntry(); + + // Get farthest result + double farthestResultDist = resultMap.lastKey(); + if (cDist > farthestResultDist) break; + + Node cNode = nodes.get(cId); + if (cNode == null || layer > cNode.level) continue; + + for (String neighborId : cNode.connections.get(layer)) { + if (visited.contains(neighborId)) continue; + visited.add(neighborId); + + Node neighbor = nodes.get(neighborId); + if (neighbor == null) continue; + double nDist = cosineDistance(query, neighbor.vector); + farthestResultDist = resultMap.lastKey(); + if (nDist < farthestResultDist || countMap(resultMap) < ef) { + addToMap(candidateMap, nDist, neighborId); + addToMap(resultMap, nDist, neighborId); + + // Trim results to ef + while (countMap(resultMap) > ef) { + Map.Entry> last = resultMap.lastEntry(); + last.getValue().remove(last.getValue().size() - 1); + if (last.getValue().isEmpty()) resultMap.pollLastEntry(); + } + } + } + } + + // Collect results sorted by distance + List sortedResults = new ArrayList<>(); + for (Map.Entry> entry : resultMap.entrySet()) { + sortedResults.addAll(entry.getValue()); + } + return sortedResults; + } + + private void addToMap(TreeMap> map, double dist, String id) { + map.computeIfAbsent(dist, k -> new ArrayList<>()).add(id); + } + + private int countMap(TreeMap> map) { + int count = 0; + for (List v : map.values()) count += v.size(); + return count; + } + + /** + * Select up to maxConn nearest neighbors from candidates. + */ + private List selectNeighbors(double[] query, List candidates, int maxConn) { + List> scored = new ArrayList<>(); + for (String id : candidates) { + Node node = nodes.get(id); + if (node != null) { + scored.add(Map.entry(id, cosineDistance(query, node.vector))); + } + } + scored.sort(Comparator.comparingDouble(Map.Entry::getValue)); + List selected = new ArrayList<>(); + for (int i = 0; i < Math.min(maxConn, scored.size()); i++) { + selected.add(scored.get(i).getKey()); + } + return selected; + } + + /** + * Prune connections of a node at a given layer to keep only M closest. + */ + private void pruneConnections(Node node, int layer) { + List conns = node.connections.get(layer); + List> scored = new ArrayList<>(); + for (String neighborId : conns) { + Node neighbor = nodes.get(neighborId); + if (neighbor != null) { + scored.add(Map.entry(neighborId, cosineDistance(node.vector, neighbor.vector))); + } + } + scored.sort(Comparator.comparingDouble(Map.Entry::getValue)); + List pruned = new ArrayList<>(); + for (int i = 0; i < Math.min(m, scored.size()); i++) { + pruned.add(scored.get(i).getKey()); + } + conns.clear(); + conns.addAll(pruned); + } + + private int randomLevel() { + return (int) Math.floor(-Math.log(ThreadLocalRandom.current().nextDouble()) * mL); + } + + static double cosineDistance(double[] a, double[] b) { + double dot = 0, normA = 0, normB = 0; + for (int i = 0; i < a.length; i++) { + dot += a[i] * b[i]; + normA += a[i] * a[i]; + normB += b[i] * b[i]; + } + if (normA == 0 || normB == 0) return 1.0; + return 1.0 - dot / (Math.sqrt(normA) * Math.sqrt(normB)); + } +} diff --git a/src/main/java/com/javaducker/server/ingestion/ImportParser.java b/src/main/java/com/javaducker/server/ingestion/ImportParser.java new file mode 100644 index 0000000..22e9665 --- /dev/null +++ b/src/main/java/com/javaducker/server/ingestion/ImportParser.java @@ -0,0 +1,84 @@ +package com.javaducker.server.ingestion; + +import org.springframework.stereotype.Component; + +import java.util.ArrayList; +import java.util.LinkedHashSet; +import java.util.List; +import java.util.Set; +import java.util.regex.Matcher; +import java.util.regex.Pattern; + +@Component +public class ImportParser { + + private static final Pattern JAVA_IMPORT = Pattern.compile("import\\s+([\\w.]+);"); + private static final Pattern JS_IMPORT_FROM = Pattern.compile("import\\s+.*?from\\s+['\"]([^'\"]+)['\"]"); + private static final Pattern JS_REQUIRE = Pattern.compile("require\\s*\\(\\s*['\"]([^'\"]+)['\"]\\s*\\)"); + private static final Pattern PYTHON_IMPORT = Pattern.compile("import\\s+([\\w.]+)"); + private static final Pattern PYTHON_FROM = Pattern.compile("from\\s+([\\w.]+)\\s+import"); + private static final Pattern GO_IMPORT = Pattern.compile("\"([^\"]+)\""); + private static final Pattern RUST_USE = Pattern.compile("use\\s+([\\w:]+)"); + + public List parseImports(String text, String fileName) { + if (text == null || fileName == null) { + return List.of(); + } + + String ext = getExtension(fileName).toLowerCase(); + Set imports = new LinkedHashSet<>(); + + switch (ext) { + case "java" -> extractAll(JAVA_IMPORT, text, imports); + case "js", "jsx", "ts", "tsx", "mjs", "cjs" -> { + extractAll(JS_IMPORT_FROM, text, imports); + extractAll(JS_REQUIRE, text, imports); + } + case "py" -> { + extractAll(PYTHON_IMPORT, text, imports); + extractAll(PYTHON_FROM, text, imports); + } + case "go" -> extractGoImports(text, imports); + case "rs" -> extractAll(RUST_USE, text, imports); + default -> { /* unsupported language */ } + } + + return new ArrayList<>(imports); + } + + private void extractAll(Pattern pattern, String text, Set results) { + Matcher matcher = pattern.matcher(text); + while (matcher.find()) { + String match = matcher.group(1).trim(); + if (!match.isEmpty()) { + results.add(match); + } + } + } + + private void extractGoImports(String text, Set results) { + // Match single imports: import "path" + // Match block imports: import ( "path1" \n "path2" ) + Pattern blockPattern = Pattern.compile("import\\s*\\(([^)]+)\\)", Pattern.DOTALL); + Matcher blockMatcher = blockPattern.matcher(text); + while (blockMatcher.find()) { + String block = blockMatcher.group(1); + Matcher quoteMatcher = GO_IMPORT.matcher(block); + while (quoteMatcher.find()) { + results.add(quoteMatcher.group(1)); + } + } + + // Single-line imports + Pattern singleImport = Pattern.compile("import\\s+\"([^\"]+)\""); + Matcher singleMatcher = singleImport.matcher(text); + while (singleMatcher.find()) { + results.add(singleMatcher.group(1)); + } + } + + private String getExtension(String fileName) { + int dot = fileName.lastIndexOf('.'); + return dot >= 0 ? fileName.substring(dot + 1) : ""; + } +} diff --git a/src/main/java/com/javaducker/server/ingestion/IngestionWorker.java b/src/main/java/com/javaducker/server/ingestion/IngestionWorker.java index 937d8fe..9f2a98b 100644 --- a/src/main/java/com/javaducker/server/ingestion/IngestionWorker.java +++ b/src/main/java/com/javaducker/server/ingestion/IngestionWorker.java @@ -4,6 +4,7 @@ import com.javaducker.server.db.DuckDBDataSource; import com.javaducker.server.model.ArtifactStatus; import com.javaducker.server.service.ArtifactService; +import com.javaducker.server.service.SearchService; import jakarta.annotation.PreDestroy; import org.slf4j.Logger; import org.slf4j.LoggerFactory; @@ -11,13 +12,11 @@ import org.springframework.stereotype.Component; import java.nio.file.Path; +import java.sql.Connection; import java.sql.PreparedStatement; import java.sql.ResultSet; import java.sql.SQLException; -import java.util.ArrayList; -import java.util.EnumMap; -import java.util.List; -import java.util.Map; +import java.util.*; import java.util.concurrent.Executors; import java.util.concurrent.ThreadPoolExecutor; import java.util.concurrent.TimeUnit; @@ -33,6 +32,7 @@ public class IngestionWorker { private final TextNormalizer textNormalizer; private final Chunker chunker; private final EmbeddingService embeddingService; + private final SearchService searchService; private final AppConfig config; private final ThreadPoolExecutor threadPool; private volatile boolean ready = false; @@ -42,15 +42,24 @@ public class IngestionWorker { private final AtomicLong lastCompletedTasks = new AtomicLong(0); private volatile long lastProgressLogMs = System.currentTimeMillis(); + private final FileSummarizer fileSummarizer; + private final ImportParser importParser; + public IngestionWorker(DuckDBDataSource dataSource, ArtifactService artifactService, TextExtractor textExtractor, TextNormalizer textNormalizer, - Chunker chunker, EmbeddingService embeddingService, AppConfig config) { + Chunker chunker, EmbeddingService embeddingService, + FileSummarizer fileSummarizer, ImportParser importParser, + SearchService searchService, + AppConfig config) { this.dataSource = dataSource; this.artifactService = artifactService; this.textExtractor = textExtractor; this.textNormalizer = textNormalizer; this.chunker = chunker; this.embeddingService = embeddingService; + this.fileSummarizer = fileSummarizer; + this.importParser = importParser; + this.searchService = searchService; this.config = config; this.threadPool = (ThreadPoolExecutor) Executors.newFixedThreadPool(config.getIngestionWorkerThreads()); } @@ -210,7 +219,7 @@ public void processArtifact(String artifactId) { dataSource.withConnection(conn -> { try (PreparedStatement ps = conn.prepareStatement( - "INSERT INTO artifact_chunks (chunk_id, artifact_id, chunk_index, chunk_text, char_start, char_end) VALUES (?, ?, ?, ?, ?, ?)")) { + "INSERT INTO artifact_chunks (chunk_id, artifact_id, chunk_index, chunk_text, char_start, char_end, line_start, line_end) VALUES (?, ?, ?, ?, ?, ?, ?, ?)")) { for (Chunker.Chunk chunk : chunks) { String chunkId = artifactId + "-" + chunk.index(); ps.setString(1, chunkId); @@ -219,6 +228,8 @@ public void processArtifact(String artifactId) { ps.setString(4, chunk.text()); ps.setLong(5, chunk.charStart()); ps.setLong(6, chunk.charEnd()); + ps.setInt(7, chunk.lineStart()); + ps.setInt(8, chunk.lineEnd()); ps.executeUpdate(); } } @@ -233,6 +244,14 @@ record ChunkEmbedding(String chunkId, double[] embedding) {} embeddings.add(new ChunkEmbedding(artifactId + "-" + chunk.index(), embeddingService.embed(chunk.text()))); } + // Add to HNSW index if available + HnswIndex hnsw = searchService.getHnswIndex(); + if (hnsw != null) { + for (ChunkEmbedding ce : embeddings) { + hnsw.insert(ce.chunkId(), ce.embedding()); + } + } + dataSource.withConnection(conn -> { for (ChunkEmbedding ce : embeddings) { StringBuilder sb = new StringBuilder("["); @@ -252,7 +271,48 @@ record ChunkEmbedding(String chunkId, double[] embedding) {} return null; }); - // Step 4: Mark indexed + // Step 4: Generate file summary + try { + Map summary = fileSummarizer.summarize(normalizedText, fileName); + dataSource.withConnection(conn -> { + try (PreparedStatement ps = conn.prepareStatement( + "INSERT INTO artifact_summaries (artifact_id, summary_text, class_names, method_names, import_count, line_count) VALUES (?, ?, ?, ?, ?, ?)")) { + ps.setString(1, artifactId); + ps.setString(2, (String) summary.get("summary_text")); + ps.setString(3, String.join(", ", (List) summary.get("class_names"))); + ps.setString(4, String.join(", ", (List) summary.get("method_names"))); + ps.setInt(5, ((List) summary.get("imports")).size()); + ps.setInt(6, (int) summary.get("line_count")); + ps.executeUpdate(); + } + return null; + }); + } catch (Exception e) { + log.warn("Summary generation failed for {}, continuing", artifactId, e); + } + + // Step 5: Parse imports + try { + List imports = importParser.parseImports(normalizedText, fileName); + if (!imports.isEmpty()) { + dataSource.withConnection(conn -> { + try (PreparedStatement ps = conn.prepareStatement( + "INSERT INTO artifact_imports (artifact_id, import_statement, resolved_artifact_id) VALUES (?, ?, ?)")) { + for (String imp : imports) { + ps.setString(1, artifactId); + ps.setString(2, imp); + ps.setString(3, resolveImport(conn, imp, fileName)); + ps.executeUpdate(); + } + } + return null; + }); + } + } catch (Exception e) { + log.warn("Import parsing failed for {}, continuing", artifactId, e); + } + + // Step 6: Mark indexed artifactService.updateStatus(artifactId, ArtifactStatus.INDEXED, null); log.info("Artifact indexed: {} ({}, {} chunks)", artifactId, fileName, chunks.size()); @@ -265,4 +325,68 @@ record ChunkEmbedding(String chunkId, double[] embedding) {} } } } + + public void buildHnswIndex() throws SQLException { + HnswIndex index = new HnswIndex(config.getEmbeddingDim(), 16, 200, 50); + dataSource.withConnection(conn -> { + try (PreparedStatement ps = conn.prepareStatement( + "SELECT chunk_id, embedding FROM chunk_embeddings"); + ResultSet rs = ps.executeQuery()) { + while (rs.next()) { + String chunkId = rs.getString("chunk_id"); + double[] embedding = extractEmbeddingArray(rs); + if (embedding != null) { + index.insert(chunkId, embedding); + } + } + } + return null; + }); + searchService.setHnswIndex(index); + log.info("HNSW index built with {} vectors", index.size()); + } + + private double[] extractEmbeddingArray(ResultSet rs) throws SQLException { + Object embObj = rs.getObject("embedding"); + if (embObj == null) return null; + if (embObj instanceof double[] arr) return arr; + if (embObj instanceof Object[] objArr) { + double[] result = new double[objArr.length]; + for (int i = 0; i < objArr.length; i++) { + result[i] = ((Number) objArr[i]).doubleValue(); + } + return result; + } + if (embObj instanceof java.sql.Array sqlArray) { + Object[] arr = (Object[]) sqlArray.getArray(); + double[] result = new double[arr.length]; + for (int i = 0; i < arr.length; i++) { + result[i] = ((Number) arr[i]).doubleValue(); + } + return result; + } + return null; + } + + private String resolveImport(Connection conn, String importStr, String fileName) { + try { + // Only resolve Java imports for now + if (!fileName.endsWith(".java")) { + return null; + } + String pathSuffix = importStr.replace('.', '/') + ".java"; + try (PreparedStatement ps = conn.prepareStatement( + "SELECT artifact_id FROM artifacts WHERE original_client_path LIKE '%' || ? AND status = 'INDEXED' LIMIT 1")) { + ps.setString(1, pathSuffix); + try (ResultSet rs = ps.executeQuery()) { + if (rs.next()) { + return rs.getString("artifact_id"); + } + } + } + } catch (Exception e) { + log.debug("Could not resolve import '{}': {}", importStr, e.getMessage()); + } + return null; + } } diff --git a/src/main/java/com/javaducker/server/rest/JavaDuckerRestController.java b/src/main/java/com/javaducker/server/rest/JavaDuckerRestController.java index 45207fc..a54557c 100644 --- a/src/main/java/com/javaducker/server/rest/JavaDuckerRestController.java +++ b/src/main/java/com/javaducker/server/rest/JavaDuckerRestController.java @@ -1,15 +1,17 @@ package com.javaducker.server.rest; -import com.javaducker.server.service.ArtifactService; -import com.javaducker.server.service.SearchService; -import com.javaducker.server.service.StatsService; -import com.javaducker.server.service.UploadService; +import com.javaducker.server.ingestion.FileWatcher; +import com.javaducker.server.service.*; import org.springframework.http.ResponseEntity; import org.springframework.web.bind.annotation.*; import org.springframework.web.multipart.MultipartFile; +import java.nio.file.Path; +import java.util.Arrays; import java.util.List; import java.util.Map; +import java.util.Set; +import java.util.stream.Collectors; @RestController @RequestMapping("/api") @@ -19,13 +21,23 @@ public class JavaDuckerRestController { private final ArtifactService artifactService; private final SearchService searchService; private final StatsService statsService; + private final ProjectMapService projectMapService; + private final StalenessService stalenessService; + private final DependencyService dependencyService; + private final FileWatcher fileWatcher; public JavaDuckerRestController(UploadService uploadService, ArtifactService artifactService, - SearchService searchService, StatsService statsService) { + SearchService searchService, StatsService statsService, + ProjectMapService projectMapService, StalenessService stalenessService, + DependencyService dependencyService, FileWatcher fileWatcher) { this.uploadService = uploadService; this.artifactService = artifactService; this.searchService = searchService; this.statsService = statsService; + this.projectMapService = projectMapService; + this.stalenessService = stalenessService; + this.dependencyService = dependencyService; + this.fileWatcher = fileWatcher; } @GetMapping("/health") @@ -82,4 +94,69 @@ public ResponseEntity> search(@RequestBody Map getSummary(@PathVariable String artifactId) throws Exception { + Map summary = artifactService.getSummary(artifactId); + if (summary == null) return ResponseEntity.notFound().build(); + return ResponseEntity.ok(summary); + } + + @GetMapping("/map") + public ResponseEntity> getProjectMap() throws Exception { + return ResponseEntity.ok(projectMapService.getProjectMap()); + } + + @SuppressWarnings("unchecked") + @PostMapping("/stale") + public ResponseEntity> checkStale(@RequestBody Map body) throws Exception { + List filePaths = (List) body.get("file_paths"); + if (filePaths == null || filePaths.isEmpty()) { + return ResponseEntity.badRequest().body(Map.of("error", "file_paths is required")); + } + return ResponseEntity.ok(stalenessService.checkStaleness(filePaths)); + } + + @GetMapping("/dependencies/{artifactId}") + public ResponseEntity getDependencies(@PathVariable String artifactId) throws Exception { + return ResponseEntity.ok(Map.of("artifact_id", artifactId, + "dependencies", dependencyService.getDependencies(artifactId))); + } + + @GetMapping("/dependents/{artifactId}") + public ResponseEntity getDependents(@PathVariable String artifactId) throws Exception { + return ResponseEntity.ok(Map.of("artifact_id", artifactId, + "dependents", dependencyService.getDependents(artifactId))); + } + + @PostMapping("/watch/start") + public ResponseEntity> startWatch(@RequestBody Map body) throws Exception { + String directory = (String) body.get("directory"); + String extensions = (String) body.getOrDefault("extensions", ""); + Set extSet = extensions.isBlank() + ? Set.of() + : Arrays.stream(extensions.split(",")) + .map(String::trim) + .filter(s -> !s.isEmpty()) + .collect(Collectors.toSet()); + fileWatcher.startWatching(Path.of(directory), extSet); + return ResponseEntity.ok(Map.of( + "status", "watching", + "directory", directory, + "extensions", extSet)); + } + + @PostMapping("/watch/stop") + public ResponseEntity> stopWatch() { + fileWatcher.stopWatching(); + return ResponseEntity.ok(Map.of("status", "stopped")); + } + + @GetMapping("/watch/status") + public ResponseEntity> watchStatus() { + return ResponseEntity.ok(Map.of( + "watching", fileWatcher.isWatching(), + "directory", fileWatcher.getWatchedDirectory() != null + ? fileWatcher.getWatchedDirectory().toString() : "")); + } } diff --git a/src/main/java/com/javaducker/server/service/ArtifactService.java b/src/main/java/com/javaducker/server/service/ArtifactService.java index 6a7bfcc..1cb4eb4 100644 --- a/src/main/java/com/javaducker/server/service/ArtifactService.java +++ b/src/main/java/com/javaducker/server/service/ArtifactService.java @@ -66,6 +66,49 @@ public Map getText(String artifactId) throws SQLException { }); } + public Map getSummary(String artifactId) throws SQLException { + return dataSource.withConnection(conn -> { + try (PreparedStatement ps = conn.prepareStatement( + "SELECT artifact_id, summary_text, class_names, method_names, import_count, line_count FROM artifact_summaries WHERE artifact_id = ?")) { + ps.setString(1, artifactId); + try (ResultSet rs = ps.executeQuery()) { + if (rs.next()) { + Map result = new HashMap<>(); + result.put("artifact_id", rs.getString("artifact_id")); + result.put("summary_text", rs.getString("summary_text")); + result.put("class_names", rs.getString("class_names")); + result.put("method_names", rs.getString("method_names")); + result.put("import_count", rs.getInt("import_count")); + result.put("line_count", rs.getInt("line_count")); + return result; + } + } + } + return null; + }); + } + + public void deleteArtifactData(String artifactId) throws SQLException { + log.info("Deleting indexed data for artifact {}", artifactId); + dataSource.withConnection(conn -> { + // chunk_embeddings uses chunk_id, not artifact_id — delete via subquery + try (PreparedStatement ps = conn.prepareStatement( + "DELETE FROM chunk_embeddings WHERE chunk_id IN (SELECT chunk_id FROM artifact_chunks WHERE artifact_id = ?)")) { + ps.setString(1, artifactId); + ps.executeUpdate(); + } + String[] tables = {"artifact_chunks", "artifact_text", "artifact_summaries", "artifact_imports", "ingestion_events"}; + for (String table : tables) { + try (PreparedStatement ps = conn.prepareStatement( + "DELETE FROM " + table + " WHERE artifact_id = ?")) { + ps.setString(1, artifactId); + ps.executeUpdate(); + } + } + return null; + }); + } + public void updateStatus(String artifactId, ArtifactStatus status, String errorMessage) throws SQLException { dataSource.withConnection(conn -> { String sql; diff --git a/src/main/java/com/javaducker/server/service/DependencyService.java b/src/main/java/com/javaducker/server/service/DependencyService.java new file mode 100644 index 0000000..8f10a43 --- /dev/null +++ b/src/main/java/com/javaducker/server/service/DependencyService.java @@ -0,0 +1,68 @@ +package com.javaducker.server.service; + +import com.javaducker.server.db.DuckDBDataSource; +import org.springframework.stereotype.Service; + +import java.sql.PreparedStatement; +import java.sql.ResultSet; +import java.sql.SQLException; +import java.util.ArrayList; +import java.util.LinkedHashMap; +import java.util.List; +import java.util.Map; + +@Service +public class DependencyService { + + private final DuckDBDataSource dataSource; + + public DependencyService(DuckDBDataSource dataSource) { + this.dataSource = dataSource; + } + + public List> getDependencies(String artifactId) throws SQLException { + return dataSource.withConnection(conn -> { + List> results = new ArrayList<>(); + try (PreparedStatement ps = conn.prepareStatement( + "SELECT artifact_id, import_statement, resolved_artifact_id FROM artifact_imports WHERE artifact_id = ?")) { + ps.setString(1, artifactId); + try (ResultSet rs = ps.executeQuery()) { + while (rs.next()) { + Map row = new LinkedHashMap<>(); + row.put("artifact_id", rs.getString("artifact_id")); + row.put("import_statement", rs.getString("import_statement")); + row.put("resolved_artifact_id", rs.getString("resolved_artifact_id")); + results.add(row); + } + } + } + return results; + }); + } + + public List> getDependents(String artifactId) throws SQLException { + return dataSource.withConnection(conn -> { + List> results = new ArrayList<>(); + try (PreparedStatement ps = conn.prepareStatement( + """ + SELECT ai.artifact_id, ai.import_statement, ai.resolved_artifact_id, a.file_name + FROM artifact_imports ai + JOIN artifacts a ON a.artifact_id = ai.artifact_id + WHERE ai.resolved_artifact_id = ? + """)) { + ps.setString(1, artifactId); + try (ResultSet rs = ps.executeQuery()) { + while (rs.next()) { + Map row = new LinkedHashMap<>(); + row.put("artifact_id", rs.getString("artifact_id")); + row.put("import_statement", rs.getString("import_statement")); + row.put("resolved_artifact_id", rs.getString("resolved_artifact_id")); + row.put("file_name", rs.getString("file_name")); + results.add(row); + } + } + } + return results; + }); + } +} diff --git a/src/main/java/com/javaducker/server/service/ProjectMapService.java b/src/main/java/com/javaducker/server/service/ProjectMapService.java new file mode 100644 index 0000000..dd9fbc1 --- /dev/null +++ b/src/main/java/com/javaducker/server/service/ProjectMapService.java @@ -0,0 +1,104 @@ +package com.javaducker.server.service; + +import com.javaducker.server.db.DuckDBDataSource; +import org.springframework.stereotype.Service; + +import java.sql.PreparedStatement; +import java.sql.ResultSet; +import java.sql.SQLException; +import java.util.*; + +@Service +public class ProjectMapService { + + private final DuckDBDataSource dataSource; + + public ProjectMapService(DuckDBDataSource dataSource) { + this.dataSource = dataSource; + } + + public Map getProjectMap() throws SQLException { + return dataSource.withConnection(conn -> { + Map result = new LinkedHashMap<>(); + Map> dirFiles = new LinkedHashMap<>(); + long totalFiles = 0; + long totalBytes = 0; + List> largest = new ArrayList<>(); + List> recent = new ArrayList<>(); + + String sql = "SELECT original_client_path, file_name, size_bytes, indexed_at " + + "FROM artifacts " + + "WHERE status = 'INDEXED' AND original_client_path IS NOT NULL " + + "ORDER BY size_bytes DESC"; + + try (PreparedStatement ps = conn.prepareStatement(sql); + ResultSet rs = ps.executeQuery()) { + while (rs.next()) { + String path = rs.getString("original_client_path"); + String fileName = rs.getString("file_name"); + long size = rs.getLong("size_bytes"); + String indexedAt = rs.getString("indexed_at"); + + totalFiles++; + totalBytes += size; + + String dir = extractParentDir(path); + dirFiles.computeIfAbsent(dir, k -> new ArrayList<>()).add(fileName); + + if (largest.size() < 5) { + largest.add(fileEntry(fileName, size, path)); + } + + recent.add(recentEntry(fileName, indexedAt, path)); + } + } + + // Sort recent by indexed_at descending, keep top 5 + recent.sort((a, b) -> String.valueOf(b.get("indexed_at")) + .compareTo(String.valueOf(a.get("indexed_at")))); + if (recent.size() > 5) { + recent = recent.subList(0, 5); + } + + // Build directory list sorted by file_count desc + List> directories = new ArrayList<>(); + for (var entry : dirFiles.entrySet()) { + Map dirEntry = new LinkedHashMap<>(); + dirEntry.put("path", entry.getKey()); + dirEntry.put("file_count", entry.getValue().size()); + dirEntry.put("files", entry.getValue()); + directories.add(dirEntry); + } + directories.sort((a, b) -> Integer.compare( + (int) b.get("file_count"), (int) a.get("file_count"))); + + result.put("total_files", totalFiles); + result.put("total_bytes", totalBytes); + result.put("directories", directories); + result.put("largest_files", largest); + result.put("recently_indexed", recent); + return result; + }); + } + + private static String extractParentDir(String path) { + int lastSlash = Math.max(path.lastIndexOf('/'), path.lastIndexOf('\\')); + return lastSlash > 0 ? path.substring(0, lastSlash) : "."; + } + + private static Map fileEntry(String name, long size, String path) { + Map m = new LinkedHashMap<>(); + m.put("file_name", name); + m.put("size_bytes", size); + m.put("original_client_path", path); + return m; + } + + private static Map recentEntry(String name, String indexedAt, String path) { + Map m = new LinkedHashMap<>(); + m.put("file_name", name); + m.put("indexed_at", indexedAt); + m.put("original_client_path", path); + return m; + } +} diff --git a/src/main/java/com/javaducker/server/service/SearchService.java b/src/main/java/com/javaducker/server/service/SearchService.java index 62ae564..3c17647 100644 --- a/src/main/java/com/javaducker/server/service/SearchService.java +++ b/src/main/java/com/javaducker/server/service/SearchService.java @@ -3,6 +3,7 @@ import com.javaducker.server.config.AppConfig; import com.javaducker.server.db.DuckDBDataSource; import com.javaducker.server.ingestion.EmbeddingService; +import com.javaducker.server.ingestion.HnswIndex; import org.slf4j.Logger; import org.slf4j.LoggerFactory; import org.springframework.stereotype.Service; @@ -20,6 +21,7 @@ public class SearchService { private final DuckDBDataSource dataSource; private final EmbeddingService embeddingService; private final AppConfig config; + private volatile HnswIndex hnswIndex; public SearchService(DuckDBDataSource dataSource, EmbeddingService embeddingService, AppConfig config) { this.dataSource = dataSource; @@ -27,12 +29,21 @@ public SearchService(DuckDBDataSource dataSource, EmbeddingService embeddingServ this.config = config; } + public void setHnswIndex(HnswIndex index) { + this.hnswIndex = index; + } + + public HnswIndex getHnswIndex() { + return hnswIndex; + } + public List> exactSearch(String phrase, int maxResults) throws SQLException { Connection conn = dataSource.getConnection(); List> results = new ArrayList<>(); try (PreparedStatement ps = conn.prepareStatement(""" - SELECT ac.chunk_id, ac.artifact_id, ac.chunk_index, ac.chunk_text, a.file_name + SELECT ac.chunk_id, ac.artifact_id, ac.chunk_index, ac.chunk_text, + ac.line_start, ac.line_end, a.file_name FROM artifact_chunks ac JOIN artifacts a ON ac.artifact_id = a.artifact_id WHERE a.status = 'INDEXED' @@ -49,6 +60,8 @@ AND LOWER(ac.chunk_text) LIKE LOWER('%' || ? || '%') hit.put("chunk_index", rs.getInt("chunk_index")); hit.put("preview", truncatePreview(rs.getString("chunk_text"), phrase)); hit.put("file_name", rs.getString("file_name")); + hit.put("line_start", rs.getObject("line_start")); + hit.put("line_end", rs.getObject("line_end")); hit.put("score", computeExactScore(rs.getString("chunk_text"), phrase)); hit.put("match_type", "EXACT"); results.add(hit); @@ -64,12 +77,48 @@ public List> semanticSearch(String phrase, int maxResults) t double[] queryEmbedding = embeddingService.embed(phrase); int limit = maxResults > 0 ? maxResults : config.getMaxSearchResults(); + // HNSW fast path + if (hnswIndex != null && !hnswIndex.isEmpty()) { + List annResults = hnswIndex.search(queryEmbedding, limit); + List> results = new ArrayList<>(); + Connection conn = dataSource.getConnection(); + try (PreparedStatement ps = conn.prepareStatement(""" + SELECT ac.chunk_id, ac.artifact_id, ac.chunk_index, ac.chunk_text, + ac.line_start, ac.line_end, a.file_name + FROM artifact_chunks ac + JOIN artifacts a ON ac.artifact_id = a.artifact_id + WHERE ac.chunk_id = ? + """)) { + for (HnswIndex.Result annResult : annResults) { + ps.setString(1, annResult.id()); + try (ResultSet rs = ps.executeQuery()) { + if (rs.next()) { + Map hit = new HashMap<>(); + hit.put("chunk_id", rs.getString("chunk_id")); + hit.put("artifact_id", rs.getString("artifact_id")); + hit.put("chunk_index", rs.getInt("chunk_index")); + String text = rs.getString("chunk_text"); + hit.put("preview", text.length() > 200 ? text.substring(0, 200) + "..." : text); + hit.put("file_name", rs.getString("file_name")); + hit.put("line_start", rs.getObject("line_start")); + hit.put("line_end", rs.getObject("line_end")); + hit.put("score", 1.0 - annResult.distance()); + hit.put("match_type", "SEMANTIC"); + results.add(hit); + } + } + } + } + return results; + } + Connection conn = dataSource.getConnection(); List> results = new ArrayList<>(); // Load all chunk embeddings and compute similarity in Java (brute force for v2) try (PreparedStatement ps = conn.prepareStatement(""" - SELECT ce.chunk_id, ce.embedding, ac.artifact_id, ac.chunk_index, ac.chunk_text, a.file_name + SELECT ce.chunk_id, ce.embedding, ac.artifact_id, ac.chunk_index, ac.chunk_text, + ac.line_start, ac.line_end, a.file_name FROM chunk_embeddings ce JOIN artifact_chunks ac ON ce.chunk_id = ac.chunk_id JOIN artifacts a ON ac.artifact_id = a.artifact_id @@ -90,6 +139,8 @@ public List> semanticSearch(String phrase, int maxResults) t ? rs.getString("chunk_text").substring(0, 200) + "..." : rs.getString("chunk_text")); hit.put("file_name", rs.getString("file_name")); + hit.put("line_start", rs.getObject("line_start")); + hit.put("line_end", rs.getObject("line_end")); hit.put("score", similarity); hit.put("match_type", "SEMANTIC"); results.add(hit); diff --git a/src/main/java/com/javaducker/server/service/StalenessService.java b/src/main/java/com/javaducker/server/service/StalenessService.java new file mode 100644 index 0000000..adf9508 --- /dev/null +++ b/src/main/java/com/javaducker/server/service/StalenessService.java @@ -0,0 +1,97 @@ +package com.javaducker.server.service; + +import com.javaducker.server.db.DuckDBDataSource; +import org.springframework.stereotype.Service; + +import java.io.IOException; +import java.nio.file.Files; +import java.nio.file.Path; +import java.sql.PreparedStatement; +import java.sql.ResultSet; +import java.sql.SQLException; +import java.sql.Timestamp; +import java.time.Instant; +import java.util.*; + +@Service +public class StalenessService { + + private final DuckDBDataSource dataSource; + + public StalenessService(DuckDBDataSource dataSource) { + this.dataSource = dataSource; + } + + public Map checkStaleness(List filePaths) throws SQLException { + List> stale = new ArrayList<>(); + int current = 0; + List notIndexed = new ArrayList<>(); + + for (String path : filePaths) { + if (path == null || path.isBlank()) continue; + + Path filePath = Path.of(path); + boolean fileExists = Files.exists(filePath); + + List> artifacts = queryArtifacts(path); + if (artifacts.isEmpty()) { + notIndexed.add(path); + continue; + } + + if (!fileExists) { + notIndexed.add(path); + continue; + } + + Instant fileMtime; + try { + fileMtime = Files.getLastModifiedTime(filePath).toInstant(); + } catch (IOException e) { + notIndexed.add(path); + continue; + } + + for (Map artifact : artifacts) { + Instant indexedAt = (Instant) artifact.get("indexed_at"); + if (indexedAt == null || fileMtime.isAfter(indexedAt)) { + artifact.put("file_modified_at", fileMtime.toString()); + stale.add(artifact); + } else { + current++; + } + } + } + + Map result = new LinkedHashMap<>(); + result.put("stale", stale); + result.put("current", current); + result.put("not_indexed", notIndexed); + result.put("total_checked", filePaths.stream().filter(p -> p != null && !p.isBlank()).count()); + return result; + } + + private List> queryArtifacts(String path) throws SQLException { + String sql = "SELECT artifact_id, file_name, original_client_path, indexed_at " + + "FROM artifacts WHERE original_client_path = ? AND status = 'INDEXED'"; + + return dataSource.withConnection(conn -> { + List> results = new ArrayList<>(); + try (PreparedStatement ps = conn.prepareStatement(sql)) { + ps.setString(1, path); + try (ResultSet rs = ps.executeQuery()) { + while (rs.next()) { + Map row = new LinkedHashMap<>(); + row.put("artifact_id", rs.getString("artifact_id")); + row.put("file_name", rs.getString("file_name")); + row.put("original_client_path", rs.getString("original_client_path")); + Timestamp ts = rs.getTimestamp("indexed_at"); + row.put("indexed_at", ts != null ? ts.toInstant() : null); + results.add(row); + } + } + } + return results; + }); + } +} diff --git a/src/main/java/com/javaducker/server/service/UploadService.java b/src/main/java/com/javaducker/server/service/UploadService.java index 404de54..32ceda3 100644 --- a/src/main/java/com/javaducker/server/service/UploadService.java +++ b/src/main/java/com/javaducker/server/service/UploadService.java @@ -25,17 +25,27 @@ public class UploadService { private static final Logger log = LoggerFactory.getLogger(UploadService.class); private final DuckDBDataSource dataSource; private final AppConfig config; + private final ArtifactService artifactService; - public UploadService(DuckDBDataSource dataSource, AppConfig config) { + public UploadService(DuckDBDataSource dataSource, AppConfig config, ArtifactService artifactService) { this.dataSource = dataSource; this.config = config; + this.artifactService = artifactService; } public String upload(String fileName, String originalClientPath, String mediaType, long sizeBytes, byte[] content) throws IOException, SQLException { String existing = findExisting(fileName, originalClientPath, sizeBytes); if (existing != null) { - log.info("Dedup: returning existing artifact {} for {} ({} bytes)", existing, fileName, sizeBytes); + log.info("Re-indexing existing artifact {} for {} ({} bytes)", existing, fileName, sizeBytes); + artifactService.deleteArtifactData(existing); + + Path intakePath = storeInIntake(existing, fileName, content); + String sha256 = computeChecksum(content); + + updateArtifactForReindex(existing, sha256, sizeBytes, intakePath.toString(), + fileName, mediaType, originalClientPath); + log.info("Reset artifact {} for re-ingestion ({}), {} bytes", existing, fileName, sizeBytes); return existing; } @@ -92,6 +102,35 @@ public static String computeChecksum(byte[] content) { } } + private void updateArtifactForReindex(String artifactId, String sha256, long sizeBytes, + String intakePath, String fileName, String mediaType, + String originalClientPath) throws SQLException { + // DuckDB UPDATE can fail with PK constraint on ART index — use DELETE+INSERT + dataSource.withConnection(conn -> { + try (PreparedStatement ps = conn.prepareStatement( + "DELETE FROM artifacts WHERE artifact_id = ?")) { + ps.setString(1, artifactId); + ps.executeUpdate(); + } + try (PreparedStatement ps = conn.prepareStatement(""" + INSERT INTO artifacts (artifact_id, file_name, media_type, original_client_path, + intake_path, size_bytes, sha256, status) + VALUES (?, ?, ?, ?, ?, ?, ?, ?) + """)) { + ps.setString(1, artifactId); + ps.setString(2, fileName); + ps.setString(3, mediaType); + ps.setString(4, originalClientPath); + ps.setString(5, intakePath); + ps.setLong(6, sizeBytes); + ps.setString(7, sha256); + ps.setString(8, ArtifactStatus.STORED_IN_INTAKE.name()); + ps.executeUpdate(); + } + return null; + }); + } + private void createArtifactRecord(String artifactId, String fileName, String mediaType, String originalClientPath, String intakePath, long sizeBytes, String sha256) throws SQLException { diff --git a/src/test/java/com/javaducker/integration/FullFlowIntegrationTest.java b/src/test/java/com/javaducker/integration/FullFlowIntegrationTest.java index 6abe479..8b21bda 100644 --- a/src/test/java/com/javaducker/integration/FullFlowIntegrationTest.java +++ b/src/test/java/com/javaducker/integration/FullFlowIntegrationTest.java @@ -52,8 +52,8 @@ static void setUp() throws Exception { config.setEmbeddingDim(128); dataSource = new DuckDBDataSource(config); - uploadService = new UploadService(dataSource, config); artifactService = new ArtifactService(dataSource); + uploadService = new UploadService(dataSource, config, artifactService); textExtractor = new TextExtractor(); textNormalizer = new TextNormalizer(); chunker = new Chunker(); @@ -61,7 +61,8 @@ static void setUp() throws Exception { searchService = new SearchService(dataSource, embeddingService, config); statsService = new StatsService(dataSource); ingestionWorker = new IngestionWorker(dataSource, artifactService, - textExtractor, textNormalizer, chunker, embeddingService, config); + textExtractor, textNormalizer, chunker, embeddingService, + new FileSummarizer(), new ImportParser(), searchService, config); schemaBootstrap = new SchemaBootstrap(dataSource, config, ingestionWorker); schemaBootstrap.bootstrap(); diff --git a/src/test/java/com/javaducker/server/db/SchemaBootstrapTest.java b/src/test/java/com/javaducker/server/db/SchemaBootstrapTest.java index c7dc992..5540cee 100644 --- a/src/test/java/com/javaducker/server/db/SchemaBootstrapTest.java +++ b/src/test/java/com/javaducker/server/db/SchemaBootstrapTest.java @@ -3,6 +3,7 @@ import com.javaducker.server.config.AppConfig; import com.javaducker.server.ingestion.*; import com.javaducker.server.service.ArtifactService; +import com.javaducker.server.service.SearchService; import org.junit.jupiter.api.AfterEach; import org.junit.jupiter.api.BeforeEach; import org.junit.jupiter.api.Test; @@ -39,9 +40,11 @@ void tearDown() { private SchemaBootstrap createBootstrap() { ArtifactService artifactService = new ArtifactService(dataSource); + SearchService searchService = new SearchService(dataSource, new EmbeddingService(config), config); IngestionWorker worker = new IngestionWorker(dataSource, artifactService, new TextExtractor(), new TextNormalizer(), new Chunker(), - new EmbeddingService(config), config); + new EmbeddingService(config), new FileSummarizer(), new ImportParser(), + searchService, config); return new SchemaBootstrap(dataSource, config, worker); } diff --git a/src/test/java/com/javaducker/server/ingestion/IngestionWorkerParallelTest.java b/src/test/java/com/javaducker/server/ingestion/IngestionWorkerParallelTest.java index 5bce862..3026ac7 100644 --- a/src/test/java/com/javaducker/server/ingestion/IngestionWorkerParallelTest.java +++ b/src/test/java/com/javaducker/server/ingestion/IngestionWorkerParallelTest.java @@ -5,6 +5,7 @@ import com.javaducker.server.db.SchemaBootstrap; import com.javaducker.server.model.ArtifactStatus; import com.javaducker.server.service.ArtifactService; +import com.javaducker.server.service.SearchService; import com.javaducker.server.service.UploadService; import org.junit.jupiter.api.*; import org.junit.jupiter.api.io.TempDir; @@ -39,10 +40,12 @@ void setUp() throws Exception { dataSource = new DuckDBDataSource(config); artifactService = new ArtifactService(dataSource); - uploadService = new UploadService(dataSource, config); + uploadService = new UploadService(dataSource, config, artifactService); + SearchService searchService = new SearchService(dataSource, new EmbeddingService(config), config); ingestionWorker = new IngestionWorker(dataSource, artifactService, new TextExtractor(), new TextNormalizer(), new Chunker(), - new EmbeddingService(config), config); + new EmbeddingService(config), new FileSummarizer(), new ImportParser(), + searchService, config); SchemaBootstrap bootstrap = new SchemaBootstrap(dataSource, config, ingestionWorker); bootstrap.bootstrap(); diff --git a/src/test/java/com/javaducker/server/rest/JavaDuckerRestControllerTest.java b/src/test/java/com/javaducker/server/rest/JavaDuckerRestControllerTest.java index a22cac2..1bc552c 100644 --- a/src/test/java/com/javaducker/server/rest/JavaDuckerRestControllerTest.java +++ b/src/test/java/com/javaducker/server/rest/JavaDuckerRestControllerTest.java @@ -1,10 +1,8 @@ package com.javaducker.server.rest; import com.fasterxml.jackson.databind.ObjectMapper; -import com.javaducker.server.service.ArtifactService; -import com.javaducker.server.service.SearchService; -import com.javaducker.server.service.StatsService; -import com.javaducker.server.service.UploadService; +import com.javaducker.server.ingestion.FileWatcher; +import com.javaducker.server.service.*; import org.junit.jupiter.api.Test; import org.springframework.beans.factory.annotation.Autowired; import org.springframework.boot.test.autoconfigure.web.servlet.WebMvcTest; @@ -30,6 +28,10 @@ class JavaDuckerRestControllerTest { @MockBean ArtifactService artifactService; @MockBean SearchService searchService; @MockBean StatsService statsService; + @MockBean ProjectMapService projectMapService; + @MockBean StalenessService stalenessService; + @MockBean DependencyService dependencyService; + @MockBean FileWatcher fileWatcher; @Test void healthReturnsOk() throws Exception { diff --git a/workflows/bug-fix.md b/workflows/bug-fix.md new file mode 100644 index 0000000..2dd7031 --- /dev/null +++ b/workflows/bug-fix.md @@ -0,0 +1,29 @@ +# Bug Fix Workflow + +## Step 1: Investigate (parallel) +Spawn parallel agents in ONE message: +- **Agent A**: Reproduce the bug — find failing test case or steps to trigger +- **Agent B**: Search codebase — grep for related patterns, read recent git log for the area + +Wait for both. Combine findings. + +## Step 2: Locate root cause +Trace from symptoms to root cause using investigation results. Read the code path involved. + +## Step 3: Fix +Make the minimal change that addresses the root cause, not the symptom. + +## Step 4: Verify (parallel) +Run in ONE message: +- Add a regression test that fails without the fix and passes with it +- Run the full test suite + +## Step 5: Document +Update `context/MEMORY.md` with what was found and fixed. + +## If fix doesn't work → closed loop +If the test still fails after the fix: +1. Re-analyze — was the root cause wrong? +2. Try a different fix +3. Re-run tests +4. Max 3 attempts before escalating to the user diff --git a/workflows/closed-loop.md b/workflows/closed-loop.md new file mode 100644 index 0000000..f7c48ff --- /dev/null +++ b/workflows/closed-loop.md @@ -0,0 +1,95 @@ +# Closed-Loop Workflow + +A repeating pipeline that runs checks, fixes issues, and re-checks until all pass or max iterations reached. + +## Prerequisites + +Before starting, you need: +- A **check command** — a script or test that produces a machine-readable report (JSON) +- A **pass condition** — what "done" looks like (e.g., 0 issues, all tests pass) +- A **max iterations** cap (default: 10) + +## Loop Protocol + +### 1. Capture baseline +``` +Run the check command → parse results → record in context/MEMORY.md: + Iteration: 0 (baseline) + Total issues: N + Pass rate: X/Y + Issue breakdown by category +``` + +### 2. Analyze +- Read the report output (JSON, test results, screenshots, etc.) +- Categorize issues by type and fix approach +- Group into independent fix batches — each batch becomes a parallel agent + +### 3. Fix (PARALLEL) +Spawn one Agent per independent fix category in a SINGLE message: +``` +Agent 1: "Fix all [category-A] issues. Files: [list]. Read each file first." +Agent 2: "Fix all [category-B] issues. Files: [list]. Read each file first." +Agent 3: "Fix all [category-C] issues. Files: [list]. Read each file first." +``` +All agents run with `run_in_background: true`. Wait for ALL to complete. + +### 4. Review agent results +- Read ALL agent outputs before proceeding +- Check for conflicts (two agents editing the same file/region) +- If conflicts exist, resolve them before re-checking + +### 5. Re-check +Run the check command again. Compare to previous iteration: +- **Improved** (fewer issues) → continue +- **Regression** (more issues) → revert immediately, log what failed, try different approach +- **Same count** but different issues → continue (progress is being made) +- **All pass** → go to Confirm + +### 6. Log iteration +Append to `context/MEMORY.md`: +``` +### Iteration N +- Pass: X/Y +- Issues: N (was M) +- Fixed: [categories] +- Regressed: [none | categories] +- Key changes: [summary] +``` + +### 7. Loop or exit +- Issues remain AND iteration < max → go to step 2 +- All pass → go to Confirm +- Max iterations reached → stop, report remaining issues + +### 8. Confirm +- Run check one final time to verify +- If visual/manual inspection is needed, do it now +- Copy confirmed outputs to final location +- Write final summary to `context/MEMORY.md` + +## Anti-patterns to avoid + +- **Don't fix everything at once** — group by category, fix in parallel batches +- **Don't retry the same fix** — if it regressed, try a different approach +- **Don't skip the re-check** — always verify before continuing +- **Don't run agents sequentially** — if fixes are independent, they MUST be parallel +- **Don't keep going past max iterations** — diminishing returns, stop and report + +## Prompting this workflow + +Tell Claude: +``` +Follow workflows/closed-loop.md. +Check command: [your command here] +Pass condition: [what success looks like] +Max iterations: [N] +``` + +Example: +``` +Follow workflows/closed-loop.md. +Check command: npx playwright test tests/qa-audit.spec.js --reporter=json +Pass condition: 0 failures in the JSON report +Max iterations: 10 +``` diff --git a/workflows/code-review.md b/workflows/code-review.md new file mode 100644 index 0000000..774b60c --- /dev/null +++ b/workflows/code-review.md @@ -0,0 +1,16 @@ +# Code Review Workflow + +1. **Read the diff holistically** — Understand the full change before commenting on details. +2. **Check each dimension:** + - Correctness — Does it do what it's supposed to? + - Security — Any injection, auth, or data exposure risks? + - Performance — Any unnecessary loops, queries, or allocations? + - Readability — Can someone else understand this in 6 months? + - Maintainability — Is it easy to change later? +3. **Rate issues by severity:** + - **Blocker** — Will cause bugs, security issues, or data loss. Must fix. + - **Major** — Significant design or logic issue. Should fix. + - **Minor** — Improvement opportunity. Fix if convenient. + - **Nit** — Style preference. Optional. +4. **Note positives** — Acknowledge good patterns and decisions. +5. **Verdict** — Approve, Approve with comments, or Request changes. diff --git a/workflows/new-feature.md b/workflows/new-feature.md new file mode 100644 index 0000000..4609b94 --- /dev/null +++ b/workflows/new-feature.md @@ -0,0 +1,32 @@ +# New Feature Workflow + +## Step 1: Understand (parallel) +Spawn parallel agents in ONE message: +- **Agent A**: Read requirements + check `context/MEMORY.md` and `context/CONVENTIONS.md` +- **Agent B**: Explore existing code — find related files, patterns, interfaces to extend + +Wait for both. Combine into implementation plan. + +## Step 2: Design +Identify affected files, interfaces, and data flow. Use `/architect` for complex designs. +Split implementation into independent pieces that can be built in parallel. + +## Step 3: Implement (parallel) +Spawn one Agent per independent piece in ONE message: +- Each agent gets: specific files to create/edit, interfaces to implement, conventions to follow +- Each agent reads existing files before editing +- All agents run with `run_in_background: true` + +Wait for all. Review for conflicts. + +## Step 4: Test (parallel) +Run in ONE message: +- Write unit tests for new logic +- Write integration tests for boundaries +- Run full test suite + +## Step 5: Review +Self-review against `/reviewer` checklist: correctness, security, readability. + +## Step 6: Document +Update `context/MEMORY.md`. Add ADRs to `context/DECISIONS.md` if architectural. diff --git a/workflows/refactor.md b/workflows/refactor.md new file mode 100644 index 0000000..8c92e29 --- /dev/null +++ b/workflows/refactor.md @@ -0,0 +1,31 @@ +# Refactoring Workflow + +## Step 1: Assess (parallel) +Spawn parallel agents in ONE message: +- **Agent A**: Identify refactoring targets — duplication, complexity, unclear naming, tight coupling +- **Agent B**: Run existing tests to establish passing baseline, note coverage gaps + +Wait for both. + +## Step 2: Plan +Group refactoring targets into independent batches. Each batch should be: +- Independently testable +- Non-conflicting with other batches (different files or non-overlapping regions) + +## Step 3: Refactor (parallel) +Spawn one Agent per independent batch in ONE message: +- Each agent makes one type of structural change +- Each agent runs tests after their changes +- All agents run with `run_in_background: true` + +Wait for all. Check for conflicts between agents. + +## Step 4: Verify (closed loop) +Run full test suite. If failures: +1. Identify which batch broke tests +2. Fix or revert that batch only +3. Re-run tests +4. Max 3 iterations + +## Step 5: Clean up +Remove dead code, update imports, delete unused files.