imgsearch is a local-first image organization and similarity search app.
It runs as a small local web app that:
- indexes images into a SQLite-backed library,
- supports text-to-image and image-to-image search,
- keeps data on your machine,
- uses a built-in default multimodal embedding setup.
- Download the latest archive from the GitHub
rollingrelease for your system (the name ends with the commit SHA). - Extract it.
- Linux: run
./run.sh. The archive bundleslibvips, the llama.cpp libraries,sqlite-vector, and the ONNX Runtime for video transcription; the wrapper sets the library path and enables transcription. - macOS: install
libvips(for examplebrew install vips), then run./run.sh(or./imgsearchwithout video transcription). - Open
http://127.0.0.1:8080/.
On first run, imgsearch downloads the default 2B embedding model into ./models/VesNFF/Qwen3-VL-Embedding-2B-GGUF/ if it is missing, and also downloads the default Gemma e4b annotator files when annotations are enabled.
With the app running, import a folder recursively:
./scripts/import_images.sh ~/Pictures/memesOr any other folder:
./scripts/import_images.sh /path/to/your/imagesThe import script uploads supported images to the local app and indexing continues in the background.
You can also import full-size pictures and webms from a 4chan thread URL:
./scripts/import_images.sh https://boards.4chan.org/v/thread/737156945For 4chan thread imports, the script pulls full files from i.4cdn.org (not thumbnails) and currently imports supported thread pictures plus .webm files.
If 4chan rate-limits requests (HTTP 429), the importer retries with Retry-After support.
If needed, tune retry behavior with IMGSEARCH_IMPORT_HTTP_MAX_ATTEMPTS and IMGSEARCH_IMPORT_HTTP_RETRY_DELAY_SECONDS.
By default, 4chan media downloads are paced (about every 5 seconds with jitter) to reduce rate-limit spikes.
scripts/import_images.sh / mise run import-images now sends an API key header by default.
Set IMGSEARCH_IMPORT_API_KEY (or IMGSEARCH_API_KEY) to override the built-in development key.
If there is no release for your system:
- Install Go, CMake, and
libvips. - Initialize the llama.cpp submodule:
git submodule update --init --recursive deps/llama.cpp
- Build llama.cpp runtime libraries:
./scripts/ensure_llama_cpp_native_build.sh
- Install sqlite-vector:
./scripts/setup_sqlite_vector.sh
- Run the app:
go run ./cmd/imgsearch
Cross-platform note:
- Keep host-native llama.cpp artifacts in
./deps/llama.cpp/buildonly. - If you build Linux artifacts from Docker on macOS, write them to an explicit separate directory such as
./build-artifacts/llama.cpp/linux-cuda13/and passIMGSEARCH_LLAMA_LIB_DIR=/absolute/path/to/.../binwhen packaging.
If you want to run on a CUDA host through Podman while keeping an Ubuntu userspace, use Containerfile.cuda.
Quick path (host-accessible):
podman build -f Containerfile.cuda -t imgsearch:cuda .
podman run -d --name imgsearch --replace --gpus=all -p 8080:8080 \
-e IMGSEARCH_ADDR=0.0.0.0:8080 \
-e IMGSEARCH_API_KEY='replace-with-a-strong-token' \
-v "$HOME/imgsearch-data:/data" \
-v "$HOME/imgsearch-models:/models" \
imgsearch:cudaThe container defaults to loopback-only bind (127.0.0.1:8080) for safer startup.
Set IMGSEARCH_ADDR=0.0.0.0:8080 only when you intentionally want remote access, keep IMGSEARCH_API_KEY set, and place the service behind a trusted reverse proxy/TLS boundary.
It also disables llama.cpp CUDA graphs by default to avoid observed 26B video-annotation OOMs on 24 GiB cards; set IMGSEARCH_CUDA_GRAPHS=1 to opt back in.
Full instructions are in docs/podman-cuda-ubuntu.md.
Model choice matters more than most runtime knobs. If you are memory constrained, stay on the default smaller search model and disable annotations before tuning GPU layers or batch size.
Search embedding models:
| Model | Dimensions | Best For | Download |
|---|---|---|---|
| Qwen3-VL-Embedding-2B-Q6_K | 2048 | Default profile for CPU-only, low-VRAM, and modest unified-memory systems | Auto-downloaded at default paths |
| Qwen3-VL-Embedding-8B-Q4_K_M | 4096 | Higher-quality search profile for good GPUs and high-memory unified-memory systems | Auto-downloaded when the 8B paths below are selected |
Annotation models:
| Model | Variant | Best For | Download |
|---|---|---|---|
| Gemma-4-E4B-Uncensored-HauhauCS-Aggressive-Q4_K_P | e4b |
Default descriptions/tags with lower memory than 26B | Auto-downloaded when annotations are enabled |
| gemma-4-26b-a4b-it-heretic.q4_k_m | 26b |
Richer descriptions/tags on high-memory systems | Auto-downloaded with -llama-native-annotator-variant 26b |
Vision models require two files: the base .gguf and the matching mmproj-*.gguf. Do not mix 2B and 8B model files, and keep -llama-native-dimensions matched to the embedding model (4096 for 8B, 2048 for 2B).
Default 2B search files are downloaded automatically when these paths are missing:
./models/VesNFF/Qwen3-VL-Embedding-2B-GGUF/Qwen3-VL-Embedding-2B-Q6_K.gguf
./models/VesNFF/Qwen3-VL-Embedding-2B-GGUF/mmproj-Qwen3-VL-Embedding-2B-f16.gguf
Manual pre-download for the default 2B search model:
mkdir -p ./models/VesNFF/Qwen3-VL-Embedding-2B-GGUF
curl -L -o ./models/VesNFF/Qwen3-VL-Embedding-2B-GGUF/Qwen3-VL-Embedding-2B-Q6_K.gguf \
https://huggingface.co/VesNFF/Qwen3-VL-Embedding-2B-GGUF/resolve/main/Qwen3-VL-Embedding-2B-Q6_K.gguf
curl -L -o ./models/VesNFF/Qwen3-VL-Embedding-2B-GGUF/mmproj-Qwen3-VL-Embedding-2B-f16.gguf \
https://huggingface.co/VesNFF/Qwen3-VL-Embedding-2B-GGUF/resolve/main/mmproj-Qwen3-VL-Embedding-2B-f16.ggufThe larger 8B search model is also auto-downloaded when you run with the 8B model path and -llama-native-dimensions 4096. Optional manual pre-download:
mkdir -p ./models/Qwen
curl -L -o ./models/Qwen/Qwen3-VL-Embedding-8B-Q4_K_M.gguf \
https://huggingface.co/lainsoykaf/Qwen3-VL-Embedding-8B-GGUF/resolve/main/Qwen3-VL-Embedding-8B-Q4_K_M.gguf
curl -L -o ./models/Qwen/mmproj-Qwen3-VL-Embedding-8B-f16.gguf \
https://huggingface.co/lainsoykaf/Qwen3-VL-Embedding-8B-GGUF/resolve/main/mmproj-Qwen3-VL-Embedding-8B-f16.ggufThe default e4b and optional 26b annotators are also downloaded automatically when their default paths are used. Skip annotation downloads and model loading with:
./imgsearch -enable-annotations=falseManual pre-download for the default e4b annotator:
mkdir -p ./models/HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive
curl -L -o ./models/HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf \
https://huggingface.co/HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive/resolve/main/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf
curl -L -o ./models/HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive/mmproj-Gemma-4-E4B-Uncensored-HauhauCS-Aggressive-f16.gguf \
https://huggingface.co/HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive/resolve/main/mmproj-Gemma-4-E4B-Uncensored-HauhauCS-Aggressive-f16.ggufManual pre-download for the 26b annotator:
mkdir -p ./models/nohurry/gemma-4-26B-A4B-it-heretic-GUFF
curl -L -o ./models/nohurry/gemma-4-26B-A4B-it-heretic-GUFF/gemma-4-26b-a4b-it-heretic.q4_k_m.gguf \
https://huggingface.co/nohurry/gemma-4-26B-A4B-it-heretic-GUFF/resolve/main/gemma-4-26b-a4b-it-heretic.q4_k_m.gguf
curl -L -o ./models/nohurry/gemma-4-26B-A4B-it-heretic-GUFF/gemma-4-26B-A4B-it-heretic-mmproj.f16.gguf \
https://huggingface.co/nohurry/gemma-4-26B-A4B-it-heretic-GUFF/resolve/main/gemma-4-26B-A4B-it-heretic-mmproj.f16.ggufIf automatic downloads fail, use the direct URLs in cmd/imgsearch/default_model_assets.go to pre-stage the files under ./models/.
imgsearch versions the embedding configuration. Changing the search model path, mmproj path, embedding dimensions, image size, image token cap, or retrieval instructions creates a new model version.
On startup after a model/config change, the app removes old embeddings and queues the library for re-embedding. Images and videos stay in the library, but search results are incomplete until the worker catches up. For large libraries, change model settings during idle time or back up ./data first.
Use the 8B search model plus the default e4b annotator:
./imgsearch \
-llama-native-model-path ./models/Qwen/Qwen3-VL-Embedding-8B-Q4_K_M.gguf \
-llama-native-mmproj-path ./models/Qwen/mmproj-Qwen3-VL-Embedding-8B-f16.gguf \
-llama-native-dimensions 4096From a source checkout, the matching developer command is:
mise run "serve:8b"This trades higher memory use for better search quality while keeping image/video search, background indexing, and generated descriptions/tags.
The default search model is already 2B. Disable annotations first if memory is tight; it is usually the biggest remaining memory reduction.
./imgsearch -enable-annotations=falseFor CPU-only, add:
-llama-native-use-gpu=false -llama-native-gpu-layers 0If a low-VRAM GPU still runs out of memory with the 2B model, add smaller runtime knobs:
-llama-native-gpu-layers 20 -llama-native-batch-size 128 -llama-native-image-max-side 320For mise run serve, no 2B embedder override is needed. To skip annotations from a source checkout, run the binary directly:
go run ./cmd/imgsearch -enable-annotations=falseUse direct ./imgsearch or go run ./cmd/imgsearch when you need flags that are not wired into a mise task.
Use this when you mainly care about similarity/text search and want to avoid loading any annotation model:
./imgsearch -enable-annotations=falseThis still embeds images and videos for search. It skips generated descriptions/tags, which is usually the largest memory and latency reduction after choosing the smaller search model.
On larger GPUs or high-memory unified-memory systems, you can try the 26B annotator for richer descriptions:
./imgsearch -llama-native-annotator-variant 26bFrom a source checkout:
mise run "serve:8b:annotator-26b"This is the heaviest local profile. If interactive search latency matters, run the UI/API without annotations and run a worker separately when you want to backfill annotations:
./imgsearch -mode=api -enable-annotations=false
./imgsearch -mode=worker -llama-native-annotator-variant 26bBoth processes must point at the same -data-dir if you split them. On a single GPU, split mode can still increase total memory if API and worker run at the same time; if memory is tight, run the worker as a batch backfill job and stop it before latency-sensitive searches.
| Symptom | First change to try |
|---|---|
| GPU out of memory on startup | Use the 2B search model, which is the default, or add -enable-annotations=false |
| GPU out of memory while embedding | Lower -llama-native-gpu-layers, then lower -llama-native-batch-size |
| System memory pressure on CPU | Use the 2B model, disable annotations, and lower -llama-native-image-max-side |
| Indexing is too slow but stable | Raise -llama-native-gpu-layers or -llama-native-batch-size one step at a time |
| Descriptions/tags are not needed | Keep -enable-annotations=false permanently |
Descriptions, titles, summaries, and tags come from an annotation backend that you can change at runtime from the Atelier Settings page or through the API, without restarting:
- Native runs the bundled Gemma annotator in-process. Pick the
e4bvariant (default, lower memory) or26b(richer output, high-memory systems). Switching variants unloads and reloads the GGUF files, downloading them on first use. If you pinned custom paths with-llama-native-annotator-model-path/-llama-native-annotator-mmproj-path, the variant selector is locked to those files. - Remote server sends images to any OpenAI-compatible chat-completions endpoint with vision support, such as llama-server, Ollama, LM Studio, vLLM, OpenAI, or OpenRouter. Configure the base URL (for example
http://127.0.0.1:11434/v1), an optional API key, the model name, a request timeout, and the number of parallel requests. While a remote backend is selected, the native annotator is unloaded and no GGUF download happens at startup.
Existing annotations are kept when you switch. Use "Re-annotate all" to refresh the library with the new backend, keeping in mind that paid remote APIs bill per image.
The same settings are available at GET/PUT /api/settings, and POST /api/settings/annotation/test checks a remote server without saving. The worker picks up saved changes between jobs, also when it runs as a separate -mode=worker process sharing the same -data-dir.
imgsearch keeps a single trust boundary: the network address it binds to.
- The default bind is loopback (
127.0.0.1:8080). If you can reach the address, you are trusted to use both the UI and the API. - Any non-API page load mints an
imgsearch_api_keycookie that the API auth middleware then accepts as equivalent to the configured API key. This is intentional: the UI is the only way to use the app, so locking the API behind a separate token from the UI would block the UI itself. The cookie isHttpOnly,SameSite=Strict, andSecurewhen served over TLS. - API clients that are not a browser (importers, scripts,
curl) authenticate withX-Imgsearch-API-Key: <token>orAuthorization: Bearer <token>. Both are honored independently of cookies. - Stored media under
/media/images/and/media/videos/sits behind the same trust boundary as/api/: requests need the cookie or an API key header, and even authenticated directory paths answer404so the library can never be listed. - Binding to a non-loopback address (
-addr 0.0.0.0:8080or a public hostname) requires an explicit strong API key via-api-key/IMGSEARCH_API_KEY; the built-in development key is rejected at startup, and aWARNINGis logged that anyone who can reach the address can use the API. - Settings saved from the UI, including a remote annotation server API key, are stored in plaintext in the local SQLite database under
./data. Treat the data directory as sensitive. - If you expose the app beyond your own machine, put it behind a trusted reverse proxy that terminates TLS, enforces auth if needed, and exposes a stable hostname (the
Set-Cookieis keyed to the configured API key, not the hostname, so a hostile same-origin page on the same hostname inherits the same trust).
This is documented in meta/issues/054-harden-ui-api-cookie-auth.md so the trust boundary stays explicit.
- The app binds to
127.0.0.1:8080by default. - Supported formats: JPEG, PNG, WEBP, and AVIF images; MP4, WebM, QuickTime, and Matroska videos.
scripts/import_images.shalso converts GIFs to MP4 with ffmpeg. - Embedding uses the in-process
llama-cpp-nativeruntime with the Qwen3-VL-Embedding-2B GGUF pair by default. - The default Qwen 2B embedding files and the default Gemma
e4bannotator files are downloaded automatically on first run when missing. - Add
-enable-annotations=falseif you want to run the API without loading the Gemma annotation model. - Add
-mode=apior-mode=workerif you want to split the HTTP server and background worker into separate processes. /api/*routes are authenticated by default.- Set
-api-key <token>(orIMGSEARCH_API_KEY) to use your own key; when unset, the server falls back to a built-in development key and logs a startup warning. - If you bind to a non-loopback address (for example
-addr 0.0.0.0:8080), startup requires an explicit strong API key; the built-in development key is rejected. - API clients can authenticate with
X-Imgsearch-API-Key: <token>orAuthorization: Bearer <token>. - Multipart uploads to
/api/uploadkeep partial-success semantics: each uploaded file returns either IDs/digest data or anerror, and mixed success/failure batches return207 Multi-Status. - Upload limits are per file: images up to 64 MiB and videos up to 2048 MiB by default (
-max-image-upload-mb,-max-video-upload-mb). An oversized file rejects the whole request with413 Payload Too Largeand a JSON body naming the file, its media type, andlimit_bytes. One upload request may run for up to-upload-timeout(default 30m) regardless of the server-wide read/write timeouts. - Video transcription (Parakeet via ONNX Runtime) is on when
-parakeet-onnxruntime-libpoints at an ONNX Runtime shared library. Release archives bundle it and therun.shwrappers pass it; the Parakeet model bundle downloads on first run. Without the flag the startup log prints one line saying transcription is disabled. - Data is stored in
./databy default. - The UI includes uploads, indexing status, gallery browsing, text and tag search, similar-image search (including by a pasted or dropped picture), a near-duplicate finder, manual title and tag editing, a similar-video Feed, and an annotation settings page.
Development-focused setup, tasks, integration checks, and lower-level runtime notes are in:
docs/development.mddocs/architecture.mddocs/mvp-plan.mddocs/decisions.md
