Turn your Android phone into a local LLM server for your home network, or chat privately on-device with text, images, PDFs, audio, and offline speech-to-text.
Chat in the app, open its Web UI on your computer, or connect a compatible OpenAI client while the model runs on your phone.
π Watch the LAN demo or connect scripts and scheduled workflows.
Install the small base APK, then download a built-in model or import your own LiteRT model (beta).
Changes since v1.5.0:
Run the model on your phone and use it from other devices on your home network; see the demo or client guide.
- π LAN and browser chat: Open the phone's password-protected Web UI on your PC while inference stays on the phone.
- ποΈ Browser history and uploads: Save, reopen, or delete chats on the phone, and upload PDFs or images from your PC.
- π OpenAI-compatible API: Connect Open WebUI, example scripts, or compatible harnesses; OpenCode tool-call round trips tested.
- βοΈ Context settings: Adjust context length per LiteRT model.
- π§© Your own model (beta): Import a local
.litertlmfile for text chat, with native image and audio input attempted when supported. - π Documents: Attach text files and PDFs, including scanned pages read with OCR.
- π Document answers: Use simple BM25 retrieval for questions or process all source chunks for summaries.
- ποΈ Speech-to-text: Transcribe offline with multilingual Whisper, including German, and edit before sending.
- π Audio attachments: Upload or record audio for native understanding on compatible models or Whisper transcription.
- β³ Long inputs: Added segmented audio processing and cancellable attachment jobs with progress.
- π¨ Browser controls: Added light/dark themes, Enter to send, and Shift+Enter for a new line.
- β‘ Runtime: Updated and pinned LiteRT-LM to 0.17.1.
- π‘οΈ Model loading: Improved file validation, downloads, cancellation, and recovery after failed initialization.
- πΌοΈ Image input: Removed FastVLM descriptions; images now use OCR or native Gemma vision.
- π§© Custom models: Beta compatibility varies by model and phone; native tool calling is unavailable.
β οΈ Memory: Large models and contexts can still fail to load or crash, including during long OpenCode sessions.- πΌοΈ Browser images: Direct image uploads require a compatible model.
- π Coming next: Improved document retrieval; this release uses lightweight BM25.
β‘οΈ See all releases
![]() Chat Inference |
![]() Image Support |
![]() New UI |
Figure: Pocket LLM showing offline chat, image input, and the updated Android UI.
- π± Local inference: Run models on your phone with ONNX or LiteRT, offline after installation.
- π¬ Chat: Stream replies, keep multi-turn history, render Markdown, and copy responses.
- ποΈ History: Reopen or delete saved conversations in the Android app or browser.
- π¨ Appearance: Choose light/dark mode, accent colors, and chat font size in the Android app.
- π Privacy: Keep inference and chat data on your phone, with no telemetry.
- π§ Built-in models: Choose Qwen2.5, Qwen3, DeepSeek R1 Distill Qwen, or Gemma 4.
- π¦ Model management: Download, switch, and delete models inside the app.
- π§© Custom models (beta): Import your own
.litertlmmodel from device storage; native image and audio availability depend on the model and device backend. - π Thinking mode: Toggle reasoning on supported Qwen3 and Gemma models.
- ποΈ Instructions: Edit model instructions or choose a prompt preset.
- βοΈ Context length: Set each LiteRT model's context size to balance history and memory use.
- π‘οΈ Recovery: Validate model files and recover from interrupted loading without losing saved chats.
- ποΈ Dictation: Turn recordings into editable text with offline multilingual Whisper, including German.
- π Audio: Use native audio understanding on compatible Gemma or imported models, or Whisper transcription when native audio is unavailable.
- π Documents: Attach text files and PDFs, including scanned PDFs read with OCR.
- π Retrieval: Ask document questions with BM25 or summarize source chunks within the context budget.
- π References: Keep document, page, or timestamp references with the owning chat.
- β³ Progress: Track or cancel long document extraction and audio transcription jobs.
- πΌοΈ Images: Extract text with OCR or send images directly to compatible Gemma or imported models.
- π· Camera: Capture, retake, crop, review, and send photos.
- π LAN server: Serve the phone's selected model to devices on the same private network.
- π Local AI home server: Use built-in or compatible imported LiteRT models as a shared inference endpoint for scheduled jobs and monitoring workflows.
- π» Web UI: Chat in a PC browser, upload PDFs or images, and save conversations on the phone.
- π OpenAI-compatible API: Connect compatible clients through model and chat-completion endpoints.
- π οΈ Agent tools: Return native tool calls from supported LiteRT models for harnesses such as OpenCode to execute.
![]() Start the server on your phone |
![]() Use the model from your browser Β· View demo GIF |
The app ships as a single smaller base APK.
β‘οΈ Download APK
Models are not bundled inside the APK. After installation, choose and download the models you want directly on device.
You can download multiple models, switch between them inside the app, and delete unused downloaded models later to free storage.
- Gemma 4 E4B LiteRT - Best for flagship mobiles
- Gemma 4 E2B LiteRT - Best for decent to mid-range mobiles
- DeepSeek R1 Distill Qwen 1.5B LiteRT - Optional reasoning model for high-RAM mobiles
- Qwen3 0.6B LiteRT - Best for low-end mobiles
- Qwen3 0.6B Q4F16 ONNX - Good for low to mid-range mobiles
- Qwen2.5 0.5B ONNX - Best for mid to high-end mobiles, full precision
- Open Manage Models β use your own model (beta) and choose a local
.litertlmfile. - The app copies the file into private storage and adds it to the model picker.
- The app detects declared inputs and attempts native image and audio support when available. The ready message reports availability on the current device/backend; declared support is not a compatibility guarantee.
- Text chat is required. Images can use OCR or Native image input when available; audio uses native input when available or optional Whisper transcription otherwise.
- Video input, generated audio, and custom thinking controls are not enabled. See the custom-model beta guide for details and device checks.
- Custom models do not expose native tool calling through the LAN API.
- Custom-model compatibility is experimental; larger models or contexts may fail to load or crash.
- OCR mode - Extract text from images
- Native image input - Send images directly to compatible Gemma or imported models
- Camera capture - Take a photo, retake, crop, review, and send it as input
Note: internet is required only for downloading models. Chat, OCR, image input, camera workflows, and inference remain fully on-device after the required models are installed.
- Text: Attach UTF-8 or UTF-16 text files up to 16 MiB.
- PDFs: Attach embedded-text or scanned PDFs up to 64 MiB and 500 pages in the Android app.
- PDF OCR: Extract embedded text first and use on-device OCR only for pages that need it.
- Questions: Use lightweight BM25 to select relevant passages that fit the model context.
- Summaries: Process all source chunks for summaries and ordered transformations.
- References: Prompt answers to retain page, section, or timestamp markers.
- Audio files: Attach Android-decodable audio up to 512 MiB and two hours, or record audio in-app.
- Native audio: Use compatible Gemma or imported models for native audio understanding, with overlapping segments for longer recordings. Imported models use segments of at most 10 seconds.
- Whisper fallback: When native audio is unavailable, transcribe with multilingual Whisper tiny int8 after a separate download of about 100 MB.
- Dictation: Tap the microphone, record, transcribe, and edit the recognized text before sending.
- Attachment limits: Keep one active document or audio source per Android chat; do not mix it with images in one send.
- Storage: Detaching keeps the source data; deleting its chat removes it from private app storage.
- Progress: Long extraction and transcription jobs show progress and can be cancelled.
- Load a model, open LAN Server in the Android app, enter a password of at least eight characters, and tap Save password.
- Tap Start, then open the displayed
/uiURL on a computer on the same Wi-Fi or LAN. - For API clients, use
http://PHONE_IP:8080/v1and the same password.
- Save, reopen, and delete browser conversations stored privately on the phone.
- Upload PDFs or images; direct images require a compatible model.
- Switch light/dark themes; use Enter to send and Shift+Enter for a new line.
- Stop the LAN server before changing the model or password.
π Client setup and examples: Python, HTTP scripts, PowerShell, OpenCode, Open WebUI, and automation.
LiteRT models expose a context-length field in the Android app's model settings. This is a device-memory trade-off: a larger context can require substantially more native and GPU memory.
- Recommended tested baseline: 8K tokens.
- Limited manual testing: 10K to 15K worked on the tested device, but this is not broad device validation.
- Configurable maximum: Qwen LiteRT models allow up to 40K and Gemma LiteRT models up to 128K. Those values are configuration ceilings, not guarantees that a phone can initialize or run safely at that size.
The app warns above 8K. A native LiteRT-LM crash was observed with Gemma 4 E2B at 20K on a tested device, consistent with high-context memory pressure. Keep the context near 8K unless you have tested the selected model on your own device.
This app supports ONNX-based Qwen models and LiteRT-based Qwen 3 and Gemma 4 models.
- ONNX backend: supports Qwen2.5 and Qwen3
- LiteRT backend: supports Qwen3, DeepSeek R1 Distill Qwen, Gemma 4, and imported
.litertlmmodels (beta), with native image/audio input when available
- Qwen3 and Gemma 4 support Thinking Mode
- The toggle is shown only for models that support it
LiteRT is a strong fit for fast local Android chat because:
- It is designed for high-performance on-device LLM deployment
- It supports hardware acceleration, including GPU and NPU acceleration on supported devices
- It helps reduce startup and generation latency for local chat workloads
- It expands the range of practical Android model builds beyond a single backend path
- It fits well with a privacy-first app design focused on fully offline usage
Note: model capability and performance still depend on the specific model build and the hardware of the target Android device.
- Android Studio
- A physical Android device for deployment and testing
- 4 GB or more RAM for smaller models
- More RAM is recommended for larger models such as Gemma 4 E2B and Gemma 4 E4B
- A temporary internet connection for downloading models inside the app
- Real hardware is preferred; emulators are mainly useful for UI checks
Before a model is loaded, Pocket LLM checks available storage and validates the downloaded files. Memory estimates are advisory only: they do not block loading or require confirmation because device memory reports cannot reliably predict LiteRT backend compatibility. The app automatically attempts initialization with the available backend fallbacks.
The app records when initialization starts and clears that record only after the model is ready. If initialization fails or the app stops during loading, the same model is not retried automatically on the next launch. A recovery prompt lets you choose another model, delete the failed files, retry manually, or continue without a loaded model. Saved chat history is kept independently and is not deleted by model-load recovery.
-
Clone this repository.
-
Install the latest Android Studio.
-
Open the Android project folder in Android Studio:
pocket_llm_src/ -
Build and install the app on your Android device.
-
Launch the app.
-
On first launch, choose a model from the built-in model picker.
-
Download the selected model directly inside the app.
-
Start chatting locally on device
local-document-intelligence A privacy-first offline document intelligence system with persistent local RAG, hybrid retrieval, and source-grounded answers.
Gemma 4 is provided by Google under the Apache License 2.0. Google's Gemma documentation also states that Gemma models are provided with open weights and support responsible commercial use.
- Gemma 4 license: https://ai.google.dev/gemma/apache_2
- Gemma 4 overview: https://ai.google.dev/gemma/docs/core
Qwen model files follow the upstream Qwen license terms.
Please review the original model license before redistribution or commercial use.




