|
jitLLM is a JVM-native LLM inference and serving engine. You write and ship plain Java; TornadoVM JIT-compiles the hot transformer kernels to CUDA, OpenCL, or Apple Metal at runtime — no JNI glue, no second toolchain, no native rebuild per GPU. One Serve it behind an OpenAI-compatible API, embed it in LangChain4j or Quarkus, or run it from the CLI in one line. |
- 🟦 Pure Java, all the way down. Transformer kernels are written in Java and accelerated by TornadoVM — no CUDA C, no hand-written JNI. Debug and build with the toolchain you already have.
- 🌍 Write once, run on any GPU. NVIDIA (CUDA), Intel & AMD (OpenCL), Apple Silicon (Metal). Backend is auto-detected from your TornadoVM SDK — switch with a flag, not a rebuild.
- 🔌 Drop-in for the Java AI stack. Official LangChain4j provider (since v1.7.1) and Quarkus inference engine.
- ⚡ Built to serve. OpenAI-compatible HTTP server, llama-bench-style benchmarking, and tensor-core (MMA) batch prefill (see Serving).
- 📦 Many models, one runtime. Llama 3, Mistral, Qwen 2.5 / Qwen 3, Phi-3, IBM Granite 3.3 / 4.0, DeepSeek-R1-Distill — all in GGUF.
# 1. Install a TornadoVM SDK (bundles the GPU runtime)
curl -s "https://get.sdkman.io" | bash && source "$HOME/.sdkman/bin/sdkman-init.sh"
sdk install tornadovm
tornado --devices # confirm your GPU is listed
# 2. Run a model on the GPU — no build required, via JBang
jbang jitllm@beehive-lab -m beehive-llama-3.2-1b-instruct-fp16.gguf -p "Explain GPU acceleration in one sentence."Grab a ready-to-run model from the Hugging Face collections below.
jitLLM is growing into a serving engine — the vLLM-style path for the JVM:
- 🌐 OpenAI-compatible server —
jitllm serveexposes/v1/chat/completionsand/v1/completionswith streaming and zero external dependencies./v1/modelsreports the served context length, so clients size their prompts instead of guessing. Point any OpenAI client atlocalhost. - 🎯 Tensor-core (MMA) batch prefill on the CUDA backend, FP16 & Q8_0 —
--with-prefill-decode --batch-prefill-size N. - 📈 llama-bench-style benchmarking —
jitllm --benchreports a pp/tg matrix with avg±stddev in md/csv/json/jsonl/sql. See Running the CLI for the flag. - 🧮 On-device greedy sampling (landing next) — argmax on the GPU keeps logits device-side, cutting device→host traffic by ~500× per token. (PR #134)
- 📚 Static batched decode (landing next) — B independent sequences per step for up to 41× aggregate throughput (Llama & Qwen3). (PR #129)
Since LangChain4j v1.7.1, jitllm is an officially supported model provider — no glue code, GPU-accelerated out of the box.
JitLLMChatModel model = JitLLMChatModel.builder()
.modelPath(modelPath)
.temperature(0.9) // more creative
.topP(0.9) // more variety
.maxTokens(2048)
.onGPU(Boolean.TRUE) // false → lightweight CPU llama3.java
.build();📖 LangChain4j docs · 🚀 Agentic workflow demo
JDK 21 (jdk21 profile, auto-activates for JDK [21,25)):
<dependency>
<groupId>io.github.beehive-lab</groupId>
<artifactId>jitllm</artifactId>
<version>1.0.2-jdk21</version>
</dependency>JDK 22+ (jdk22plus profile, auto-activates for JDK [22,)):
<dependency>
<groupId>io.github.beehive-lab</groupId>
<artifactId>jitllm</artifactId>
<version>1.0.2-jdk22plus</version>
</dependency>JDK 21:
implementation 'io.github.beehive-lab:jitllm:1.0.2-jdk21'JDK 22+:
implementation 'io.github.beehive-lab:jitllm:1.0.2-jdk22plus'The artifact is a library jar: it carries jitllm's own classes and nothing else. TornadoVM and
its dependencies come from the TornadoVM SDK at runtime, so start the JVM with
@$TORNADOVM_HOME/tornado-argfile --add-modules jdk.incubator.vector. Code that uses TornadoVM
types directly declares tornado-api itself, with provided scope.
- Java 21 or 22+ — required for the Vector API & TornadoVM. Each line has its own artifact (
-jdk21/-jdk22plus), matching TornadoVM's; both launchers work on either, whilejitllm4jitself needs Java 25 to run. - TornadoVM with an OpenCL, CUDA, or Metal backend.
jitllm/jitllm4jauto-detect whichever backend your installed SDK was built with. - GCC/G++ 13+ — to build TornadoVM's native components.
jitllm compiles against, and runs on, the TornadoVM SDK that TORNADOVM_HOME points at. Any SDK
from TornadoVM 7.0.0 on works: a released one, from the
official website or SDKMAN!
(sdk install tornadovm), or a local build of TornadoVM develop. Its JDK line must match the JDK
you build with: a -jdk21 SDK for JDK 21, a -jdk22plus SDK for JDK 22 and newer. The build
checks both that TORNADOVM_HOME is set and that the line matches.
git clone https://github.com/beehive-lab/jitllm.git && cd jitllm
export TORNADOVM_HOME=/path/to/tornadovm-sdk # e.g. ~/.sdkman/candidates/tornadovm/current
export PATH="$TORNADOVM_HOME/bin:$PATH"
mvn clean package -DskipTests # or ./mvnw; drop -DskipTests to run the unit tests
./jitllm --gpu --model model.gguf --prompt "..."Nothing is resolved from a Maven repository for TornadoVM, so there is no version to pass: switch
SDKs by changing TORNADOVM_HOME, and clean when you do.
| task | command |
|---|---|
| IDE support (IntelliJ IDEA, VS Code, Eclipse) | enable the ide Maven profile; it adds the SDK's jars as dependencies. For an SDK other than the release in pom.xml (tornadovm.release.version), run scripts/ide-setup.sh first (and again after switching TORNADOVM_HOME): it records the SDK's jar version in .mvn/maven.config |
Build against TornadoVM develop |
scripts/tornadovm-dev.sh setup --backend cuda --jdk 21, then eval "$(scripts/tornadovm-dev.sh env)" (exports TORNADOVM_HOME) and build as above |
Advance to the latest develop |
scripts/tornadovm-dev.sh refresh --backend cuda --jdk 21 |
Reproduce an exact develop revision |
scripts/tornadovm-dev.sh setup --ref <40-hex commit> --backend cuda --jdk 21 |
| What is prepared | scripts/tornadovm-dev.sh status |
Remove old develop installations |
scripts/tornadovm-dev.sh prune --yes (explicit; never automatic) |
| Build a jitllm release | ./mvnw -P release -Dtornadovm.release.version=<X.Y.Z> clean package — a release depends on a published TornadoVM release from Maven Central, not on TORNADOVM_HOME |
scripts/tornadovm-dev.sh is what CI uses to build TornadoVM develop. Installations live under
~/.jitllm/tornadovm/<backend>-jdk<N>/<commit>-r<recipe>/ and are immutable; current is only a
convenience pointer to the last one prepared, and nothing is deleted unless you run
scripts/tornadovm-dev.sh prune --yes (which removes every installation on that line except
current; do not run it while a build or an inference process is using an older one).
Use the jitllm script with --gpu. The backend (OpenCL, CUDA, or Metal) is auto-detected from
your installed TornadoVM SDK (TORNADOVM_HOME/etc/tornado.backend) — no need to select it manually. If your
SDK was built with more than one backend, force one with --opencl, --cuda (NVIDIA), or --metal
(Apple Silicon); forcing a backend that isn't part of the installed SDK errors out.
# Basic GPU inference — backend auto-detected
./jitllm --gpu --verbose \
--model beehive-llama-3.2-1b-instruct-fp16.gguf \
--prompt "Explain the benefits of GPU acceleration."
# Force a specific backend (only needed for multi-backend SDKs)
./jitllm --gpu --cuda \
--model beehive-llama-3.2-1b-instruct-fp16.gguf \
--prompt "Explain the benefits of GPU acceleration."Swap in any tested model — e.g. beehive-llama-3.2-3b-instruct-fp16.gguf or ...-8b-....
Same backend auto-detection as jitllm. A single-file Java 25 launcher that replaces the Python
script (needs java 25+ on your PATH):
./jitllm4j --gpu --verbose-init --metal \
--model Mistral-7B-Instruct-v0.3.Q8_0.gguf --prompt "what is java"Script-like startup à la Jlama, powered by JBang:
curl -Ls https://sh.jbang.dev | bash -s - app setup
# From the catalog
jbang jitllm@beehive-lab -m model.gguf -p "Tell me a joke"
jbang app install jitllm@beehive-lab && jitllm -m model.gguf -p "Hello!"Runs
JitllmCli.javaon the GPU through the TornadoVM SDK inTORNADOVM_HOME(JDK 22+). The pre-renamegpullama3@beehive-labalias runs the same script.
Default device allocation is 14GB. Larger models need more — raise it with --gpu-memory:
| Model size | Recommended | Flag |
|---|---|---|
| 1B | 14GB (default) | — |
| 3–7B | 15GB+ | --gpu-memory 15GB |
| 8B+ | 20GB+ | --gpu-memory 20GB |
./jitllm --gpu --model beehive-llama-3.2-3b-instruct-fp16.gguf \
--prompt "Tell me a joke" --gpu-memory 15GBStill out of memory? Use Q4_0 instead of Q8_0, or close other GPU apps. The error to look for:
org.beehive.jitllm.api.InsufficientDeviceMemoryException: [GPUL-MEM-001] This configuration
needs about 13826.6 MiB of device memory but the configured budget is 1024.0 MiB
(short by 12802.6 MiB).
Dominant component: weights (per-layer) at 13313.0 MiB
Raise -Dtornado.device.memory, reduce the context length, or select a smaller quantization.
It is followed by a per-component device memory plan, so you can see what is actually consuming the budget before changing anything.
With -v / --verbose, the CLI prints an aligned summary to stderr, after preparing the session and before
printing generated text. It shows the model filename, parameter count and file size,
GGUF quantization alongside loaded weight types, device, TornadoVM version on GPU runs, execution mode, context
capacity, batched prefill chunk width, KV-cache precision, MMA/tensor-core selection,
native libraries, CUDA graphs, staged transfers, GPU allocation budget, and sampling settings.
GPU memory estimates split weights, KV cache, and workspace (including staging and
control buffers), with a total. These predict allocation-budget charges, not physical
VRAM use; conservative predictions are labeled. CPU runs instead show common-pool worker
parallelism plus the caller and the configured tensor/Q4 Vector API settings.
Tensor-core and native-library details describe the selected prefill path; native
libraries choose their own kernel algorithms.
Startup timings separate model loading, plan construction, TornadoVM JIT precompilation, and initial device setup. Device setup includes uploads and execution (including CUDA graph capture when enabled), so it is not a pure transfer measurement. “Ready to generate” measures elapsed startup time through session preparation. Prefill and decode performance still appear after generation.
Without --verbose, startup diagnostics are hidden and session preparation stays lazy;
errors, warnings, generated text, and final performance metrics remain visible.
Verbose output includes SDK location, dimensions/head counts, training context,
execution path, and memory-estimate assumptions. Startup timings are printed once;
the ending performance block contains only request metrics.
--verbose-init remains a hidden deprecated alias for --verbose.
For direct Java launches, use -Djitllm.verbose=true. The legacy
-Djitllm.EnableTimingForTornadoVMInit=true setting still enables the report and, in
addition, the older per-stage initialization log lines; --verbose no longer sets it.
The full Java command is printed by --show-command, not by --verbose.
The summary is CLI-only; library callers can explicitly use GenerationSession.prepare()
to prepare a session and obtain its execution settings without advancing its position.
--show-command prints the exact Java + JVM invocation used under the hood, so you can replicate it in IntelliJ, Maven, Gradle, or any launcher:
jitllm --gpu --model beehive-llama-3.2-1b-instruct-fp16.gguf \
--prompt "tell me a joke" --show-commandEach command has focused help:
./jitllm --help
./jitllm run --help
./jitllm chat --help
./jitllm serve --help
./jitllm bench --help| Command | Purpose | Key options |
|---|---|---|
run |
Generate one response, then exit | --prompt, --system-prompt, --max-new-tokens |
chat |
Terminal conversation using one persistent session | --system-prompt, --max-new-tokens per turn |
serve |
OpenAI-compatible HTTP API | --host, --port; experimental --continuous-batching |
bench |
Repeated prefill/decode workloads | --pp, --tg, --depth, --repetitions, --output |
./jitllm --help lists every command and option; ./jitllm COMMAND --help shows one command's.
Options are grouped as Engine Configuration (model, prompt, sampling, context, prefill mode,
-v/--verbose), the command's own group (Server, Benchmark), Hardware Configuration, Debug
and Profiling, TornadoVM Execution Verbose, and Advanced Options, where experimental options
are tagged [experimental].
./jitllm run -m model.gguf --gpu -c 4096 --max-new-tokens 128 --prompt "Explain SIMD."
./jitllm chat -m model.gguf --gpu -c 4096 --max-new-tokens 128
./jitllm serve -m model.gguf --gpu -c 4096 --host 127.0.0.1 --port 8080 -v
./jitllm bench -m model.gguf --gpu --pp 128,512 --tg 64 --depth 0,4096 --repetitions 3 --output jsonKey/value cache precision. The KV cache is stored in FP16 by default (accumulation
stays FP32). --fp32-kv-cache selects FP32, the compatibility and numerical-reference
choice. A configuration whose kernels do not implement the FP16 cache is refused before
the model's cache or plan is built, naming the combination and pointing at
--fp32-kv-cache; it never falls back to FP32 silently. The supported set is tracked in
docs/architecture/kv-cache-support.md. The old
--fp16-kv-cache flag was removed and is refused with this migration.
Existing flag-based invocations remain supported: default/--instruct → run,
--interactive/--chat/-i → chat, --server → serve, and --bench → bench.
Conflicting modes and options for another command are rejected. --max-tokens/-n
remains a deprecated context-capacity alias; it has not been repurposed as an output
limit. HTTP max_tokens keeps its existing generated-token meaning. Legacy
--bench-args="..." still accepts benchmark arguments, including quoted values.
Serving binds to loopback by default and prints its actual address and port when ready.
Use --host 0.0.0.0 explicitly for all IPv4 interfaces. Clients submit conversation
history in each HTTP request; chat retains terminal conversation history locally.
# Peek at what TornadoVM is doing
./jitllm --gpu --model model.gguf --prompt "..." --print-kernel # generated GPU kernel
./jitllm --gpu --model model.gguf --prompt "..." --print-bytecodes # TornadoVM bytecodes
./jitllm --gpu --model model.gguf --prompt "..." --debug --full-dump # everythingPartially funded by EU Horizon Europe & UKRI grants (most recent first): AERO 101092850 · P2CODE 101093069 · ENCRYPT 101070670 · TANGO 101070052.

