Skip to content

Latest commit

 

History

History
382 lines (285 loc) · 18.5 KB

File metadata and controls

382 lines (285 loc) · 18.5 KB

jitLLM: LLM inference & serving for the JVM, on any GPU

build JDK21 Maven Central Java 21 Java 25 LangChain4j NVIDIA OpenCL Apple Docker DeepWiki License: MIT


Think vLLM — but pure Java, and it runs on any GPU.

jitLLM is a JVM-native LLM inference and serving engine. You write and ship plain Java; TornadoVM JIT-compiles the hot transformer kernels to CUDA, OpenCL, or Apple Metal at runtime — no JNI glue, no second toolchain, no native rebuild per GPU.

One .jar runs the same model on NVIDIA, Intel, AMD, and Apple Silicon, from a laptop to an RTX 5090.

Serve it behind an OpenAI-compatible API, embed it in LangChain4j or Quarkus, or run it from the CLI in one line.


Why jitLLM

  • 🟦 Pure Java, all the way down. Transformer kernels are written in Java and accelerated by TornadoVM — no CUDA C, no hand-written JNI. Debug and build with the toolchain you already have.
  • 🌍 Write once, run on any GPU. NVIDIA (CUDA), Intel & AMD (OpenCL), Apple Silicon (Metal). Backend is auto-detected from your TornadoVM SDK — switch with a flag, not a rebuild.
  • 🔌 Drop-in for the Java AI stack. Official LangChain4j provider (since v1.7.1) and Quarkus inference engine.
  • ⚡ Built to serve. OpenAI-compatible HTTP server, llama-bench-style benchmarking, and tensor-core (MMA) batch prefill (see Serving).
  • 📦 Many models, one runtime. Llama 3, Mistral, Qwen 2.5 / Qwen 3, Phi-3, IBM Granite 3.3 / 4.0, DeepSeek-R1-Distill — all in GGUF.

⏱️ Quickstart (60 seconds)

# 1. Install a TornadoVM SDK (bundles the GPU runtime)
curl -s "https://get.sdkman.io" | bash && source "$HOME/.sdkman/bin/sdkman-init.sh"
sdk install tornadovm
tornado --devices        # confirm your GPU is listed

# 2. Run a model on the GPU — no build required, via JBang
jbang jitllm@beehive-lab -m beehive-llama-3.2-1b-instruct-fp16.gguf -p "Explain GPU acceleration in one sentence."

Grab a ready-to-run model from the Hugging Face collections below.


🧩 Serving: OpenAI-compatible (preview)

jitLLM is growing into a serving engine — the vLLM-style path for the JVM:

  • 🌐 OpenAI-compatible server — jitllm serve exposes /v1/chat/completions and /v1/completions with streaming and zero external dependencies. /v1/models reports the served context length, so clients size their prompts instead of guessing. Point any OpenAI client at localhost.
  • 🎯 Tensor-core (MMA) batch prefill on the CUDA backend, FP16 & Q8_0 — --with-prefill-decode --batch-prefill-size N.
  • 📈 llama-bench-style benchmarking — jitllm --bench reports a pp/tg matrix with avg±stddev in md/csv/json/jsonl/sql. See Running the CLI for the flag.
  • 🧮 On-device greedy sampling (landing next) — argmax on the GPU keeps logits device-side, cutting device→host traffic by ~500× per token. (PR #134)
  • 📚 Static batched decode (landing next) — B independent sequences per step for up to 41× aggregate throughput (Llama & Qwen3). (PR #129)

LangChain4j LangChain4j & Quarkus

Since LangChain4j v1.7.1, jitllm is an officially supported model provider — no glue code, GPU-accelerated out of the box.

JitLLMChatModel model = JitLLMChatModel.builder()
        .modelPath(modelPath)
        .temperature(0.9)      // more creative
        .topP(0.9)             // more variety
        .maxTokens(2048)
        .onGPU(Boolean.TRUE)   // false → lightweight CPU llama3.java
        .build();

📖 LangChain4j docs · 🚀 Agentic workflow demo

📦 Maven

JDK 21 (jdk21 profile, auto-activates for JDK [21,25)):

<dependency>
    <groupId>io.github.beehive-lab</groupId>
    <artifactId>jitllm</artifactId>
    <version>1.0.2-jdk21</version>
</dependency>

JDK 22+ (jdk22plus profile, auto-activates for JDK [22,)):

<dependency>
    <groupId>io.github.beehive-lab</groupId>
    <artifactId>jitllm</artifactId>
    <version>1.0.2-jdk22plus</version>
</dependency>

📦 Gradle

JDK 21:

implementation 'io.github.beehive-lab:jitllm:1.0.2-jdk21'

JDK 22+:

implementation 'io.github.beehive-lab:jitllm:1.0.2-jdk22plus'

The artifact is a library jar: it carries jitllm's own classes and nothing else. TornadoVM and its dependencies come from the TornadoVM SDK at runtime, so start the JVM with @$TORNADOVM_HOME/tornado-argfile --add-modules jdk.incubator.vector. Code that uses TornadoVM types directly declares tornado-api itself, with provided scope.


[Interactive mode] — RTX 5090, with nvtop tracking GPU utilization and memory

Demo


🛠️ Install & build

Prerequisites

  • Java 21 or 22+ — required for the Vector API & TornadoVM. Each line has its own artifact (-jdk21 / -jdk22plus), matching TornadoVM's; both launchers work on either, while jitllm4j itself needs Java 25 to run.
  • TornadoVM with an OpenCL, CUDA, or Metal backend. jitllm/jitllm4j auto-detect whichever backend your installed SDK was built with.
  • GCC/G++ 13+ — to build TornadoVM's native components.

TornadoVM

jitllm compiles against, and runs on, the TornadoVM SDK that TORNADOVM_HOME points at. Any SDK from TornadoVM 7.0.0 on works: a released one, from the official website or SDKMAN! (sdk install tornadovm), or a local build of TornadoVM develop. Its JDK line must match the JDK you build with: a -jdk21 SDK for JDK 21, a -jdk22plus SDK for JDK 22 and newer. The build checks both that TORNADOVM_HOME is set and that the line matches.

Build from source

git clone https://github.com/beehive-lab/jitllm.git && cd jitllm

export TORNADOVM_HOME=/path/to/tornadovm-sdk     # e.g. ~/.sdkman/candidates/tornadovm/current
export PATH="$TORNADOVM_HOME/bin:$PATH"
mvn clean package -DskipTests                    # or ./mvnw; drop -DskipTests to run the unit tests

./jitllm --gpu --model model.gguf --prompt "..."

Nothing is resolved from a Maven repository for TornadoVM, so there is no version to pass: switch SDKs by changing TORNADOVM_HOME, and clean when you do.

task command
IDE support (IntelliJ IDEA, VS Code, Eclipse) enable the ide Maven profile; it adds the SDK's jars as dependencies. For an SDK other than the release in pom.xml (tornadovm.release.version), run scripts/ide-setup.sh first (and again after switching TORNADOVM_HOME): it records the SDK's jar version in .mvn/maven.config
Build against TornadoVM develop scripts/tornadovm-dev.sh setup --backend cuda --jdk 21, then eval "$(scripts/tornadovm-dev.sh env)" (exports TORNADOVM_HOME) and build as above
Advance to the latest develop scripts/tornadovm-dev.sh refresh --backend cuda --jdk 21
Reproduce an exact develop revision scripts/tornadovm-dev.sh setup --ref <40-hex commit> --backend cuda --jdk 21
What is prepared scripts/tornadovm-dev.sh status
Remove old develop installations scripts/tornadovm-dev.sh prune --yes (explicit; never automatic)
Build a jitllm release ./mvnw -P release -Dtornadovm.release.version=<X.Y.Z> clean package — a release depends on a published TornadoVM release from Maven Central, not on TORNADOVM_HOME

scripts/tornadovm-dev.sh is what CI uses to build TornadoVM develop. Installations live under ~/.jitllm/tornadovm/<backend>-jdk<N>/<commit>-r<recipe>/ and are immutable; current is only a convenience pointer to the last one prepared, and nothing is deleted unless you run scripts/tornadovm-dev.sh prune --yes (which removes every installation on that line except current; do not run it while a build or an inference process is using an older one).


▶️ Running the CLI

Use the jitllm script with --gpu. The backend (OpenCL, CUDA, or Metal) is auto-detected from your installed TornadoVM SDK (TORNADOVM_HOME/etc/tornado.backend) — no need to select it manually. If your SDK was built with more than one backend, force one with --opencl, --cuda (NVIDIA), or --metal (Apple Silicon); forcing a backend that isn't part of the installed SDK errors out.

# Basic GPU inference — backend auto-detected
./jitllm --gpu --verbose \
  --model beehive-llama-3.2-1b-instruct-fp16.gguf \
  --prompt "Explain the benefits of GPU acceleration."

# Force a specific backend (only needed for multi-backend SDKs)
./jitllm --gpu --cuda \
  --model beehive-llama-3.2-1b-instruct-fp16.gguf \
  --prompt "Explain the benefits of GPU acceleration."

Swap in any tested model — e.g. beehive-llama-3.2-3b-instruct-fp16.gguf or ...-8b-....

jitllm4j — zero-dependency Java 25 script

Same backend auto-detection as jitllm. A single-file Java 25 launcher that replaces the Python script (needs java 25+ on your PATH):

./jitllm4j --gpu --verbose-init --metal \
  --model Mistral-7B-Instruct-v0.3.Q8_0.gguf --prompt "what is java"

🚀 JBang — run without building

Script-like startup à la Jlama, powered by JBang:

curl -Ls https://sh.jbang.dev | bash -s - app setup

# From the catalog
jbang jitllm@beehive-lab -m model.gguf -p "Tell me a joke"
jbang app install jitllm@beehive-lab && jitllm -m model.gguf -p "Hello!"

Runs JitllmCli.java on the GPU through the TornadoVM SDK in TORNADOVM_HOME (JDK 22+). The pre-rename gpullama3@beehive-lab alias runs the same script.


Tested Models


💾 GPU memory

Default device allocation is 14GB. Larger models need more — raise it with --gpu-memory:

Model size Recommended Flag
1B 14GB (default) —
3–7B 15GB+ --gpu-memory 15GB
8B+ 20GB+ --gpu-memory 20GB
./jitllm --gpu --model beehive-llama-3.2-3b-instruct-fp16.gguf \
  --prompt "Tell me a joke" --gpu-memory 15GB

Still out of memory? Use Q4_0 instead of Q8_0, or close other GPU apps. The error to look for:

org.beehive.jitllm.api.InsufficientDeviceMemoryException: [GPUL-MEM-001] This configuration
needs about 13826.6 MiB of device memory but the configured budget is 1024.0 MiB
(short by 12802.6 MiB).
  Dominant component: weights (per-layer) at 13313.0 MiB
  Raise -Dtornado.device.memory, reduce the context length, or select a smaller quantization.

It is followed by a per-component device memory plan, so you can see what is actually consuming the budget before changing anything.


Run configuration at startup

With -v / --verbose, the CLI prints an aligned summary to stderr, after preparing the session and before printing generated text. It shows the model filename, parameter count and file size, GGUF quantization alongside loaded weight types, device, TornadoVM version on GPU runs, execution mode, context capacity, batched prefill chunk width, KV-cache precision, MMA/tensor-core selection, native libraries, CUDA graphs, staged transfers, GPU allocation budget, and sampling settings. GPU memory estimates split weights, KV cache, and workspace (including staging and control buffers), with a total. These predict allocation-budget charges, not physical VRAM use; conservative predictions are labeled. CPU runs instead show common-pool worker parallelism plus the caller and the configured tensor/Q4 Vector API settings. Tensor-core and native-library details describe the selected prefill path; native libraries choose their own kernel algorithms.

Startup timings separate model loading, plan construction, TornadoVM JIT precompilation, and initial device setup. Device setup includes uploads and execution (including CUDA graph capture when enabled), so it is not a pure transfer measurement. “Ready to generate” measures elapsed startup time through session preparation. Prefill and decode performance still appear after generation.

Without --verbose, startup diagnostics are hidden and session preparation stays lazy; errors, warnings, generated text, and final performance metrics remain visible. Verbose output includes SDK location, dimensions/head counts, training context, execution path, and memory-estimate assumptions. Startup timings are printed once; the ending performance block contains only request metrics.

--verbose-init remains a hidden deprecated alias for --verbose. For direct Java launches, use -Djitllm.verbose=true. The legacy -Djitllm.EnableTimingForTornadoVMInit=true setting still enables the report and, in addition, the older per-stage initialization log lines; --verbose no longer sets it. The full Java command is printed by --show-command, not by --verbose. The summary is CLI-only; library callers can explicitly use GenerationSession.prepare() to prepare a session and obtain its execution settings without advancing its position.


🔧 Embed in your own tools

--show-command prints the exact Java + JVM invocation used under the hood, so you can replicate it in IntelliJ, Maven, Gradle, or any launcher:

jitllm --gpu --model beehive-llama-3.2-1b-instruct-fp16.gguf \
  --prompt "tell me a joke" --show-command

Each command has focused help:

./jitllm --help
./jitllm run --help
./jitllm chat --help
./jitllm serve --help
./jitllm bench --help
Command Purpose Key options
run Generate one response, then exit --prompt, --system-prompt, --max-new-tokens
chat Terminal conversation using one persistent session --system-prompt, --max-new-tokens per turn
serve OpenAI-compatible HTTP API --host, --port; experimental --continuous-batching
bench Repeated prefill/decode workloads --pp, --tg, --depth, --repetitions, --output

./jitllm --help lists every command and option; ./jitllm COMMAND --help shows one command's. Options are grouped as Engine Configuration (model, prompt, sampling, context, prefill mode, -v/--verbose), the command's own group (Server, Benchmark), Hardware Configuration, Debug and Profiling, TornadoVM Execution Verbose, and Advanced Options, where experimental options are tagged [experimental].

./jitllm run -m model.gguf --gpu -c 4096 --max-new-tokens 128 --prompt "Explain SIMD."
./jitllm chat -m model.gguf --gpu -c 4096 --max-new-tokens 128
./jitllm serve -m model.gguf --gpu -c 4096 --host 127.0.0.1 --port 8080 -v
./jitllm bench -m model.gguf --gpu --pp 128,512 --tg 64 --depth 0,4096 --repetitions 3 --output json

Key/value cache precision. The KV cache is stored in FP16 by default (accumulation stays FP32). --fp32-kv-cache selects FP32, the compatibility and numerical-reference choice. A configuration whose kernels do not implement the FP16 cache is refused before the model's cache or plan is built, naming the combination and pointing at --fp32-kv-cache; it never falls back to FP32 silently. The supported set is tracked in docs/architecture/kv-cache-support.md. The old --fp16-kv-cache flag was removed and is refused with this migration.

Existing flag-based invocations remain supported: default/--instruct → run, --interactive/--chat/-i → chat, --server → serve, and --bench → bench. Conflicting modes and options for another command are rejected. --max-tokens/-n remains a deprecated context-capacity alias; it has not been repurposed as an output limit. HTTP max_tokens keeps its existing generated-token meaning. Legacy --bench-args="..." still accepts benchmark arguments, including quoted values.

Serving binds to loopback by default and prints its actual address and port when ready. Use --host 0.0.0.0 explicitly for all IPv4 interfaces. Clients submit conversation history in each HTTP request; chat retains terminal conversation history locally.

# Peek at what TornadoVM is doing
./jitllm --gpu --model model.gguf --prompt "..." --print-kernel      # generated GPU kernel
./jitllm --gpu --model model.gguf --prompt "..." --print-bytecodes   # TornadoVM bytecodes
./jitllm --gpu --model model.gguf --prompt "..." --debug --full-dump # everything

🙏 Acknowledgments

Partially funded by EU Horizon Europe & UKRI grants (most recent first): AERO 101092850 · P2CODE 101093069 · ENCRYPT 101070670 · TANGO 101070052.

License

MIT