Archived (2026-08-23). The runtime library continues as erllama (hex package
erllama, docs at hexdocs.pm/erllama): Erlang-native llama.cpp inference with supervised model processes, the byte-exact tiered KV cache, streaming, chat and tool calls, all as a library you embed in your own application. The HTTP daemon and CLI in this repository are frozen at this commit (runtime at its vendored state, llama.cpp b10068) and are not developed further. If you need an HTTP inference server rather than a library, use Ollama.
OTP-native LLM inference for the BEAM: dirty NIFs over llama.cpp, supervised
per-model processes, and a byte-exact tiered KV cache, with an
OpenAI/Anthropic/Ollama-compatible HTTP daemon on top.
Inference as a first-class OTP citizen, not a Python sidecar. The wedge is supervision, per-model queues, and cancel-on-disconnect, with the cache more warm state than fits in RAM.
A rebar3 umbrella; each app is a separately publishable Hex package and the repo is versioned as a whole.
| App | What it is |
|---|---|
apps/barrel_inference |
The runtime: dirty NIFs over llama.cpp, supervised model processes, byte-exact tiered KV cache. |
apps/barrel_inference_server |
The API daemon: OpenAI-, Anthropic-, and Ollama-compatible HTTP, model registry, per-model queues, keep-alive, metrics. |
apps/barrel_inference_cli |
The barrel-inference CLI: serve boots the daemon; pull/run/ps/rm drive a running one over HTTP. |
A distributed control plane (barrel_inference_cluster: routing, cache-aware
placement, node discovery) is a planned follow-up.
rebar3 compile # builds the NIF (vendored llama.cpp via cmake)
rebar3 as prod release # the barrel_inference_server daemon release
rebar3 escriptize # the barrel-inference CLI
Requires Erlang/OTP 28 and rebar3 3.25+, plus cmake and a C/C++ toolchain for the NIF. See each app's README for the public API and configuration.
barrel-inference serve # start the API server
barrel-inference pull <model> # fetch a model
barrel-inference run <model> "hello" # one-shot completion
barrel-inference ps # list loaded models
Or with Docker:
docker compose up
One site covers the whole project, for both operators and contributors:
https://barrel-platform.github.io/barrel_inference/. It is a mkdocs build at
the repo root (mkdocs.yml, docs/) that surfaces each app's guides in place.
Function-level API reference ships per package on hexdocs: barrel_inference and barrel_inference_server.
Build the site locally:
pip install -r docs-requirements.txt
mkdocs serve
MIT. Part of the barrel-platform project.