Local LLM Stack (llama.cpp)
Status: Active Tags: Deployment Infrastructure Inference Service Created: 2026-08-30 Updated: 2026-09-05 Related: honcho-local-deployment, cloudflare-pages-deploy, hermes-cron-operations, realtime-voice-assistant
Definition
The local inference stack on this host: llama.cpp llama-server running in docker-compose, serving the models that power Honcho and the Hermes agent’s local provider. As of 2026-09-04 there are three serving endpoints on the box: the main llama.cpp stack (:8000), a vLLM test server (:8010), and a dedicated small-LLM container for the voice assistant (:8020).
How It Works
- docker-compose setup with a
llama-serverservice and anembeddingservice (nomic-embed-text). - Dual AMD GPUs, addressed with
HIP_VISIBLE_DEVICES=0,1(ROCm/HIP, not CUDA). - Host port 8000 — compose maps host 8000 → container 8001.
- Models live at
/home/romain/models/. - vLLM test server (added 2026-09-03):
vllm-06b.service— a user-level systemd unit (enabled; active since 2026-09-03 20:44; a staging copy of the unit sits in/home/romain/workspace/). Runs therocm/vllm-dev:nightlydocker containervllm-06bservingQwen3-0.6B(a test model, stored separately at/home/romain/models-vllm/Qwen3-0.6B) as OpenAI-compatibleqwen3-0.6bon:8010,--max-model-len 8192. Key flags:HIP_VISIBLE_DEVICES=0(GPU 0 only),HSA_OVERRIDE_GFX_VERSION=11.0.0,--gpu-memory-utilization 0.27(deliberately small so it coexists with llama.cpp on the same GPUs),--enforce-eager,--attention-backend TRITON_ATTN,--no-enable-prefix-caching,--dtype bfloat16,--shm-size=16g,--network=host.ExecStartPrepolls the docker daemon (up to 60 s) and force-removes any stalevllm-06bcontainer first. An older system-levelvllm-server.service(2026-08-16) also exists in/etc/systemd/system/. - Voice LLM container (added 2026-09-04):
llama-server-voice4b— a separate llama.cpp serving container on:8020(GPU 1 only,unless-stopped), servingQwen3-4B-Instruct-2507-Q4_K_M.gguf(~2.5 GB, at/home/romain/workspace/voice-rt/) with 4 slots. It is the dedicated LLM for the real-time voice assistant (see realtime-voice-assistant). It is a small model deliberately kept off the main:8000stack so the voice path has low, predictable latency and never competes with the 27 B reasoning server.
Key Parameters
| Item | Value |
|---|---|
| Main endpoint | http://localhost:8000/v1 (OpenAI-compatible) |
| Main models | Qwen3.6-35B, Qwen3.8-27B (Qwen3.8-27B-UD-Q6_K_XL.gguf) |
| Embedding model | nomic-embed-text-v1.5 (768 dims) |
| GPU env | HIP_VISIBLE_DEVICES=0,1 |
| Context | 200K (llama-server) |
| vLLM test endpoint | http://localhost:8010/v1 (qwen3-0.6b, 8192 ctx, GPU 0, 27 % mem) |
| Voice LLM endpoint | http://localhost:8020/v1 (Qwen3-4B Q4_K_M, GPU 1, 4 slots — serves realtime-voice-assistant) |
When To Use
- Any consumer needing local inference: Honcho (LLM + embeddings), Hermes agent local provider, and the voice assistant (
:8020). - Qwen3.8-27B is a reasoning model (~2.5 tok/s effective): thinking traces consume the output budget — disable thinking for JSON-returning callers (see honcho-local-deployment).
Risks & Pitfalls
- The llama.cpp image has no python3 — container healthchecks must use bash (e.g.
bash -c '</dev/tcp/...'), not Python. - The embedding service needs host ROCm libs mounted into the container.
- Qwen thinking mode left on → empty content / JSON repair failures for structured callers; raise
max_output_tokens(16384 safe at 200K ctx) AND disable thinking. - Stale-token floor: the
qwen3model slug carried a 180 s stale check that killed long agent runs mid-task. Fixed 2026-08-30 withproviders.custom.stale_timeout_seconds=900in the Hermes config — the CLI splits dotted model names, so the key applies provider-wide (every model onproviders.custom), not just one model. See hermes-cron-operations. - vLLM on ROCm: the nightly image needs
HSA_OVERRIDE_GFX_VERSION=11.0.0,--enforce-eagerand theTRITON_ATTNbackend to run reliably on this host’s AMD GPUs; keep--gpu-memory-utilizationlow (0.27) so it shares a GPU with the llama.cpp stack. Models for vLLM live in a separate dir (/home/romain/models-vllm/), not/home/romain/models/. - Three coexisting GPU servers: the main stack (
:8000, GPU 0+1), vLLM (:8010, GPU 0, 27 % mem) and the voice4b container (:8020, GPU 1) all share the dual GPUs — keep each one’s memory footprint small (low--gpu-memory-utilization/ small model) so none starves the others.
Related Concepts
- honcho-local-deployment — main consumer of the
:8000stack - realtime-voice-assistant — the voice assistant that uses the dedicated
:8020voice4b LLM - cloudflare-pages-deploy — unrelated deploy tooling on the same host
- hermes-cron-operations — timeout limits (stale floor, idle) that interact with this provider
Sources
/home/romain/workspacedocker-compose for llama-server (host 8000 → 8001)/media/data/honcho/.env(deployment comments, 2026-08-29 deriver fix)/home/romain/workspace/voice-rt/RUNBOOK.md(voice4b container on :8020, 2026-09-04)