Local LLM Stack (llama.cpp)

Status: Active Tags: Deployment Infrastructure Inference Service Created: 2026-08-30 Updated: 2026-09-05 Related: honcho-local-deployment, cloudflare-pages-deploy, hermes-cron-operations, realtime-voice-assistant

Definition

The local inference stack on this host: llama.cpp llama-server running in docker-compose, serving the models that power Honcho and the Hermes agent’s local provider. As of 2026-09-04 there are three serving endpoints on the box: the main llama.cpp stack (:8000), a vLLM test server (:8010), and a dedicated small-LLM container for the voice assistant (:8020).

How It Works

  • docker-compose setup with a llama-server service and an embedding service (nomic-embed-text).
  • Dual AMD GPUs, addressed with HIP_VISIBLE_DEVICES=0,1 (ROCm/HIP, not CUDA).
  • Host port 8000 — compose maps host 8000 → container 8001.
  • Models live at /home/romain/models/.
  • vLLM test server (added 2026-09-03): vllm-06b.service — a user-level systemd unit (enabled; active since 2026-09-03 20:44; a staging copy of the unit sits in /home/romain/workspace/). Runs the rocm/vllm-dev:nightly docker container vllm-06b serving Qwen3-0.6B (a test model, stored separately at /home/romain/models-vllm/Qwen3-0.6B) as OpenAI-compatible qwen3-0.6b on :8010, --max-model-len 8192. Key flags: HIP_VISIBLE_DEVICES=0 (GPU 0 only), HSA_OVERRIDE_GFX_VERSION=11.0.0, --gpu-memory-utilization 0.27 (deliberately small so it coexists with llama.cpp on the same GPUs), --enforce-eager, --attention-backend TRITON_ATTN, --no-enable-prefix-caching, --dtype bfloat16, --shm-size=16g, --network=host. ExecStartPre polls the docker daemon (up to 60 s) and force-removes any stale vllm-06b container first. An older system-level vllm-server.service (2026-08-16) also exists in /etc/systemd/system/.
  • Voice LLM container (added 2026-09-04): llama-server-voice4b — a separate llama.cpp serving container on :8020 (GPU 1 only, unless-stopped), serving Qwen3-4B-Instruct-2507-Q4_K_M.gguf (~2.5 GB, at /home/romain/workspace/voice-rt/) with 4 slots. It is the dedicated LLM for the real-time voice assistant (see realtime-voice-assistant). It is a small model deliberately kept off the main :8000 stack so the voice path has low, predictable latency and never competes with the 27 B reasoning server.

Key Parameters

ItemValue
Main endpointhttp://localhost:8000/v1 (OpenAI-compatible)
Main modelsQwen3.6-35B, Qwen3.8-27B (Qwen3.8-27B-UD-Q6_K_XL.gguf)
Embedding modelnomic-embed-text-v1.5 (768 dims)
GPU envHIP_VISIBLE_DEVICES=0,1
Context200K (llama-server)
vLLM test endpointhttp://localhost:8010/v1 (qwen3-0.6b, 8192 ctx, GPU 0, 27 % mem)
Voice LLM endpointhttp://localhost:8020/v1 (Qwen3-4B Q4_K_M, GPU 1, 4 slots — serves realtime-voice-assistant)

When To Use

  • Any consumer needing local inference: Honcho (LLM + embeddings), Hermes agent local provider, and the voice assistant (:8020).
  • Qwen3.8-27B is a reasoning model (~2.5 tok/s effective): thinking traces consume the output budget — disable thinking for JSON-returning callers (see honcho-local-deployment).

Risks & Pitfalls

  • The llama.cpp image has no python3 — container healthchecks must use bash (e.g. bash -c '</dev/tcp/...'), not Python.
  • The embedding service needs host ROCm libs mounted into the container.
  • Qwen thinking mode left on → empty content / JSON repair failures for structured callers; raise max_output_tokens (16384 safe at 200K ctx) AND disable thinking.
  • Stale-token floor: the qwen3 model slug carried a 180 s stale check that killed long agent runs mid-task. Fixed 2026-08-30 with providers.custom.stale_timeout_seconds=900 in the Hermes config — the CLI splits dotted model names, so the key applies provider-wide (every model on providers.custom), not just one model. See hermes-cron-operations.
  • vLLM on ROCm: the nightly image needs HSA_OVERRIDE_GFX_VERSION=11.0.0, --enforce-eager and the TRITON_ATTN backend to run reliably on this host’s AMD GPUs; keep --gpu-memory-utilization low (0.27) so it shares a GPU with the llama.cpp stack. Models for vLLM live in a separate dir (/home/romain/models-vllm/), not /home/romain/models/.
  • Three coexisting GPU servers: the main stack (:8000, GPU 0+1), vLLM (:8010, GPU 0, 27 % mem) and the voice4b container (:8020, GPU 1) all share the dual GPUs — keep each one’s memory footprint small (low --gpu-memory-utilization / small model) so none starves the others.

Sources

  • /home/romain/workspace docker-compose for llama-server (host 8000 → 8001)
  • /media/data/honcho/.env (deployment comments, 2026-08-29 deriver fix)
  • /home/romain/workspace/voice-rt/RUNBOOK.md (voice4b container on :8020, 2026-09-04)