Tags

Tags give the ability to mark specific points in history as being important
  • v9.0.0

    v9.0.0
    
    **Updates Audiolla CPU and CUDA to v2.0.0 with configurable staged-file cleanup.**
    
    - Audiolla staged uploads and outputs now expire after 24 hours by file modification time, including files retained before upgrading. Set `AUDIOLLA_FILES_TTL=0` in `.env` before upgrading to preserve indefinite retention. Model caches are excluded.
    
    - Pins both Audiolla variants to v2.0.0, which uses shared Torchbase images. Exposes `AUDIOLLA_FILES_TTL` to both services and documents expiry, active-operation protection, and the upgrade path.
    - Adds `make restart-audiolla` to recreate only enabled Audiolla variants from locally available pinned images, without pulling, rebuilding, or restarting the rest of the stack.
    - Updates the agent setup reference and Codex plugin version for the new storage default.
  • v8.0.0

    v8.0.0
    
    **Fixes local GPU deployment and browser startup, and removes the GPU Dolphin Phi alias.**
    
    - Removes `local-ollama-cuda-dolphin-phi` from the model catalog, CUDA pulls, and fallback chains. Switch callers to `local-ollama-cpu-dolphin-phi`, `local-ollama-cuda-qwen3-abliterated-16b`, or `local-ollama-cuda-gemma4-abliterated-e4b`. Existing downloaded model files remain untouched.
    
    - Pins the CPU, CUDA, and pull-sidecar Ollama images to v0.34.4 with a digest so they can download the configured Gemma 4 models. The CUDA server keeps its `q8_0` KV cache; `OLLAMA_CUDA_KV_CACHE` exposes the setting in `.env.example`.
    - CUDA Ollama eviction checks HTTP status codes and reports failed unloads instead of logging success. Tests cover the unload request, downstream failure, and shared CUDA lock.
    - sd.cpp CUDA builds target `all-major` by default instead of relying on GPU detection during compilation. `SDCPP_CUDA_ARCHITECTURES` accepts narrower build targets when needed.
    - Browser replicas default to 1 GiB RAM and 2 GiB total RAM plus swap. The previous 256 MiB limit could stall browser startup and leave the proxy returning 503. `SAB_MEM_LIMIT` and `SAB_MEMSWAP_LIMIT` are documented in `.env.example`.
    
    - Updates the model catalog, fallback configuration, smoke tests, and deployment docs for CPU-only Dolphin Phi. Generated-config tests verify that the CPU alias remains available and the GPU alias is absent.
  • v7.0.0

    v7.0.0
    
    **Enables CLM by default and coordinates Decidealot's local inference with the shared hardware lock.**
    
    - Starting either Decidealot variant with `make run-bg` now enables its CUDA Qwen3-8B encoder and generated embeddings route. Existing CPU-only deployments must set `DECIDEALOT_CLM_ENABLED=false` before upgrading to avoid that NVIDIA GPU requirement.
    
    - Decidealot enables Laya, Von, and CLM by default. CLM uses the internal LiteLLM Qwen3-8B embeddings route. Hosted Jev still enables when its TypeSafe key is configured.
    - Aigate builds thin CPU and CUDA launcher images from Decidealot v0.6.0. The launchers require Redis admission for Laya and Von; unavailable admission fails the request rather than running unlocked. The upstream Decidealot images and embeddings API remain unchanged.
    
    - REST, MCP and batch items take the same shared hardware lock per Laya/Von inference. CUDA admission waits for the configured local encoder to unload before allocating the model. Failed encoder eviction returns an unavailable-provider response without loading Laya or Von.
    - CLM leaves hardware-lock ownership to its nested LiteLLM embeddings request, preventing a recursive MCP lock. When CUDA Decidealot and its internal encoder share the GPU, Aigate serializes CLM against local provider swaps even when `DECIDEALOT_CLM_PARALLEL_WITH_LOCAL_MODELS=true`. Remote encoder URLs do not use that shared-encoder lane.
    
    - `make test-config` checks provider activation for default, explicit, disabled, and standalone CLM encoder configurations.
    - `make build-decidealot` builds both thin launcher images without reinstalling Torch or CUDA. `make test-decidealot-coordination` checks admission, eviction ordering, cancellation and failure handling with a throwaway Redis and simulated inference.
    - `AIGATE_DECIDEALOT_ENCODER_URL` selects the internal encoder unload API. `AIGATE_DECIDEALOT_UNLOAD_TIMEOUT_SECONDS` sets its timeout, default 30 seconds and maximum 300 seconds.
  • v6.3.0

    v6.3.0
    
    **Updates Decidealot to v0.6.0 for hosted Jev decisions and independent batch requests.**
    
    - Set `DECIDEALOT_TYPESAFE_API_KEY` in the gitignored `.env` to include the account's live Jev selectors in Decidealot's model catalog. The key is passed to both CPU and CUDA variants. Jev requests leave the stack for TypeSafe's API; local Laya, Von, and CLM remain available without that key.
    - `POST /decidealot/v1/systemone/batch` and its CUDA counterpart accept a list of complete decision requests. The `system_one_batch` MCP tool joins both the direct servers and LiteLLM's aggregated `/mcp/`, bringing each variant to four tools.
    - `DECIDEALOT_MAX_BATCH_REQUESTS`, `DECIDEALOT_MAX_BATCH_CONCURRENCY`, `DECIDEALOT_MAX_RESIDENT_LOCAL_PROVIDERS`, and `DECIDEALOT_CLM_PARALLEL_WITH_LOCAL_MODELS` expose Decidealot's batch and residency controls. Existing single-request behavior and the one-local-provider default remain unchanged.
    
    - Both Decidealot images now use v0.6.0. The service docs, MCP descriptions, agent guidance, and smoke-test tool list match the new API.
  • v6.2.0

    v6.2.0
    
    Adds the optional Decidealot CLM provider to the AIGate stack.
    
    - `DECIDEALOT_CLM_ENABLED=true` enables Decidealot CLM. With `DECIDEALOT=1` or `DECIDEALOT_CUDA=1` and `LLAMACPP_CUDA=1`, AIGate starts the CUDA llama.cpp profile and routes CLM embeddings through `local-llamacpp-cuda-qwen3-8b` on its internal LiteLLM network.
    - The Decidealot CPU and CUDA services use `psyb0t/decidealot:v0.5.0`, which includes Laya, Von, and the optional CLM provider. Laya and Von remain enabled by default. CLM stays opt-in.
    
    - The source-controlled Codex plugin metadata now reports the current AIGate release version.
  • v6.1.0

    v6.1.0
    
    Adds a CUDA Qwen3 8B embeddings endpoint for Contrastive-LM CLM and makes llama.cpp model downloads artifact-pinned.
    
    - `local-llamacpp-cuda-qwen3-8b` serves Qwen3 8B Q8_0 at `/v1/embeddings` with last-token pooling, L2 normalization, and 4096 dimensions. It is the Qwen encoder used by Decidealot CLM.
    - CPU and CUDA llama.cpp profiles now use separate pull sidecars. Each downloads only the immutable, checksum-pinned artifacts in its own registry and serializes access to the shared model directory.
    
    - The llama.cpp wrapper drops unset OpenAI request fields before forwarding an embeddings request. LiteLLM sends `encoding_format: null`, which llama.cpp rejects.
  • v6.0.0

    v6.0.0
    
    **Moves the hardware lock to Redis so it holds across every LiteLLM worker, turns on LiteLLM's response cache, splits Redis into per-service ACL users, replaces the CPU vLLM embedding model that could not run, and makes every memory limit a fixed number instead of a share of the host's RAM.**
    
    - **The Redis `default` user is disabled.** Redis now loads an ACL file. proxq connects as the `proxq` user with `REDIS_PASSWORD`, and LiteLLM connects as the `litellm` user, which can only touch the hardware lock and response cache keys. Anything else that connected to aigate's Redis with only `REDIS_PASSWORD` must now also send the username `proxq`. Redis ACL passwords cannot contain whitespace, so check `REDIS_PASSWORD` before upgrading. Existing Redis data is kept.
    - **`local-vllm-nomic-embed-v2` is removed.** Nomic Embed v2 is a Mixture-of-Experts model, and the vLLM CPU build has no MoE kernels, so every request to this alias failed. The CPU variant now serves `local-vllm-bge-m3` (multilingual, 8192-token context, 1024 dimensions) and `local-vllm-nomic-embed-v1.5` (English, 2048-token context, 768 dimensions). Vectors from different models are not comparable, so anything indexed with the old alias needs to be embedded again. `local-vllm-cuda-nomic-embed-v2` is unchanged.
    - **`make limits` no longer sizes services from a share of the host's RAM, and `MAXUSE` is gone.** Every service's memory and CPU limit is now the fixed default in `docker-compose.yml`, overridable per service in `.env`. `make limits` prints the enabled services with their limits, warns when the worst case does not fit in RAM, and writes only CPU caps into `.env.limits` for services whose default exceeds the host's core count. **Run `make limits` after upgrading.** An `.env.limits` written by an earlier version still holds percentage-based memory limits that override the new defaults. On a 96 GB host the old script gave talkies-cuda 5.4g, below what its NeMo models need to load, so they were OOM-killed on first use.
    
    - The resource manager's hardware lock is a Redis lock shared by every LiteLLM worker process (`litellm/callbacks/hardware_lock.py`). It replaces the per-process `asyncio.Semaphore`, which only serialized requests inside one worker, so with `LITELLM_WORKERS=4` up to four CUDA jobs could run at once. The holder stores a random token, refreshes the expiry while it holds the lock, and deletes the key only while the token is still its own. A crashed worker's lock expires after `RESOURCE_LOCK_TTL_SECONDS` (default `60`), and a lock whose release was missed stops being refreshed after `RESOURCE_LOCK_MAX_HOLD_SECONDS` (default `86400`). When Redis cannot be reached the request fails instead of running unlocked.
    - MCP inference tools on predictalot and decidealot take the hardware lock and evict competing groups, the same as LiteLLM-routed models. Listing tools and `unload_models` skip the lock.
    - LiteLLM's response cache is on. An identical request within 10 minutes returns the stored response from Redis without calling the model. The cache settings were in `general_settings`, where LiteLLM ignores them, so no earlier version cached anything. They now sit in `litellm_settings`, and the cache keys live under `aigate:cache:`.
    - `LITELLM_REDIS_PASSWORD` sets the `litellm` Redis user's password, used by the hardware lock and the response cache. It falls back to `REDIS_PASSWORD`.
    - `make test-unit` runs the lock and resource manager unit tests in the LiteLLM image against a throwaway Redis that uses the ACL from `docker-compose.yml`. It needs no running stack.
    - llama.cpp model entries accept `"--threads", "auto"`, which resolves to the container's CPU quota. The CPU Surya entry uses it.
    
    - proxq `v0.9.1` to `v0.11.1`, which adds the Redis username setting and takes dependency security fixes.
    - talkies-cuda's memory limit defaults to `10g` (was `12g`). Measured peaks while loading: canary-1b-flash 7.4 GB, canary-qwen-2.5b 7.3 GB, parakeet-tdt-0.6b-v3 5.5 GB.
    - The CPU llama.cpp request timeout defaults to `1800` seconds (was `300`), because OCR of one page on CPU can take several minutes. The CUDA variant keeps `300`.
    - Every Dockerfile base image is pinned by digest. The vLLM CPU image moves from the floating `latest-x86_64` tag to `v0.21.0`, the same release as the CUDA image.
    - vLLM embedding models use `--runner pooling`, which replaced `--task embed` in current vLLM. Nomic Embed v2's context is set to 512 tokens, its trained maximum.
    - The MCP `generate_image` tool tries local models first. Hugging Face no longer serves `hf-flux-schnell`, so every call without an explicit model failed.
    - The MCP `generate_tts` tool picks a default voice per model, `af_heart` for Kokoro and `alloy` for the others. Kokoro rejects OpenAI's bare voice names.
    
    - The llama.cpp CPU image ran `llama-server` with one thread per host core, ignoring the container's CPU limit. The kernel throttled the extra threads at every sync point and Surya decoded at about 0.06 tokens per second. With `--threads auto` it decodes at about 3.5.
    - The llama.cpp and vLLM idle sweepers measured idle time from the start of the last request, so a request running longer than the idle TTL had its model killed under it. A running request now counts as use.
    - The vLLM wrapper forwarded `null` for unset optional fields, which vLLM rejects. LiteLLM sends `"encoding_format": null` on every embedding request. The wrapper drops null fields before forwarding and rejects a JSON body that is not an object with `400`.
    - Ollama vision requests failed because the LiteLLM image does not ship Pillow. The image now installs `pillow==12.3.0`, pinned by hash.
    - `vllm-pull` downloads the `nomic-ai/nomic-bert-2048` model code that the Nomic models load at startup. The vLLM containers run offline and could not load them without it. ONNX exports are skipped.
    - The test suite sends the same token fallbacks as `docker-compose.yml`, expects a registered model only when its provider is enabled, checks claudebox at `/healthz`, and uses timeouts and model names that match the current services.
  • v5.7.0

    v5.7.0
    
    **Brings predictalot and decidealot under the resource manager and moves predictalot to `v1.2.1`, which adds a model unload endpoint.**
    
    - The resource manager (`litellm/callbacks/resource_manager.py`) now treats predictalot and decidealot as competing groups on CUDA and CPU. Before any LiteLLM-routed local model runs (Ollama, sd.cpp, talkies, vLLM, llama.cpp), it calls `POST /v1/models/unload` on both, with each service's own token. A `409` means the service is mid-request. The resource manager logs it as a skipped unload and continues, and the service frees its model on its own idle timer.
    - `POST /v1/unload/cuda` and `POST /v1/unload/cpu` include predictalot and decidealot. A service busy with a request reports `"status": "busy"` and keeps its model.
    - The `mcp` service receives `PREDICTALOT_AUTH_TOKEN` and `DECIDEALOT_AUTH_TOKEN`, with the same `AIGATE_TOKEN` fallback the services use.
    - Tests cover the unload endpoint on both services, the new `unload_models` predictalot MCP tool, and eviction by the resource manager on each variant. Each eviction test loads a model, sends a local Ollama embedding through LiteLLM, and checks that LiteLLM logs the unload.
    
    - predictalot and predictalot-cuda bumped `v1.0.1` to `v1.2.1`. v1.2.0 added `POST /v1/models/unload` and the `unload_models` MCP tool, for 27 tools in total. It also made `"unload": true` on a forecast wait for concurrent forecasts on the same model, made the Sundial sidecar free its weights, and turned Moirai-2 multivariate forecasts past 64 steps into a `400` instead of a `503`. v1.2.1 only changes how the images are built. No `PREDICTALOT_*` variable changed.
    - Direct requests to audiolla, flickies, predictalot, or decidealot still do not evict LiteLLM-routed models. The docs now say so.
    
    - `tests/test_predictalot.sh` sent `$PREDICTALOT_AUTH_TOKEN` directly, which is empty unless set in `.env`, so every authenticated predictalot test failed with `401` on a default setup. The tests now fall back to `AIGATE_TOKEN`, matching the service.
  • v5.6.0

    v5.6.0
    
    **Moves decidealot to `v0.4.1`, which lets aigate configure the MCP host allowlist instead of overriding the `Host` header in LiteLLM.**
    
    - `decidealot` and `decidealot-cuda` bumped `v0.4.0` to `v0.4.1`. The new image keeps the MCP SDK's DNS-rebinding protection on and reads its accepted `Host` and `Origin` values from configuration.
    - Each container now sets `DECIDEALOT_MCP_ALLOWED_HOSTS` to loopback plus its own service name (`decidealot` or `decidealot-cuda`). The LiteLLM MCP fragments no longer override `Host`, and LiteLLM calls each container by its real name. Upstream's default list names only `decidealot`, so the CUDA service needs its own list.
    - The nginx routes still send `Host: 127.0.0.1:8080` upstream, because aigate cannot know the public hostname a client uses, such as a tailnet name or a tunnel domain. Direct MCP through nginx keeps working from any of them without extra config.
    
    - `DECIDEALOT_MCP_ALLOWED_HOSTS` and `DECIDEALOT_CUDA_MCP_ALLOWED_HOSTS` override the per-variant host allowlists. `DECIDEALOT_MCP_ALLOWED_ORIGINS` sets the browser `Origin` allowlist for both, defaulting to loopback HTTP origins. Add your public origin there for a browser-based MCP client.
    - `tests/test_decidealot.sh` calls `list_models` through the aggregated `/mcp/` for each variant, which fails if a service name drops out of its container's allowlist.
  • v5.5.0

    v5.5.0
    
    **Adds decidealot, local typed decisions with the Laya and Von models, as an opt-in CPU and CUDA service with REST and MCP.**
    
    - `decidealot` (`DECIDEALOT=1`) and `decidealot-cuda` (`DECIDEALOT_CUDA=1`) on `psyb0t/decidealot:v0.4.0`, routed at `/decidealot/` and `/decidealot-cuda/`. `POST /v1/systemone` takes a `model` selector, a `state`, and named `choice`, `score`, or `noul` questions, and returns typed answers with model probabilities. The routes also serve `GET /v1/models`, `POST /v1/models/unload`, and an unauthenticated `GET /health`. See `docs/services/decidealot.md`.
    - MCP at `/decidealot/mcp` and `/decidealot-cuda/mcp` with `system_one`, `list_models`, and `unload_models`. The tools also join the aggregated `/mcp/` as `decidealot-*` and `decidealot_cuda-*`.
    - `DECIDEALOT_AUTH_TOKEN`, defaulting to `AIGATE_TOKEN` for both the services and LiteLLM's MCP client.
    - Both containers run as `DECIDEALOT_UID:DECIDEALOT_GID` (default `1000:1000`) with a read-only root filesystem, all capabilities dropped, `no-new-privileges`, `init`, and memory, CPU, and PID limits. The CUDA variant adds an exec-allowed `/var/cache` tmpfs, because Triton loads the kernels it compiles at runtime from there.
    - `DATA_DIR_DECIDEALOT` (default `.data/decidealot`) holds the two model bundles (~5.3 GB), shared by both variants. The first start downloads them before `/health` passes, and the health check allows 15 minutes for it. The repo ships `.data/decidealot/models/` so a fresh clone gets it owned by the cloning user. `HF_TOKEN` is passed through to raise the Hugging Face rate limit.
    - Per-route rate limits (`RATELIMIT_DECIDEALOT[_CUDA][_BURST]`, default `120r/m`), a shared `TIMEOUT_DECIDEALOT` (default `10m`), and variables for idle unload, request and start timeouts, body size, and log level. All are listed in `.env.example`.
    - `tests/test_decidealot.sh` covers health, bearer enforcement, the model catalog, a live decision, the 401, 422, and 413 rejections, direct MCP from a non-local hostname, and the aggregated tools for each variant.
    
    - decidealot v0.4.0 leaves the MCP SDK's DNS-rebinding protection at its localhost-only default, so its MCP endpoint answers `421` to any other `Host` header. The nginx routes and the LiteLLM MCP fragments send `Host: 127.0.0.1:8080` upstream, so MCP works from the Docker network, a tailnet name, or a tunnel domain.
  • v5.4.1

    v5.4.1
    
    **Fixes `claudebox-*` models silently answering from pibox-zai instead of Claude when called through the gateway.**
    
    - `claudebox-*` models failed through the gateway unless `CLAUDEBOX_API_TOKEN` was set in `.env`. The claudebox container falls back to `AIGATE_TOKEN` for its API token, but LiteLLM read `CLAUDEBOX_API_TOKEN` with no fallback, sent an empty key, and got `401`. The request then moved down the fallback chain to `pibox-zai-glm-5.3-flash` or `pibox-zai-glm-5.3` and still returned `200`, so the only sign was the `model` field in the response. LiteLLM now gets `CLAUDEBOX_API_TOKEN` and `PIBOX_ZAI_API_TOKEN` with the same `AIGATE_TOKEN` fallback the agent services use. `PIBOX_ZAI_API_TOKEN` had the same gap and only worked when set explicitly.
    - Recreate the `litellm` container to pick up the change.
  • v5.4.0

    v5.4.0
    
    **Moves claudebox to `v2.4.5` and both pibox services to `v0.18.4`. Streaming chat completions on all three can now carry the agent's native event records.**
    
    - Streaming requests to `/claudebox/openai/v1/chat/completions`, `/pibox-zai/openai/v1/chat/completions`, and `/pibox/openai/v1/chat/completions` accept `"stream_options": {"include_aicodebox_events": true}`. The response then adds named `aicodebox.native` SSE events next to the normal OpenAI chunks, each wrapping a raw record from the agent. Content chunks and `[DONE]` are unchanged, and a stream without the option carries no extra events. Sending the option without `"stream": true` returns `400`. See `docs/services/claudebox.md`.
    
    - claudebox image bumped `v2.3.10` to `v2.4.5`. A container created from the new image installs Claude Code `2.1.280` on first start. Claude Code lives in the container filesystem, not the config volume, so the new version arrives when the container is recreated.
    - pibox image bumped `v0.16.2` to `v0.18.4` for both `pibox-zai` and `pibox`. This updates pi-coding-agent to `0.85.1`.
    - No environment variables changed. Existing `.env` files work as they are.
  • v5.3.0

    **Adds five current OpenRouter free models to the gateway.**
    
    - `or-nemotron-ultra`, backed by NVIDIA Nemotron 3 Ultra.
    - `or-qwen3.8-27b`, backed by Qwen 3.8 27B.
    - `or-ling-3-vl`, `or-gemma-4-31b`, and `or-inkling` for multimodal requests.
    - General and multimodal fallback chains for the new aliases.
  • v5.2.0

    **Gives `pibox` the same tailnet egress wiring `claudebox` and `pibox-zai`
    already had.**
    
    - `pibox` joins the tailnet egress overlay, so it gets the split DNS and route
      helper the other two agent containers get when `TAILSCALE=1` is on. Without
      it the service reached the tailnet only on hosts that already route
      `100.64.0.0/10` themselves, and had no tailnet access at all on a host that is
      not a tailnet node.
  • v5.1.1

    **Bumps both pibox services to `v0.16.2`, which fixes every advertised model
    except the default one failing on each request.**
    
    - pibox image bumped `v0.16.1` to `v0.16.2`. On v0.16.1 only
      `PIBOX_PROVIDER_MODEL` was registered in Pi's provider model list, so any
      other model in `PIBOX_AVAILABLE_MODELS` fell back to Pi's default API shape,
      disagreed with the configured provider base URL, and failed with
      `Stream ended without finish_reason`. Both services shipped in v5.1.0 with
      this fault: `pibox-zai` served `glm-5.3-flash` but not `glm-5.3`, and `pibox`
      served only its default model out of everything listed in `PIBOX_MODELS`.
      Picking a model from the advertised list is the point of these services, so
      v5.1.0 should be skipped. v0.16.2 registers every advertised model under the
      configured `PIBOX_PROVIDER_API`, and warns at startup when the provider base
      URL and API protocol describe different protocols rather than surfacing the
      mismatch as a truncated stream at request time.
  • v5.1.0

    **Adds `pibox`, an agent that runs on the models this stack already serves. No
    extra provider account and no second subscription. A local Ollama or vLLM model,
    or a free cloud model, drives the agent loop.**
    
    - `pibox` service, opt-in with `PIBOX=1`, reachable at `/pibox/` with its MCP
      server at `/pibox/mcp/`. It is [pi-coding-agent](https://github.com/earendil-works/pi-mono)
      in API mode with its upstream pointed at this stack's own LiteLLM, so any model
      in `/v1/models` becomes an agent backend with shell, file, and MCP tool use. It
      exposes the same REST API, OpenAI-compatible endpoint, `/files/*` CRUD, and MCP
      server as pibox-zai, in its own container with its own workspace.
      - `PIBOX_MODELS` lists the models it offers and `PIBOX_DEFAULT_MODEL` picks the
        one used when a caller names none. Both ship with defaults spanning groq,
        Ollama CUDA, Cohere, HuggingFace, and OpenRouter. Trim them to what you have
        enabled.
      - A listed model has to be able to call tools. The agent loop is tool driven,
        so a model that answers with text instead of a tool call stalls on the first
        turn. Capability does not track size: `local-ollama-cuda-qwen3-8b` calls
        tools and the larger `local-ollama-cuda-qwen3-30b-a3b` does not, while
        `local-ollama-cuda-deepseek-coder-v2-16b` predates Ollama's tool support and
        reports no `tools` capability at all. `ollama show <model>` lists what a
        given model supports.
      - No GPU is required. Tool capability belongs to the model file, so a
        `local-ollama-cpu-*` model calls tools exactly as well as the CUDA copy of
        the same tag and only speed differs. An agent loop is many turns, so a CPU
        run is slow rather than impossible. llamacpp and vllm cannot back this
        agent: llamacpp serves only an OCR model, and vllm serves a 0.6B chat model
        and an embedding model.
      - **Do not list `claudebox-*`, `pibox-*`, or any model whose fallback chain
        reaches one.** Those route back into an agent and the run recurses.
      - `PIBOX_UPSTREAM_KEY` sets the key the agent presents to LiteLLM. It defaults
        to `LITELLM_MASTER_KEY`, which itself defaults to `AIGATE_TOKEN`, so it works
        unconfigured. Point it at a LiteLLM virtual key to track this agent's
        requests and spend separately from the rest of the gateway.
    
    - pibox image bumped `v0.15.12` to `v0.16.1` for both services. v0.16.0 added
      generic upstream provider configuration, which is what lets an agent target a
      LiteLLM endpoint.
    - pibox-zai moved from the `ANTHROPIC_*` variables to `PIBOX_PROVIDER_*`. Same
      endpoint, same models, same Anthropic Messages protocol. The new path keeps an
      environment-variable reference to the key in Pi's provider config instead of
      writing the key value into `models.json`. `PIBOX_ZAI_BASE_URL` overrides the
      endpoint; z.ai also serves an OpenAI-compatible Coding Plan endpoint at
      `https://api.z.ai/api/coding/paas/v4`.
    
    - `make down` now tears down the `piston`, `llamacpp`, and `llamacpp-cuda`
      profiles. They were never listed, so those containers survived a `make down`
      and had to be stopped by hand.
    - `make limits` sizes the new `pibox` service. Without it the service would have
      kept the fixed compose fallback on every machine while every other service got
      limits scaled to the host.
  • v5.0.0

    **Reverses the `docker-compose.yml` change from v4.0.0. The base compose file is
    tracked again, and local changes belong in `docker-compose.override.yml`.**
    
    v4.0.0 made `docker-compose.yml` a local untracked file so an update could not
    overwrite it. That was the wrong mechanism for this file. It carries 46 service
    definitions, 22 nginx routes, and 19 rate-limit zones that have to move together
    with the provider configs, the Makefile profiles, and the LiteLLM config
    builder. A frozen copy silently breaks: a later release adds a service, ships
    its provider YAML and profile flag, and the local compose has no matching
    service, so LiteLLM registers a model pointing at a host that does not resolve
    and nginx has no route for it.
    
    Compose already solves this with an override file, which is what this release
    uses.
    
    - **`docker-compose.yml` is tracked again and an update overwrites edits to it.**
      Put local changes in `docker-compose.override.yml`, which is gitignored and
      merged last, so it wins over the base and over any bundled overlay. Write only
      the keys being changed:
    
      ```yaml
      services:
        claudebox:
          mem_limit: 8g
      ```
    
      **Upgrading from v4.0.0 needs one manual step.** v4.0.0 left an untracked
      `docker-compose.yml` in the working tree and this release adds that same path
      as a tracked file, so Git refuses the checkout with `untracked working tree
      file would be overwritten`. Move the file aside first, then pull, then port
      any edits into `docker-compose.override.yml`:
    
      ```bash
      mv docker-compose.yml docker-compose.yml.mine
      git pull
      diff -u docker-compose.yml docker-compose.yml.mine   # port what you changed
      ```
    
      Upgrading from v3.24.0 or earlier needs nothing; the file is tracked in both.
    - **`docker-compose.yml.example` is removed.** The base file is the shipped
      default again, so the copy served no purpose.
    
    - `COMPOSE_FILE` is now assembled by the Makefile in merge order: the base file,
      then `docker-compose.tailscale.yml` when `TAILSCALE=1`, then
      `docker-compose.override.yml` when it exists. Compose only auto-loads an
      override file when `COMPOSE_FILE` is unset, and the tailnet overlay sets it, so
      the override is appended explicitly rather than relying on that default.
    - `make bootstrap` creates `.env` from `.env.example` and prints the active
      compose file chain. It no longer creates `docker-compose.yml`.
    - `.env` is unchanged: still created from `.env.example` on first run, still
      gitignored. That file is settings rather than wiring, so a local copy cannot
      drift out of step with the rest of the repository.
  • v4.0.0

    **`docker-compose.yml` and `.env` are now local files created from tracked
    `.example` copies, so an update never overwrites them. Every cloud provider's
    model list was audited against the live APIs and the dead entries removed.**
    
    - **`docker-compose.yml` is no longer tracked.** The repository ships
      `docker-compose.yml.example`; `make` copies it to `docker-compose.yml` on
      first run, and `.gitignore` covers the copy. Pulling this release deletes the
      tracked file from your checkout. If you had local edits to it, save them
      first, then reapply them after any `make` target recreates the file. To change
      the shipped defaults for everyone, edit `docker-compose.yml.example` instead;
      edits to `docker-compose.yml` cannot be committed.
    - **Model aliases removed.** Each of these returned a result before and no
      longer resolves. The ones marked as already failing answered with a fallback
      model behind an HTTP 200 rather than an error, so callers may not have
      noticed.
      - pibox-zai, all of which worked because z.ai auto-routed them:
        `pibox-zai-glm-5.2`, `pibox-zai-glm-5.1`, `pibox-zai-glm-5-turbo`,
        `pibox-zai-glm-5`, `pibox-zai-glm-4.7`, `pibox-zai-glm-4.6`,
        `pibox-zai-glm-4.5`, `pibox-zai-glm-4.5-air`. Replace with
        `pibox-zai-glm-5.3` (for 5.2, 5.1, 5) or `pibox-zai-glm-5.3-flash` (for 4.7
        and below).
      - OpenRouter, already failing: `or-hermes-3-405b`, `or-qwen3-coder`,
        `or-qwen3-80b`, `or-llama-3.3-70b`, `or-gpt-oss-120b`, `or-gpt-oss-20b`,
        `or-nemotron-ultra-550b`, `or-nemotron-nano-9b`, `or-nemotron-nano-30b`.
      - HuggingFace, already failing: `hf-qwq-32b`, replaced by `hf-qwen3-32b`, and
        `hf-qwen3-vl-8b`, replaced by `hf-gemma-3-27b`.
      - Cerebras, archived upstream: `cerebras-glm-4.7`.
    - **`PIBOX_ZAI_AVAILABLE_MODELS` and `PIBOX_ZAI_DEFAULT_MODEL` defaults
      changed** to `glm-5.3,glm-5.3-flash` and `glm-5.3-flash`. An `.env` pinning
      the old values will fail. Drop the override or set the new ids.
    
    - `make bootstrap` creates `.env` and `docker-compose.yml` from their `.example`
      files. Every other target seeds them first, so `make run` on a fresh clone
      works with no manual copy step. The copy runs while `make` parses the
      Makefile, before it reads `.env`, so profile flags in a freshly created `.env`
      take effect on that same invocation.
    - OpenRouter: nine free models replacing the nine that no longer resolve.
      `or-nemotron-lightning` (1M context), `or-nemotron-120b`, `or-dots-3-note`
      (text and image), `or-nemotron-omni-30b` (omni, reasoning),
      `or-north-mini-code`, `or-lfm-2.5-2.6b`, `or-ling-3-sante`, `or-ling-3-fin`,
      and `or-nemotron-content-safety`. The last three answer by name but stay out
      of the general fallback chains, where a domain-tuned model or a classifier
      would answer off-target.
    - Groq: `groq-qwen3.8-27b`, `groq-allam-2-7b` for Arabic, and the
      `groq-prompt-guard-22m` and `groq-prompt-guard-86m` prompt-injection
      classifiers, which return a probability score rather than chat text.
    - Cohere: `cohere-command-a-plus`, `cohere-command-a-reasoning`,
      `cohere-command-a-vision`, `cohere-command-a-translate`,
      `cohere-north-mini-code`, `cohere-command-r7b-arabic`,
      `cohere-aya-vision-32b`, and `cohere-tiny-aya-global`, `-earth`, `-fire`,
      `-water`.
    - HuggingFace: `hf-qwen3-235b` and `hf-gemma-3-27b`.
    - Cerebras: `cerebras-qwen3.8-27b` and `cerebras-gemma-4-31b`.
    
    - **`groq-compound` and `groq-compound-mini` reached the wrong vendor.** Groq
      namespaced both ids under `groq/` upstream, so the old pins stopped resolving
      and the fallback chain answered with a Cohere model behind an HTTP 200. The
      pins are now `groq/groq/compound` and `groq/groq/compound-mini`: the first
      `groq/` selects the provider, the second belongs to the model id.
    - Fallback chains no longer point at models that do not exist. Every chain key
      and every fallback target resolves to a registered model. That includes
      `local-talkies-cuda-qwen3-tts-0.6b`, a long-standing typo for
      `local-talkies-cuda-qwen3-tts`.
    
    - pibox-zai is documented as running on a [GLM Coding
      Plan](https://z.ai/subscribe) rather than generic z.ai credits, and the plan's
      two models are the only exposed aliases. The provider page carries the
      token-to-credit formula and the per-model rates, with a note that those rates
      are conversion factors and not multipliers against an older baseline.
    - The routing tier previously called "flat-rate" is now "subscription". The
      claim that it costs the subscription with "no extra per-call charge" is
      replaced with the accurate statement that the allowance is metered.
    - Cerebras is documented as requiring a paid plan. At the last audit every model
      returned `Payment required to access this resource` on a free account, so the
      free-tier claims in the README, the provider page, and `.env.example` were
      wrong.
  • v3.24.0

    **Bump the claudebox and pibox agent images. Both are rebuilt on the aicodebox
    v0.14.6 base, which adds native full-event retention on `POST /run`.**
    
    - claudebox image bumped `v2.0.13` to `v2.3.10`. As of the image's v2.3.0,
      Claude Code is no longer baked into the image (Anthropic's CLI carries no
      redistribution grant); it installs from npm on first container start, so a
      fresh claudebox needs outbound network and a few extra seconds on first boot.
      The version is pinned via `CLAUDEBOX_CLAUDE_VERSION` and warm restarts skip
      the install. The v0.14.6 base adds native full-event retention on `POST /run`
      (`eventMode`) and Claude Code's native `--json-schema` flag, and the chain
      since v2.0.13 inherited an API-mode restart-loop fix (the agent subprocess is
      spawned in its own session, so its signals no longer reach the uvicorn PID 1).
    - pibox image bumped `v0.15.11` to `v0.15.12`, rebuilt on the same aicodebox
      v0.14.6 base for native full-event retention on `POST /run`.
  • v3.23.0

    **Outbound tailnet access for the claudebox and pibox-zai agent containers.
    When `TAILSCALE=1` runs alongside `CLAUDEBOX=1` or `PIBOX_ZAI=1`, those
    containers can reach machines on your tailnet over any protocol, on top of the
    existing inbound `tailscale serve` proxy.**
    
    - Tailnet egress overlay (`docker-compose.tailscale.yml`), loaded by the
      Makefile only when `TAILSCALE=1`. It routes the `100.64.0.0/10` Tailscale
      CGNAT range out through the existing tailscale node and configures split DNS,
      so claudebox and pibox-zai resolve tailnet names and connect to tailnet peers
      while public and sibling-container names keep resolving. Both containers stay
      on the bridge, so nginx and LiteLLM still reach them.
      - The tailscale container becomes a NAT gateway (`MASQUERADE` on `tailscale0`,
        kept present by a sidecar). Each agent container gets the tailnet route
        installed in its network namespace by a per-container sidecar that re-adds it
        after a restart. All egress sidecars run with `no-new-privileges:true` and
        `cap_drop: [ALL]`, adding back only `NET_ADMIN` (plus `NET_RAW` on the
        gateway sidecar for the NAT rule).
      - DNS: Tailscale MagicDNS (`100.100.100.100`) answers tailnet names and
        SERVFAILs everything else; Docker's embedded resolver falls through to a
        public resolver for public names.
    - `TS_MAGICDNS_SUFFIX` (optional): your tailnet's MagicDNS suffix, so the agent
      containers resolve bare tailnet names in addition to FQDNs. Leave unset to
      require fully-qualified names.
    - `TS_FALLBACK_DNS` (optional, default `1.1.1.1`): resolver the agent containers
      use for non-tailnet names. Set your own to keep public DNS on your
      infrastructure.
    
    Scope: IPv4 tailnet peers by their `100.64.0.0/10` address or MagicDNS name. A
    LAN behind a subnet router and tailnet IPv6 are not routed. See
    [docs/services/tailscale.md](docs/services/tailscale.md).