v6.1.0

Adds a CUDA Qwen3 8B embeddings endpoint for Contrastive-LM CLM and makes llama.cpp model downloads artifact-pinned.

- `local-llamacpp-cuda-qwen3-8b` serves Qwen3 8B Q8_0 at `/v1/embeddings` with last-token pooling, L2 normalization, and 4096 dimensions. It is the Qwen encoder used by Decidealot CLM.
- CPU and CUDA llama.cpp profiles now use separate pull sidecars. Each downloads only the immutable, checksum-pinned artifacts in its own registry and serializes access to the shared model directory.

- The llama.cpp wrapper drops unset OpenAI request fields before forwarding an embeddings request. LiteLLM sends `encoding_format: null`, which llama.cpp rejects.