Tags give the ability to mark specific points in history as being important
-
26.10.0-rc2
34ec2289 · ·Ryax 26.10.0-rc2 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ # Ryax 26.10.0 ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ GPU requests by model, memory and share of a card, a random initial admin password, a broker and a filestore that no longer depend on Bitnami, and a round of fixes for imports, stuck executions and failed deployments. ## New features ### GPU scheduling by model, memory and share of a card Ryax used to know a GPU only as a count, plus an opaque MIG profile name that had to match exactly. This release describes GPUs the way users think about them, from the action that asks to the node pool that answers. - **Actions can ask for a kind of GPU.** In `ryax_metadata.yaml`, `spec.resources.gpu` still takes a number, and now also takes an object with a `count`, a `brand`, a `model`, a `memory` (with a unit, for example `40G`) and a `compute_fraction` (a share of one card, between 0 and 1, exclusive of 0). Everything in it is optional and describes **one** GPU: `count: 2` with `memory: 40G` means two cards of 40 GB each, not 80 GB across two. `gpu: 2` keeps its meaning of two GPUs of any kind. Brand and model must match the catalog; an action that names one no node pool has simply finds no place to run. A request with an unknown key, or a memory without a unit, is rejected when the action is scanned. - **Node pools declare their card.** A GPU node pool now records its catalog model next to its partition (`full` or a MIG profile), so Ryax knows how much memory and compute one GPU of the pool really gives. The new `GET /api/runner/gpu-models?search=a100 1g` lists the catalog as one flat, searchable list (brand, model, partition, memory, compute share), and the node pool form in the web interface now picks the GPU with one searchable select. Existing node pools keep working: one with no recorded model stays eligible for every request, and is simply treated as of unknown size, never as too small. Fill in the model to get precise matching. - **Placement takes any partition that is big enough, and the tightest one by default.** A recommendation of "3g.40gb" used to demand exactly that profile and fail when no pool had it; now any partition that meets the request will do. New setting `RYAX_SCHEDULER_GPU_FIT_POLICY` on the Runner: `best_fit` (the default) picks the pool that wastes the least of a card, leaving large partitions free for the actions that need them; `first_fit` keeps the previous behaviour of taking the best-scoring pool that fits. GPU requests are also now checked against the pool's CPU, memory and time, which they never were. - **IntelliScale recommends a share of a card, not a MIG profile.** It now recommends GPU memory (the observed peak plus 10%) and a compute share, and Ryax rounds them up to a partition your pools actually offer. This also works on cards that do not split in sevenths such as the A30. - **Recommendations are learned per piece of hardware, not per site.** A recommendation is only valid for the machine it was measured on, so IntelliScale now keeps one model per GPU model (for the GPU share) and per instance type (for CPU and memory). Two node pools of one site with different cards no longer contaminate each other, and two sites with the same card now learn together. The Runner asks for the recommendation after it has picked a candidate pool, one answer per pool, instead of merging every site's answer beforehand. Nothing fragments on upgrade: pools with no recorded GPU model, and HPC sites, share one model as before. - **Autoscaled GPU nodes can be kept out of use until they are ready.** An autoscaled GPU node is reported `Ready` long before its driver and MIG layout are in place, so actions landing on it got a whole GPU instead of their slice. The Kubernetes worker chart can now hold new GPU nodes behind a startup taint and release them only when the GPU stack is usable and the MIG layout is the one the pool asked for. Off by default: see `gpuReadiness` and the new [GPU node pools and MIG](https://docs.ryax.tech/howto/gpu_node_pools/) guide. ([roadmap#1421](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1421), [roadmap#1458](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1458), [roadmap#1469](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1469), [roadmap#1428](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1428)) ### Security and operations - **The initial admin password is random.** Every installation used to boot with the same `user1` / `pass1` — hard-coded in the service, never set by the chart, and published in the README and the install guide. The chart now generates one per installation into the `ryax-admin-credentials` secret, and `helm install` prints the command to read it back. The account is `admin`. Existing installations keep their users and passwords: the secret is only ever read when the user table is empty, which is the very first start. The Edit password window now warns that changing the password does not update that secret. ([roadmap#1423](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1423)) - **`helm install` prints its notes.** `NOTES.txt` had always sat at the chart root rather than in `templates/`, where Helm is the only place it looks, so no install had ever printed anything. It now carries the admin credentials command and the Grafana one. - **The broker no longer runs on a Bitnami image.** It ran on `bitnamilegacy/rabbitmq`, which Bitnami no longer updates. It is now a `RabbitmqCluster` run by the official [RabbitMQ Cluster Operator](https://www.rabbitmq.com/kubernetes/operator/operator-overview) (2.23.0) on the official `rabbitmq:4.3.6-management` image. The chart deploys the operator in the Ryax namespace, watching that namespace only, and it needs neither cluster-wide rights nor cert-manager. The services reach the broker at the same address with the same credentials. **This needs one command before upgrading, see below.** ([roadmap#1320](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1320)) - **The filestore no longer runs MinIO.** MinIO's community edition is no longer maintained, and the chart ran it from the Bitnami `bitnamilegacy/minio` image. The filestore is now [versitygw](https://github.com/versity/versitygw) v1.8.0, a small S3 gateway that keeps every object as a plain file on its volume. On a test cluster it answers Ryax's requests faster than MinIO did (about 3x on small writes, 1.3x on small reads) with a tenth of its memory. The services reach it at the same address (`ryax-minio:9000`) with the same credentials (`ryax-minio-secret`), and the upgrade copies the existing objects into it. ([roadmap#1471](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1471)) - **The `Authorization: Bearer <token>` header works on every service.** The Authorization and Repository services answered 401 to the form our own README documents, while Runner and Studio accepted it. All four now agree, and a header that is only `Bearer` answers 401 on Studio instead of 500. ([roadmap#1453](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1453)) - **Deploying a workflow that is already deploying returns 409 with its state.** `POST /workflows/{id}/deploy` answered 400 for this, the same code as for an invalid workflow, so clients had to match the error text. It now answers 409 `{"error": "Workflow can't be deployed at this status", "deployment_status": "Deploying"}`. An invalid workflow still gets 400. See the upgrade notes. ([roadmap#1467](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1467)) ## Bug fixes and Improvements - **An install reached through a proxy or by IP answers again.** 26.9.0 restricted the Ingresses to `global.tls.hostname`, so any other name got a 404. They are now restricted only to `global.ingress.hosts`, which is empty (any host) by default. The documentation gained a how-to for [running Ryax behind a proxy](https://docs.ryax.tech/). ([roadmap#1448](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1448)) - **Importing a workflow keeps its addons.** An imported workflow was valid and showed Deployed, but its HTTP services had no Ingress and every URL answered 404, because addons were dropped. Export now writes the editable addon values (`addons_inputs_values`), and import restores them; packages exported before this release still import and get the addons' defaults. A value for an addon parameter the action does not have now answers 400 instead of being silently dropped. ([roadmap#1464](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1464)) - **Importing a workflow with a numeric or directory input value no longer fails with a 500.** A value such as `max_tokens: 1024` or a directory input in a re-imported export now imports correctly, and a bad package answers 400. The deprecated `table` input/output type is removed; existing values are converted to `file`. ([roadmap#1465](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1465), [roadmap#1468](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1468)) - **An HTTP service trigger is no longer killed at startup.** Addon inputs and resources were frozen at the first deploy of an action and shared by every workflow using it, so a workflow using the HTTP addon on an action first deployed without it lost the addon's inputs. Each workflow action now owns its own, and two workflows can use one action with different addons and resources. Addon values are merged parameter by parameter: a value set in the interface beats the action's, which beats the addon default. ([roadmap#1463](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1463)) - **A workflow can be undeployed when it has several running deployments.** Undeploy used to fail and leave Studio in "Undeploying", with the portal and trigger views answering 500. It now stops all of them, and a stopped deployment is never brought back to running by a late trigger event. The leftover `PAUSED` state, which nothing has set since 26.7.0, is retired. ([roadmap#1337](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1337)) - **A failed deployment says why.** A deployment whose trigger failed (no site registered, trigger error, trigger cancelled by itself, deployment that could not even be created) stayed in "Deploying" forever or went back to "not deployed" with no explanation. It now fails and Studio shows `The trigger '<name>' failed: <reason>`. ([roadmap#1445](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1445), [roadmap#1466](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1466)) - **Executions that outlive their time allotment are cancelled.** An execution whose Worker crashed or never reported its end stayed running forever, holding its resources. The Runner now cancels any action execution still running 5 minutes past its allotment (`RYAX_EXECUTION_REAPER_GRACE_SECONDS`, checked every 60 seconds, `RYAX_EXECUTION_REAPER_INTERVAL_SECONDS`; a grace of 0 disables it). Triggers, which have no allotment, are not touched. ([roadmap#1299](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1299)) - **Stopping a workflow ends all its executions.** Trigger runs of an http trigger stayed RUNNING when the trigger ended gracefully or crashed, and executions stayed STOPPING when the Worker did not know them or had restarted. They now end as cancelled, and no longer enter the retry chain. Executions already stuck are not repaired and need a one-off cleanup. ([roadmap#1473](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1473)) - **Execution logs come out in the order the action wrote them.** The Workers now number each batch of log lines and the Runner reads them in that order. Logs stored before the upgrade keep their previous order. ([roadmap#1449](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1449)) - **Upgrades no longer deadlock on Grafana and MinIO.** Their `ReadWriteOnce` volume stopped the new pod from starting while the old one still held it, so `helm upgrade --wait` timed out and marked the release failed. Both now replace their pod instead of rolling it. ([roadmap#1451](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1451)) - **The registry garbage collector works again.** It pointed at a configuration path that does not exist since the registry moved to v3, and hid the failure behind a success message, so no image was ever collected. A failing run now fails its Job. ([roadmap#1294](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1294)) - **Security and dependencies.** The web interface's advisories are cleared (Angular 21.2.25, axios 1.20, undici 6.29, piscina 5.3, adm-zip 0.6.1, brace-expansion 2.1.6; the unpatched `braces` advisory only affects the build tooling and is tracked), and Python dependencies and nixpkgs (26.05) are refreshed across all services. Observability charts: kube-prometheus-stack 89.x (Grafana 13), Alloy 1.13, Traefik 41.6. - **Documentation.** New GPU node pools guide, rewritten IntelliScale reference, `global.ingress.hosts`, proxy configuration and the RabbitMQ operator in the install, airgap and ArgoCD guides. ## Upgrade to this version Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **⚠️ If you set `global.tls.hostname`, Ryax answers for every host again.** 26.9.0 restricted the Ingresses to that name; they are now restricted only to the names in `global.ingress.hosts`, which is empty (any host) by default. To keep Ryax to its own names, typically on a cluster shared with other applications, list them: ```yaml global: ingress: hosts: ["ryax.example.com"] ``` If you keep `global.tls.hostname` as well, it must be one of these hosts, or the chart refuses to render. Installs that set neither value need nothing. More in [Install Ryax with ArgoCD › Routing](https://docs.ryax.tech/howto/install_ryax_argocd/#routing). - **A new GitOps install needs one more secret.** With `global.secrets.create=false`, create `ryax-admin-credentials` (keys `admin-user` and `admin-password`) before the first sync, or set `authorization.adminUsername` and `authorization.adminPassword`. Without it the authorization pod stops with `No initial admin password configured`. An **existing** installation needs nothing — it never seeds again, and both `secretKeyRef`s are `optional`. The service also no longer has a built-in fallback for the initial user name (now `admin` when unset, not `user1`) nor for the JWT signing key: the chart has always injected the latter from `api-jwt-secret-key`, but a deployment that starts the services without the chart must now set `RYAX_JWT_SECRET_KEY`. - **⚠️ Apply the RabbitMQ CRD before upgrading.** The broker moves to the RabbitMQ Cluster Operator, whose `RabbitmqCluster` CRD ships in the chart, and Helm never installs a CRD on an upgrade. Without it the upgrade stops before changing anything, and prints this command: ```sh kubectl apply --server-side -f https://gitlab.com/ryax-tech/ryax/ryax-engine/-/raw/26.10.0/charts/ryax/subcharts/rabbitmq/crds/rabbitmqclusters.rabbitmq.com.yaml ``` Offline, take it from the chart package instead: ```sh tar -xzOf ryax-engine-26.10.0.tgz ryax-engine/charts/rabbitmq/crds/rabbitmqclusters.rabbitmq.com.yaml \ | kubectl apply --server-side -f - ``` `--server-side` is required: the CRD is too large for a client-side apply. ArgoCD applies the CRD itself, with the `ServerSideApply=true` the reference Application already sets. - **The broker is replaced, not upgraded, and starts empty.** Helm removes the Bitnami StatefulSet and the operator starts `ryax-broker-server-0` in its place, about a minute later. The services keep the same address (`ryax-broker:5672`) and the same credentials (`ryax-broker-secret`), and reconnect by themselves. There is nothing to migrate: Ryax publishes its messages as transient, so a broker restart has always dropped whatever was still queued. As for any upgrade, run it while no workflow is running. The services retry on their own during the switch, which took about a minute on a test cluster. Once Ryax is back, delete the old broker volume, if there is one (there is none with `rabbitmq.persistence.enabled: false`, as in `minimal.yaml`): ```sh kubectl -n ryaxns delete pvc data-ryax-broker-0 ``` - **The `rabbitmq:` values now configure the new broker.** `persistence.*`, `resources`, `tolerations`, `nodeSelector`, `affinity`, `priorityClassName` and `metrics.enabled` keep their meaning. Every other Bitnami key (`auth.*`, `image.*`, `clustering`, `plugins`, ...) is ignored: remove them. A `rabbitmq.image.repository` still naming a Bitnami image stops the render. The broker password is always the one in `ryax-broker-secret`, so a `rabbitmq.auth.password` set by the old troubleshooting guide does nothing. The broker also follows `global.tolerations`, `nodeSelector` and `affinity` now. The broker blocks publishers when its volume has less than `rabbitmq.diskFreeLimit` free (default `250Mi`, RabbitMQ units); keep it well under `rabbitmq.persistence.size` if you shrink the volume. - **Uninstalling now takes one more step.** `helm uninstall` removes the operator at the same time as the broker, so nothing clears the `RabbitmqCluster` finalizer: the broker pod keeps running, and `helm uninstall --wait` times out. Delete the broker first, while the operator still runs: ```sh kubectl -n ryaxns delete rabbitmqcluster ryax-broker helm uninstall ryax -n ryaxns ``` If an uninstall is already stuck, clear the finalizer: `kubectl -n ryaxns patch rabbitmqcluster ryax-broker --type merge -p '{"metadata":{"finalizers":[]}}'`. - **If the cluster already runs a RabbitMQ Cluster Operator** that watches the Ryax namespace, set `rabbitmq.operator.enabled: false`. Two operators would reconcile the same broker. - **Multi-site with Skupper:** the `ryax-broker-ext` and `ryax-minio-ext` connectors copied the pod selectors of the old broker and of MinIO when they were created. After the upgrade the first matches no pod, and the second matches the old MinIO, which refuses their connections. Recreate both on the main site right after the upgrade, before running workflows on remote sites: ```sh skupper -n ryaxns connector delete ryax-broker-ext skupper -n ryaxns connector create ryax-broker-ext 5672 --workload service/ryax-broker skupper -n ryaxns connector delete ryax-minio-ext skupper -n ryaxns connector create ryax-minio-ext 9000 --workload service/ryax-minio ``` - **GitOps (`global.secrets.create=false`):** `ryax-broker-cookie` is no longer read, delete it whenever you like. `ryax-broker-secret` keeps its keys; its `broker-user` and `rabbitmq-password` now seed the broker's user, so they must match the `broker` URL as before. With automated pruning off, prune the old broker resources in the same sync: the operator cannot create its `ryax-broker` Service while the Bitnami one is still there. - **⚠️ The filestore moves from MinIO to versitygw, and the upgrade copies its objects.** Helm keeps the old MinIO, `ryax-minio`, running on its volume, and marks that volume so that neither Helm nor ArgoCD ever deletes it. The new filestore, `ryax-filestore`, gets a volume of its own, the size of MinIO's. Before it starts serving, its pod copies every object out of MinIO, checks that both sides list the same objects and sizes, and leaves a marker so that it never copies again. The address and the credentials do not change. **Downtime:** the services cannot reach the filestore until the copy is over. Without a pre-copy (below), count the time to read the whole MinIO volume once: on a test cluster, 750 MiB in 3,000 objects took a few seconds, and a large volume on network storage takes minutes. On a local k3s upgrade of a 26.9.0 install holding 58 MiB in 410 objects, the copy itself took a second, the filestore was unreachable for 15 to 50 seconds while MinIO restarted, and the whole of Ryax was back within 3 minutes, the broker switch being the longest part. The runner and studio restart while they wait, and their restart back-off can add up to five minutes once the copy is over. They reconnect on their own; to skip the back-off, restart them once the filestore is Ready: ```sh kubectl -n ryaxns rollout status deploy/ryax-filestore --timeout=24h kubectl -n ryaxns rollout restart deploy/ryax-runner deploy/ryax-studio ``` Follow the copy with `kubectl -n ryaxns logs -f deploy/ryax-filestore -c migrate-from-minio`. The new volume's storage class must support user extended attributes, as ext4 and xfs do. If MinIO ran without persistence (`minio.persistence.enabled: false`), it has no volume to copy from, and its objects go away with it, as they did whenever its pod restarted. The `minio:` values are ignored now: remove them, but carry a `minio.persistence.storageClass` over as `filestore.migration.legacy.persistence.storageClass`, as the class of MinIO's volume cannot change. To give the filestore more room than MinIO had, set `filestore.persistence.size`. - **Optional: pre-copy the objects to shorten the upgrade.** While 26.9.0 still runs, this copies the objects into the volume the filestore will use, and the upgrade then only copies what changed since. It can run several times, and it can be interrupted: the copy at the upgrade makes the result exact whatever it left. Use the same values file as for the upgrade: ```sh SIZE=$(kubectl -n ryaxns get pvc ryax-minio -o jsonpath='{.spec.resources.requests.storage}') helm template ryax oci://registry.ryax.org/release-charts/ryax-engine --version 26.10.0 \ -n ryaxns -f values.yaml \ --set filestore.migration.precopy.enabled=true \ --set filestore.persistence.size="$SIZE" \ --show-only charts/filestore/templates/precopy.yaml > precopy.yaml kubectl -n ryaxns delete job ryax-filestore-precopy --ignore-not-found kubectl -n ryaxns apply -f precopy.yaml kubectl -n ryaxns wait --for=condition=complete job/ryax-filestore-precopy --timeout=24h kubectl -n ryaxns logs job/ryax-filestore-precopy | tail -n 1 kubectl -n ryaxns delete job ryax-filestore-precopy ``` It creates the `ryax-filestore` volume, which the upgrade then takes over, and a Job that reads MinIO while Ryax keeps working. The upgrade refuses to start while that Job runs. Delete the Job once it is done, as above: as long as it exists, the volume cannot be deleted, which a rollback or an uninstall does. The pre-copy only shortens the upgrade when the volume is large: on the test cluster above, both copies took about a second. Offline, render it from the chart package (`helm template ryax ryax-engine-26.10.0.tgz ...`): it only uses the MinIO image 26.9.0 already ran. - **Once 26.10.0 runs, decommission the old MinIO.** It only serves the copy, and makes a rollback possible until then. First check that the copy is done: ```sh kubectl -n ryaxns exec deploy/ryax-filestore -c versitygw -- cat /data/.ryax-migrated-from-minio ``` It prints the date of the copy. Then remove MinIO, and delete its volume: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.10.0 \ -n ryaxns --reuse-values --set filestore.migration.enabled=false kubectl -n ryaxns delete pvc ryax-minio ``` The chart refuses to remove MinIO while no filestore pod has finished the copy (`filestore.migration.allowDecommissionWithoutCopy=true` forces it), and refuses `filestore.migration.enabled=false` on an upgrade from 26.9.0, which would delete MinIO's volume before the copy. Like any upgrade, the decommission restarts the runner and studio, and it restarts the filestore, unreachable for a few seconds. **After the decommission, a rollback to 26.9.0 is no longer possible** without restoring MinIO's volume from a backup. - **Rolling back to 26.9.0 before the decommission** brings MinIO back on its volume, untouched, and deletes the filestore and its volume. Objects written since the upgrade are lost, unless you copy them back into MinIO first, while 26.10.0 still runs and no workflow runs. That copy rewrites every object, not only the new ones, so it takes about as long as the copy at the upgrade. Then delete the broker while its operator still runs, for the Bitnami broker cannot take its `ryax-broker` Service back otherwise, and roll back: ```sh kubectl -n ryaxns exec deploy/ryax-minio -- sh -c 'export HOME=/tmp mc=/opt/bitnami/minio-client/bin/mc $mc alias set new http://ryax-minio:9000 "$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD" >/dev/null $mc alias set old http://localhost:9000 "$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD" >/dev/null $mc mirror --overwrite new/ryax-filestore old/ryax-filestore' kubectl -n ryaxns delete rabbitmqcluster ryax-broker helm rollback ryax <the 26.9.0 revision> -n ryaxns ``` Without the `delete rabbitmqcluster`, the rollback stops part-way on `no Service with the name "ryax-broker" found`. Do not run a failed rollback again: each attempt leaves more of both versions behind. Upgrade to 26.10.0 again with the same values file, wait for every pod to be Ready, and roll back as above. Upgrading again later copies MinIO afresh. A `ryax-worker-k8s` already on 26.10.0 can stay there: on the test cluster it ran workflows with the rolled-back engine. - **GitOps (ArgoCD, Flux):** the chart cannot see the cluster, so it always renders the old MinIO and its volume while `filestore.migration.enabled` is on, the default. The upgrade copies as above. If you changed `minio.persistence.size` or `storageClass`, set them as `filestore.migration.legacy.persistence.size` and `storageClass`, or the sync fails on the volume, which cannot shrink. To decommission, check the marker as above, then set `filestore.migration.enabled: false`: ArgoCD removes MinIO but keeps its volume (`Prune=false`), which you delete by hand. **A new GitOps install** has nothing to migrate: set `filestore.migration.enabled: false` from the start. - **`global.security.allowInsecureImages` is gone.** It only served the Bitnami charts, and the engine chart no longer has any. Remove it from your values; it is ignored. - **Grafana 13 (kube-prometheus-stack 89).** The bundled Grafana runs the `-distroless` image with a read-only root filesystem and an `emptyDir` on `/tmp`. `GF_*__FILE` environment variables and `GF_INSTALL_PLUGINS` are no longer supported. The Ryax values set none of them, and the plugins (`grafana-piechart-panel`, `grafana-clock-panel`, `vonage-status-panel`) are now preinstalled synchronously at startup, so no values change is needed, but check that the Grafana pod starts. If your own overrides set `GF_*__FILE`, `GF_INSTALL_PLUGINS` or an `extraVolumeMounts` on `/tmp`, migrate them first. The rollout also replaces the Grafana and MinIO pods instead of rolling them, so both are briefly unavailable during the upgrade. - **Kubernetes worker: the MIG auto-labeler is removed.** The values `config.MIG` and `labeler` are gone, along with the node-labeler DaemonSet. If you relied on it to turn a `gpu-pool-mig-*` node label into `nvidia.com/mig.config`, label the GPU nodes with `nvidia.com/mig.config` yourself (cloud node-pool labels are the usual way), as the NVIDIA GPU Operator's MIG Manager expects. Remove both keys from your worker values. To keep actions off GPU nodes that are not ready yet, see the new opt-in `gpuReadiness` values: enable it **before** adding the startup taint to the node pool, as a tainted pool with the gate off never runs anything. - **GPU node pools: record the card, and mind the renamed field.** Existing pools keep working without it, but record the catalog `gpu_model` on each GPU pool (web interface, or `GET /api/runner/gpu-models` then the node pool API) to get precise matching and per-hardware recommendations. If your clusters are all A30 (4 compute slices), also set `intelliscale.config.algorithm_configs.simple_mig_recommender.total_compute_slices` to 4 for pools with no recorded model. Placement now defaults to `best_fit`; to keep the previous behaviour set `RYAX_SCHEDULER_GPU_FIT_POLICY=first_fit` through `runner.extraEnv`. - **API users:** - `POST /workflows/{id}/deploy` returns **409** (with `deployment_status` in the body) instead of 400 when the workflow is already deploying, deployed or undeploying. A client that matched the 400 text must handle 409 instead. - Node pools take and return `gpu_count`, not `gpu` (the stored value is migrated). A client that still sends `gpu` when creating a node pool gets a 422. - New root endpoints `/api/runner/node-pools` (filters on GPU, site, CPU and memory; the GPU is named by `gpu_config_id`) and `GET /api/runner/gpu-models` returns a flat list. The nested `/sites/{id}/node-pools` routes remain, marked deprecated. - `GET /modules/{id}` and workflow action views now return `gpu_brand`, `gpu_model`, `gpu_memory_gb` and `gpu_compute_fraction` in `resources`. - The `table` input/output type no longer exists. - `Authorization: Bearer <token>` is accepted by all services. - **Registry garbage collection now actually runs.** The first scheduled run after the upgrade removes every untagged image left in the registry, and a failing run now shows as a failed Job. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.10.0 \ -n ryaxns \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.10.0 \ -n ryaxns \ -f worker.yaml ``` -
26.10.0-rc1
4b7c03a2 · ·Ryax 26.10.0-rc1 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ # Ryax 26.10.0 ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ GPU requests by model, memory and share of a card, a random initial admin password, a broker and a filestore that no longer depend on Bitnami, and a round of fixes for imports, stuck executions and failed deployments. ## New features ### GPU scheduling by model, memory and share of a card Ryax used to know a GPU only as a count, plus an opaque MIG profile name that had to match exactly. This release describes GPUs the way users think about them, from the action that asks to the node pool that answers. - **Actions can ask for a kind of GPU.** In `ryax_metadata.yaml`, `spec.resources.gpu` still takes a number, and now also takes an object with a `count`, a `brand`, a `model`, a `memory` (with a unit, for example `40G`) and a `compute_fraction` (a share of one card, between 0 and 1, exclusive of 0). Everything in it is optional and describes **one** GPU: `count: 2` with `memory: 40G` means two cards of 40 GB each, not 80 GB across two. `gpu: 2` keeps its meaning of two GPUs of any kind. Brand and model must match the catalog; an action that names one no node pool has simply finds no place to run. A request with an unknown key, or a memory without a unit, is rejected when the action is scanned. - **Node pools declare their card.** A GPU node pool now records its catalog model next to its partition (`full` or a MIG profile), so Ryax knows how much memory and compute one GPU of the pool really gives. The new `GET /api/runner/gpu-models?search=a100 1g` lists the catalog as one flat, searchable list (brand, model, partition, memory, compute share), and the node pool form in the web interface now picks the GPU with one searchable select. Existing node pools keep working: one with no recorded model stays eligible for every request, and is simply treated as of unknown size, never as too small. Fill in the model to get precise matching. - **Placement takes any partition that is big enough, and the tightest one by default.** A recommendation of "3g.40gb" used to demand exactly that profile and fail when no pool had it; now any partition that meets the request will do. New setting `RYAX_SCHEDULER_GPU_FIT_POLICY` on the Runner: `best_fit` (the default) picks the pool that wastes the least of a card, leaving large partitions free for the actions that need them; `first_fit` keeps the previous behaviour of taking the best-scoring pool that fits. GPU requests are also now checked against the pool's CPU, memory and time, which they never were. - **IntelliScale recommends a share of a card, not a MIG profile.** It now recommends GPU memory (the observed peak plus 10%) and a compute share, and Ryax rounds them up to a partition your pools actually offer. This also works on cards that do not split in sevenths such as the A30. - **Recommendations are learned per piece of hardware, not per site.** A recommendation is only valid for the machine it was measured on, so IntelliScale now keeps one model per GPU model (for the GPU share) and per instance type (for CPU and memory). Two node pools of one site with different cards no longer contaminate each other, and two sites with the same card now learn together. The Runner asks for the recommendation after it has picked a candidate pool, one answer per pool, instead of merging every site's answer beforehand. Nothing fragments on upgrade: pools with no recorded GPU model, and HPC sites, share one model as before. - **Autoscaled GPU nodes can be kept out of use until they are ready.** An autoscaled GPU node is reported `Ready` long before its driver and MIG layout are in place, so actions landing on it got a whole GPU instead of their slice. The Kubernetes worker chart can now hold new GPU nodes behind a startup taint and release them only when the GPU stack is usable and the MIG layout is the one the pool asked for. Off by default: see `gpuReadiness` and the new [GPU node pools and MIG](https://docs.ryax.tech/howto/gpu_node_pools/) guide. ([roadmap#1421](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1421), [roadmap#1458](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1458), [roadmap#1469](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1469), [roadmap#1428](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1428)) ### Security and operations - **The initial admin password is random.** Every installation used to boot with the same `user1` / `pass1` — hard-coded in the service, never set by the chart, and published in the README and the install guide. The chart now generates one per installation into the `ryax-admin-credentials` secret, and `helm install` prints the command to read it back. The account is `admin`. Existing installations keep their users and passwords: the secret is only ever read when the user table is empty, which is the very first start. The Edit password window now warns that changing the password does not update that secret. ([roadmap#1423](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1423)) - **`helm install` prints its notes.** `NOTES.txt` had always sat at the chart root rather than in `templates/`, where Helm is the only place it looks, so no install had ever printed anything. It now carries the admin credentials command and the Grafana one. - **The broker no longer runs on a Bitnami image.** It ran on `bitnamilegacy/rabbitmq`, which Bitnami no longer updates. It is now a `RabbitmqCluster` run by the official [RabbitMQ Cluster Operator](https://www.rabbitmq.com/kubernetes/operator/operator-overview) (2.23.0) on the official `rabbitmq:4.3.6-management` image. The chart deploys the operator in the Ryax namespace, watching that namespace only, and it needs neither cluster-wide rights nor cert-manager. The services reach the broker at the same address with the same credentials. **This needs one command before upgrading, see below.** ([roadmap#1320](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1320)) - **The filestore no longer runs MinIO.** MinIO's community edition is no longer maintained, and the chart ran it from the Bitnami `bitnamilegacy/minio` image. The filestore is now [versitygw](https://github.com/versity/versitygw) v1.8.0, a small S3 gateway that keeps every object as a plain file on its volume. On a test cluster it answers Ryax's requests faster than MinIO did (about 3x on small writes, 1.3x on small reads) with a tenth of its memory. The services reach it at the same address (`ryax-minio:9000`) with the same credentials (`ryax-minio-secret`), and the upgrade copies the existing objects into it. ([roadmap#1471](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1471)) - **The `Authorization: Bearer <token>` header works on every service.** The Authorization and Repository services answered 401 to the form our own README documents, while Runner and Studio accepted it. All four now agree, and a header that is only `Bearer` answers 401 on Studio instead of 500. ([roadmap#1453](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1453)) - **Deploying a workflow that is already deploying returns 409 with its state.** `POST /workflows/{id}/deploy` answered 400 for this, the same code as for an invalid workflow, so clients had to match the error text. It now answers 409 `{"error": "Workflow can't be deployed at this status", "deployment_status": "Deploying"}`. An invalid workflow still gets 400. See the upgrade notes. ([roadmap#1467](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1467)) ## Bug fixes and Improvements - **An install reached through a proxy or by IP answers again.** 26.9.0 restricted the Ingresses to `global.tls.hostname`, so any other name got a 404. They are now restricted only to `global.ingress.hosts`, which is empty (any host) by default. The documentation gained a how-to for [running Ryax behind a proxy](https://docs.ryax.tech/). ([roadmap#1448](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1448)) - **Importing a workflow keeps its addons.** An imported workflow was valid and showed Deployed, but its HTTP services had no Ingress and every URL answered 404, because addons were dropped. Export now writes the editable addon values (`addons_inputs_values`), and import restores them; packages exported before this release still import and get the addons' defaults. A value for an addon parameter the action does not have now answers 400 instead of being silently dropped. ([roadmap#1464](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1464)) - **Importing a workflow with a numeric or directory input value no longer fails with a 500.** A value such as `max_tokens: 1024` or a directory input in a re-imported export now imports correctly, and a bad package answers 400. The deprecated `table` input/output type is removed; existing values are converted to `file`. ([roadmap#1465](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1465), [roadmap#1468](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1468)) - **An HTTP service trigger is no longer killed at startup.** Addon inputs and resources were frozen at the first deploy of an action and shared by every workflow using it, so a workflow using the HTTP addon on an action first deployed without it lost the addon's inputs. Each workflow action now owns its own, and two workflows can use one action with different addons and resources. Addon values are merged parameter by parameter: a value set in the interface beats the action's, which beats the addon default. ([roadmap#1463](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1463)) - **A workflow can be undeployed when it has several running deployments.** Undeploy used to fail and leave Studio in "Undeploying", with the portal and trigger views answering 500. It now stops all of them, and a stopped deployment is never brought back to running by a late trigger event. The leftover `PAUSED` state, which nothing has set since 26.7.0, is retired. ([roadmap#1337](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1337)) - **A failed deployment says why.** A deployment whose trigger failed (no site registered, trigger error, trigger cancelled by itself, deployment that could not even be created) stayed in "Deploying" forever or went back to "not deployed" with no explanation. It now fails and Studio shows `The trigger '<name>' failed: <reason>`. ([roadmap#1445](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1445), [roadmap#1466](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1466)) - **Executions that outlive their time allotment are cancelled.** An execution whose Worker crashed or never reported its end stayed running forever, holding its resources. The Runner now cancels any action execution still running 5 minutes past its allotment (`RYAX_EXECUTION_REAPER_GRACE_SECONDS`, checked every 60 seconds, `RYAX_EXECUTION_REAPER_INTERVAL_SECONDS`; a grace of 0 disables it). Triggers, which have no allotment, are not touched. ([roadmap#1299](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1299)) - **Stopping a workflow ends all its executions.** Trigger runs of an http trigger stayed RUNNING when the trigger ended gracefully or crashed, and executions stayed STOPPING when the Worker did not know them or had restarted. They now end as cancelled, and no longer enter the retry chain. Executions already stuck are not repaired and need a one-off cleanup. ([roadmap#1473](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1473)) - **Execution logs come out in the order the action wrote them.** The Workers now number each batch of log lines and the Runner reads them in that order. Logs stored before the upgrade keep their previous order. ([roadmap#1449](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1449)) - **Upgrades no longer deadlock on Grafana and MinIO.** Their `ReadWriteOnce` volume stopped the new pod from starting while the old one still held it, so `helm upgrade --wait` timed out and marked the release failed. Both now replace their pod instead of rolling it. ([roadmap#1451](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1451)) - **The registry garbage collector works again.** It pointed at a configuration path that does not exist since the registry moved to v3, and hid the failure behind a success message, so no image was ever collected. A failing run now fails its Job. ([roadmap#1294](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1294)) - **Security and dependencies.** The web interface's advisories are cleared (Angular 21.2.25, axios 1.20, undici 6.29, piscina 5.3, adm-zip 0.6.1, brace-expansion 2.1.6; the unpatched `braces` advisory only affects the build tooling and is tracked), and Python dependencies and nixpkgs (26.05) are refreshed across all services. Observability charts: kube-prometheus-stack 89.x (Grafana 13), Alloy 1.13, Traefik 41.6. - **Documentation.** New GPU node pools guide, rewritten IntelliScale reference, `global.ingress.hosts`, proxy configuration and the RabbitMQ operator in the install, airgap and ArgoCD guides. ## Upgrade to this version Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **⚠️ If you set `global.tls.hostname`, Ryax answers for every host again.** 26.9.0 restricted the Ingresses to that name; they are now restricted only to the names in `global.ingress.hosts`, which is empty (any host) by default. To keep Ryax to its own names, typically on a cluster shared with other applications, list them: ```yaml global: ingress: hosts: ["ryax.example.com"] ``` If you keep `global.tls.hostname` as well, it must be one of these hosts, or the chart refuses to render. Installs that set neither value need nothing. More in [Install Ryax with ArgoCD › Routing](https://docs.ryax.tech/howto/install_ryax_argocd/#routing). - **A new GitOps install needs one more secret.** With `global.secrets.create=false`, create `ryax-admin-credentials` (keys `admin-user` and `admin-password`) before the first sync, or set `authorization.adminUsername` and `authorization.adminPassword`. Without it the authorization pod stops with `No initial admin password configured`. An **existing** installation needs nothing — it never seeds again, and both `secretKeyRef`s are `optional`. The service also no longer has a built-in fallback for the initial user name (now `admin` when unset, not `user1`) nor for the JWT signing key: the chart has always injected the latter from `api-jwt-secret-key`, but a deployment that starts the services without the chart must now set `RYAX_JWT_SECRET_KEY`. - **⚠️ Apply the RabbitMQ CRD before upgrading.** The broker moves to the RabbitMQ Cluster Operator, whose `RabbitmqCluster` CRD ships in the chart, and Helm never installs a CRD on an upgrade. Without it the upgrade stops before changing anything, and prints this command: ```sh kubectl apply --server-side -f https://gitlab.com/ryax-tech/ryax/ryax-engine/-/raw/26.10.0/charts/ryax/subcharts/rabbitmq/crds/rabbitmqclusters.rabbitmq.com.yaml ``` Offline, take it from the chart package instead: ```sh tar -xzOf ryax-engine-26.10.0.tgz ryax-engine/charts/rabbitmq/crds/rabbitmqclusters.rabbitmq.com.yaml \ | kubectl apply --server-side -f - ``` `--server-side` is required: the CRD is too large for a client-side apply. ArgoCD applies the CRD itself, with the `ServerSideApply=true` the reference Application already sets. - **The broker is replaced, not upgraded, and starts empty.** Helm removes the Bitnami StatefulSet and the operator starts `ryax-broker-server-0` in its place, about a minute later. The services keep the same address (`ryax-broker:5672`) and the same credentials (`ryax-broker-secret`), and reconnect by themselves. There is nothing to migrate: Ryax publishes its messages as transient, so a broker restart has always dropped whatever was still queued. As for any upgrade, run it while no workflow is running. The services retry on their own during the switch, which took about a minute on a test cluster. Once Ryax is back, delete the old broker volume, if there is one (there is none with `rabbitmq.persistence.enabled: false`, as in `minimal.yaml`): ```sh kubectl -n ryaxns delete pvc data-ryax-broker-0 ``` - **The `rabbitmq:` values now configure the new broker.** `persistence.*`, `resources`, `tolerations`, `nodeSelector`, `affinity`, `priorityClassName` and `metrics.enabled` keep their meaning. Every other Bitnami key (`auth.*`, `image.*`, `clustering`, `plugins`, ...) is ignored: remove them. A `rabbitmq.image.repository` still naming a Bitnami image stops the render. The broker password is always the one in `ryax-broker-secret`, so a `rabbitmq.auth.password` set by the old troubleshooting guide does nothing. The broker also follows `global.tolerations`, `nodeSelector` and `affinity` now. - **Uninstalling now takes one more step.** `helm uninstall` removes the operator at the same time as the broker, so nothing clears the `RabbitmqCluster` finalizer: the broker pod keeps running, and `helm uninstall --wait` times out. Delete the broker first, while the operator still runs: ```sh kubectl -n ryaxns delete rabbitmqcluster ryax-broker helm uninstall ryax -n ryaxns ``` If an uninstall is already stuck, clear the finalizer: `kubectl -n ryaxns patch rabbitmqcluster ryax-broker --type merge -p '{"metadata":{"finalizers":[]}}'`. - **If the cluster already runs a RabbitMQ Cluster Operator** that watches the Ryax namespace, set `rabbitmq.operator.enabled: false`. Two operators would reconcile the same broker. - **Multi-site with Skupper:** the `ryax-broker-ext` and `ryax-minio-ext` connectors copied the pod selectors of the old broker and of MinIO when they were created. After the upgrade the first matches no pod, and the second matches the old MinIO, which refuses their connections. Recreate both on the main site right after the upgrade, before running workflows on remote sites: ```sh skupper -n ryaxns connector delete ryax-broker-ext skupper -n ryaxns connector create ryax-broker-ext 5672 --workload service/ryax-broker skupper -n ryaxns connector delete ryax-minio-ext skupper -n ryaxns connector create ryax-minio-ext 9000 --workload service/ryax-minio ``` - **GitOps (`global.secrets.create=false`):** `ryax-broker-cookie` is no longer read, delete it whenever you like. `ryax-broker-secret` keeps its keys; its `broker-user` and `rabbitmq-password` now seed the broker's user, so they must match the `broker` URL as before. With automated pruning off, prune the old broker resources in the same sync: the operator cannot create its `ryax-broker` Service while the Bitnami one is still there. - **⚠️ The filestore moves from MinIO to versitygw, and the upgrade copies its objects.** Helm keeps the old MinIO, `ryax-minio`, running on its volume, and marks that volume so that neither Helm nor ArgoCD ever deletes it. The new filestore, `ryax-filestore`, gets a volume of its own, the size of MinIO's. Before it starts serving, its pod copies every object out of MinIO, checks that both sides list the same objects and sizes, and leaves a marker so that it never copies again. The address and the credentials do not change. **Downtime:** the services cannot reach the filestore until the copy is over. Without a pre-copy (below), count the time to read the whole MinIO volume once: on a test cluster, 750 MiB in 3,000 objects took a few seconds, and a large volume on network storage takes minutes. On a local k3s upgrade of a 26.9.0 install holding 58 MiB in 410 objects, the copy itself took a second, the filestore was unreachable for 15 to 50 seconds while MinIO restarted, and the whole of Ryax was back within 3 minutes, the broker switch being the longest part. The runner and studio restart while they wait, and their restart back-off can add up to five minutes once the copy is over. They reconnect on their own; to skip the back-off, restart them once the filestore is Ready: ```sh kubectl -n ryaxns rollout status deploy/ryax-filestore --timeout=24h kubectl -n ryaxns rollout restart deploy/ryax-runner deploy/ryax-studio ``` Follow the copy with `kubectl -n ryaxns logs -f deploy/ryax-filestore -c migrate-from-minio`. The new volume's storage class must support user extended attributes, as ext4 and xfs do. If MinIO ran without persistence (`minio.persistence.enabled: false`), it has no volume to copy from, and its objects go away with it, as they did whenever its pod restarted. The `minio:` values are ignored now: remove them, but carry a `minio.persistence.storageClass` over as `filestore.migration.legacy.persistence.storageClass`, as the class of MinIO's volume cannot change. To give the filestore more room than MinIO had, set `filestore.persistence.size`. - **Optional: pre-copy the objects to shorten the upgrade.** While 26.9.0 still runs, this copies the objects into the volume the filestore will use, and the upgrade then only copies what changed since. It can run several times, and it can be interrupted: the copy at the upgrade makes the result exact whatever it left. Use the same values file as for the upgrade: ```sh SIZE=$(kubectl -n ryaxns get pvc ryax-minio -o jsonpath='{.spec.resources.requests.storage}') helm template ryax oci://registry.ryax.org/release-charts/ryax-engine --version 26.10.0 \ -n ryaxns -f values.yaml \ --set filestore.migration.precopy.enabled=true \ --set filestore.persistence.size="$SIZE" \ --show-only charts/filestore/templates/precopy.yaml > precopy.yaml kubectl -n ryaxns delete job ryax-filestore-precopy --ignore-not-found kubectl -n ryaxns apply -f precopy.yaml kubectl -n ryaxns wait --for=condition=complete job/ryax-filestore-precopy --timeout=24h kubectl -n ryaxns logs job/ryax-filestore-precopy | tail -n 1 kubectl -n ryaxns delete job ryax-filestore-precopy ``` It creates the `ryax-filestore` volume, which the upgrade then takes over, and a Job that reads MinIO while Ryax keeps working. The upgrade refuses to start while that Job runs. Delete the Job once it is done, as above: as long as it exists, the volume cannot be deleted, which a rollback or an uninstall does. The pre-copy only shortens the upgrade when the volume is large: on the test cluster above, both copies took about a second. Offline, render it from the chart package (`helm template ryax ryax-engine-26.10.0.tgz ...`): it only uses the MinIO image 26.9.0 already ran. - **Once 26.10.0 runs, decommission the old MinIO.** It only serves the copy, and makes a rollback possible until then. First check that the copy is done: ```sh kubectl -n ryaxns exec deploy/ryax-filestore -c versitygw -- cat /data/.ryax-migrated-from-minio ``` It prints the date of the copy. Then remove MinIO, and delete its volume: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.10.0 \ -n ryaxns --reuse-values --set filestore.migration.enabled=false kubectl -n ryaxns delete pvc ryax-minio ``` The chart refuses to remove MinIO while no filestore pod has finished the copy (`filestore.migration.allowDecommissionWithoutCopy=true` forces it), and refuses `filestore.migration.enabled=false` on an upgrade from 26.9.0, which would delete MinIO's volume before the copy. Like any upgrade, the decommission restarts the runner and studio, and it restarts the filestore, unreachable for a few seconds. **After the decommission, a rollback to 26.9.0 is no longer possible** without restoring MinIO's volume from a backup. - **Rolling back to 26.9.0 before the decommission** brings MinIO back on its volume, untouched, and deletes the filestore and its volume. Objects written since the upgrade are lost, unless you copy them back into MinIO first, while 26.10.0 still runs and no workflow runs. That copy rewrites every object, not only the new ones, so it takes about as long as the copy at the upgrade. Then delete the broker while its operator still runs, for the Bitnami broker cannot take its `ryax-broker` Service back otherwise, and roll back: ```sh kubectl -n ryaxns exec deploy/ryax-minio -- sh -c 'export HOME=/tmp mc=/opt/bitnami/minio-client/bin/mc $mc alias set new http://ryax-minio:9000 "$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD" >/dev/null $mc alias set old http://localhost:9000 "$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD" >/dev/null $mc mirror --overwrite new/ryax-filestore old/ryax-filestore' kubectl -n ryaxns delete rabbitmqcluster ryax-broker helm rollback ryax <the 26.9.0 revision> -n ryaxns ``` Without the `delete rabbitmqcluster`, the rollback stops part-way on `no Service with the name "ryax-broker" found`. Do not run a failed rollback again: each attempt leaves more of both versions behind. Upgrade to 26.10.0 again with the same values file, wait for every pod to be Ready, and roll back as above. Upgrading again later copies MinIO afresh. A `ryax-worker-k8s` already on 26.10.0 can stay there: on the test cluster it ran workflows with the rolled-back engine. - **GitOps (ArgoCD, Flux):** the chart cannot see the cluster, so it always renders the old MinIO and its volume while `filestore.migration.enabled` is on, the default. The upgrade copies as above. If you changed `minio.persistence.size` or `storageClass`, set them as `filestore.migration.legacy.persistence.size` and `storageClass`, or the sync fails on the volume, which cannot shrink. To decommission, check the marker as above, then set `filestore.migration.enabled: false`: ArgoCD removes MinIO but keeps its volume (`Prune=false`), which you delete by hand. **A new GitOps install** has nothing to migrate: set `filestore.migration.enabled: false` from the start. - **`global.security.allowInsecureImages` is gone.** It only served the Bitnami charts, and the engine chart no longer has any. Remove it from your values; it is ignored. - **Grafana 13 (kube-prometheus-stack 89).** The bundled Grafana runs the `-distroless` image with a read-only root filesystem and an `emptyDir` on `/tmp`. `GF_*__FILE` environment variables and `GF_INSTALL_PLUGINS` are no longer supported. The Ryax values set none of them, and the plugins (`grafana-piechart-panel`, `grafana-clock-panel`, `vonage-status-panel`) are now preinstalled synchronously at startup, so no values change is needed, but check that the Grafana pod starts. If your own overrides set `GF_*__FILE`, `GF_INSTALL_PLUGINS` or an `extraVolumeMounts` on `/tmp`, migrate them first. The rollout also replaces the Grafana and MinIO pods instead of rolling them, so both are briefly unavailable during the upgrade. - **Kubernetes worker: the MIG auto-labeler is removed.** The values `config.MIG` and `labeler` are gone, along with the node-labeler DaemonSet. If you relied on it to turn a `gpu-pool-mig-*` node label into `nvidia.com/mig.config`, label the GPU nodes with `nvidia.com/mig.config` yourself (cloud node-pool labels are the usual way), as the NVIDIA GPU Operator's MIG Manager expects. Remove both keys from your worker values. To keep actions off GPU nodes that are not ready yet, see the new opt-in `gpuReadiness` values: enable it **before** adding the startup taint to the node pool, as a tainted pool with the gate off never runs anything. - **GPU node pools: record the card, and mind the renamed field.** Existing pools keep working without it, but record the catalog `gpu_model` on each GPU pool (web interface, or `GET /api/runner/gpu-models` then the node pool API) to get precise matching and per-hardware recommendations. If your clusters are all A30 (4 compute slices), also set `intelliscale.config.algorithm_configs.simple_mig_recommender.total_compute_slices` to 4 for pools with no recorded model. Placement now defaults to `best_fit`; to keep the previous behaviour set `RYAX_SCHEDULER_GPU_FIT_POLICY=first_fit` through `runner.extraEnv`. - **API users:** - `POST /workflows/{id}/deploy` returns **409** (with `deployment_status` in the body) instead of 400 when the workflow is already deploying, deployed or undeploying. A client that matched the 400 text must handle 409 instead. - Node pools take and return `gpu_count`, not `gpu` (the stored value is migrated). A client that still sends `gpu` when creating a node pool gets a 422. - New root endpoints `/api/runner/node-pools` (filters on GPU, site, CPU and memory; the GPU is named by `gpu_config_id`) and `GET /api/runner/gpu-models` returns a flat list. The nested `/sites/{id}/node-pools` routes remain, marked deprecated. - `GET /modules/{id}` and workflow action views now return `gpu_brand`, `gpu_model`, `gpu_memory_gb` and `gpu_compute_fraction` in `resources`. - The `table` input/output type no longer exists. - `Authorization: Bearer <token>` is accepted by all services. - **Registry garbage collection now actually runs.** The first scheduled run after the upgrade removes every untagged image left in the registry, and a failing run now shows as a failed Job. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.10.0 \ -n ryaxns \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.10.0 \ -n ryaxns \ -f worker.yaml ``` -
26.10.0-rc0
be4d62b9 · ·Ryax 26.10.0-rc0 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ # Ryax 26.10.0 ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ GPU requests by model, memory and share of a card, a random initial admin password, a broker that no longer depends on Bitnami, and a round of fixes for imports, stuck executions and failed deployments. ## New features ### GPU scheduling by model, memory and share of a card Ryax used to know a GPU only as a count, plus an opaque MIG profile name that had to match exactly. This release describes GPUs the way users think about them, from the action that asks to the node pool that answers. - **Actions can ask for a kind of GPU.** In `ryax_metadata.yaml`, `spec.resources.gpu` still takes a number, and now also takes an object with a `count`, a `brand`, a `model`, a `memory` (with a unit, for example `40G`) and a `compute_fraction` (a share of one card, between 0 and 1, exclusive of 0). Everything in it is optional and describes **one** GPU: `count: 2` with `memory: 40G` means two cards of 40 GB each, not 80 GB across two. `gpu: 2` keeps its meaning of two GPUs of any kind. Brand and model must match the catalog; an action that names one no node pool has simply finds no place to run. A request with an unknown key, or a memory without a unit, is rejected when the action is scanned. - **Node pools declare their card.** A GPU node pool now records its catalog model next to its partition (`full` or a MIG profile), so Ryax knows how much memory and compute one GPU of the pool really gives. The new `GET /api/runner/gpu-models?search=a100 1g` lists the catalog as one flat, searchable list (brand, model, partition, memory, compute share), and the node pool form in the web interface now picks the GPU with one searchable select. Existing node pools keep working: one with no recorded model stays eligible for every request, and is simply treated as of unknown size, never as too small. Fill in the model to get precise matching. - **Placement takes any partition that is big enough, and the tightest one by default.** A recommendation of "3g.40gb" used to demand exactly that profile and fail when no pool had it; now any partition that meets the request will do. New setting `RYAX_SCHEDULER_GPU_FIT_POLICY` on the Runner: `best_fit` (the default) picks the pool that wastes the least of a card, leaving large partitions free for the actions that need them; `first_fit` keeps the previous behaviour of taking the best-scoring pool that fits. GPU requests are also now checked against the pool's CPU, memory and time, which they never were. - **IntelliScale recommends a share of a card, not a MIG profile.** It now recommends GPU memory (the observed peak plus 10%) and a compute share, and Ryax rounds them up to a partition your pools actually offer. This also works on cards that do not split in sevenths such as the A30. - **Recommendations are learned per piece of hardware, not per site.** A recommendation is only valid for the machine it was measured on, so IntelliScale now keeps one model per GPU model (for the GPU share) and per instance type (for CPU and memory). Two node pools of one site with different cards no longer contaminate each other, and two sites with the same card now learn together. The Runner asks for the recommendation after it has picked a candidate pool, one answer per pool, instead of merging every site's answer beforehand. Nothing fragments on upgrade: pools with no recorded GPU model, and HPC sites, share one model as before. - **Autoscaled GPU nodes can be kept out of use until they are ready.** An autoscaled GPU node is reported `Ready` long before its driver and MIG layout are in place, so actions landing on it got a whole GPU instead of their slice. The Kubernetes worker chart can now hold new GPU nodes behind a startup taint and release them only when the GPU stack is usable and the MIG layout is the one the pool asked for. Off by default: see `gpuReadiness` and the new [GPU node pools and MIG](https://docs.ryax.tech/howto/gpu_node_pools/) guide. ([roadmap#1421](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1421), [roadmap#1458](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1458), [roadmap#1469](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1469), [roadmap#1428](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1428)) ### Security and operations - **The initial admin password is random.** Every installation used to boot with the same `user1` / `pass1` — hard-coded in the service, never set by the chart, and published in the README and the install guide. The chart now generates one per installation into the `ryax-admin-credentials` secret, and `helm install` prints the command to read it back. The account is `admin`. Existing installations keep their users and passwords: the secret is only ever read when the user table is empty, which is the very first start. The Edit password window now warns that changing the password does not update that secret. ([roadmap#1423](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1423)) - **`helm install` prints its notes.** `NOTES.txt` had always sat at the chart root rather than in `templates/`, where Helm is the only place it looks, so no install had ever printed anything. It now carries the admin credentials command and the Grafana one. - **The broker no longer runs on a Bitnami image.** It ran on `bitnamilegacy/rabbitmq`, which Bitnami no longer updates. It is now a `RabbitmqCluster` run by the official [RabbitMQ Cluster Operator](https://www.rabbitmq.com/kubernetes/operator/operator-overview) (2.23.0) on the official `rabbitmq:4.3.6-management` image. The chart deploys the operator in the Ryax namespace, watching that namespace only, and it needs neither cluster-wide rights nor cert-manager. The services reach the broker at the same address with the same credentials. **This needs one command before upgrading, see below.** ([roadmap#1320](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1320)) - **The `Authorization: Bearer <token>` header works on every service.** The Authorization and Repository services answered 401 to the form our own README documents, while Runner and Studio accepted it. All four now agree, and a header that is only `Bearer` answers 401 on Studio instead of 500. ([roadmap#1453](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1453)) - **Deploying a workflow that is already deploying returns 409 with its state.** `POST /workflows/{id}/deploy` answered 400 for this, the same code as for an invalid workflow, so clients had to match the error text. It now answers 409 `{"error": "Workflow can't be deployed at this status", "deployment_status": "Deploying"}`. An invalid workflow still gets 400. See the upgrade notes. ([roadmap#1467](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1467)) ## Bug fixes and Improvements - **An install reached through a proxy or by IP answers again.** 26.9.0 restricted the Ingresses to `global.tls.hostname`, so any other name got a 404. They are now restricted only to `global.ingress.hosts`, which is empty (any host) by default. The documentation gained a how-to for [running Ryax behind a proxy](https://docs.ryax.tech/). ([roadmap#1448](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1448)) - **Importing a workflow keeps its addons.** An imported workflow was valid and showed Deployed, but its HTTP services had no Ingress and every URL answered 404, because addons were dropped. Export now writes the editable addon values (`addons_inputs_values`), and import restores them; packages exported before this release still import and get the addons' defaults. A value for an addon parameter the action does not have now answers 400 instead of being silently dropped. ([roadmap#1464](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1464)) - **Importing a workflow with a numeric or directory input value no longer fails with a 500.** A value such as `max_tokens: 1024` or a directory input in a re-imported export now imports correctly, and a bad package answers 400. The deprecated `table` input/output type is removed; existing values are converted to `file`. ([roadmap#1465](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1465), [roadmap#1468](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1468)) - **An HTTP service trigger is no longer killed at startup.** Addon inputs and resources were frozen at the first deploy of an action and shared by every workflow using it, so a workflow using the HTTP addon on an action first deployed without it lost the addon's inputs. Each workflow action now owns its own, and two workflows can use one action with different addons and resources. Addon values are merged parameter by parameter: a value set in the interface beats the action's, which beats the addon default. ([roadmap#1463](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1463)) - **A workflow can be undeployed when it has several running deployments.** Undeploy used to fail and leave Studio in "Undeploying", with the portal and trigger views answering 500. It now stops all of them, and a stopped deployment is never brought back to running by a late trigger event. The leftover `PAUSED` state, which nothing has set since 26.7.0, is retired. ([roadmap#1337](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1337)) - **A failed deployment says why.** A deployment whose trigger failed (no site registered, trigger error, trigger cancelled by itself, deployment that could not even be created) stayed in "Deploying" forever or went back to "not deployed" with no explanation. It now fails and Studio shows `The trigger '<name>' failed: <reason>`. ([roadmap#1445](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1445), [roadmap#1466](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1466)) - **Executions that outlive their time allotment are cancelled.** An execution whose Worker crashed or never reported its end stayed running forever, holding its resources. The Runner now cancels any action execution still running 5 minutes past its allotment (`RYAX_EXECUTION_REAPER_GRACE_SECONDS`, checked every 60 seconds, `RYAX_EXECUTION_REAPER_INTERVAL_SECONDS`; a grace of 0 disables it). Triggers, which have no allotment, are not touched. ([roadmap#1299](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1299)) - **Stopping a workflow ends all its executions.** Trigger runs of an http trigger stayed RUNNING when the trigger ended gracefully or crashed, and executions stayed STOPPING when the Worker did not know them or had restarted. They now end as cancelled, and no longer enter the retry chain. Executions already stuck are not repaired and need a one-off cleanup. ([roadmap#1473](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1473)) - **Execution logs come out in the order the action wrote them.** The Workers now number each batch of log lines and the Runner reads them in that order. Logs stored before the upgrade keep their previous order. ([roadmap#1449](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1449)) - **Upgrades no longer deadlock on Grafana and MinIO.** Their `ReadWriteOnce` volume stopped the new pod from starting while the old one still held it, so `helm upgrade --wait` timed out and marked the release failed. Both now replace their pod instead of rolling it. ([roadmap#1451](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1451)) - **The registry garbage collector works again.** It pointed at a configuration path that does not exist since the registry moved to v3, and hid the failure behind a success message, so no image was ever collected. A failing run now fails its Job. ([roadmap#1294](https://gitlab.com/ryax-tech/ryax/roadmap/-/issues/1294)) - **Security and dependencies.** The web interface's advisories are cleared (Angular 21.2.25, axios 1.20, undici 6.29, piscina 5.3, adm-zip 0.6.1, brace-expansion 2.1.6; the unpatched `braces` advisory only affects the build tooling and is tracked), and Python dependencies and nixpkgs (26.05) are refreshed across all services. Observability charts: kube-prometheus-stack 89.x (Grafana 13), Alloy 1.13, Traefik 41.6. - **Documentation.** New GPU node pools guide, rewritten IntelliScale reference, `global.ingress.hosts`, proxy configuration and the RabbitMQ operator in the install, airgap and ArgoCD guides. ## Upgrade to this version Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **⚠️ If you set `global.tls.hostname`, Ryax answers for every host again.** 26.9.0 restricted the Ingresses to that name; they are now restricted only to the names in `global.ingress.hosts`, which is empty (any host) by default. To keep Ryax to its own names, typically on a cluster shared with other applications, list them: ```yaml global: ingress: hosts: ["ryax.example.com"] ``` If you keep `global.tls.hostname` as well, it must be one of these hosts, or the chart refuses to render. Installs that set neither value need nothing. More in [Install Ryax with ArgoCD › Routing](https://docs.ryax.tech/howto/install_ryax_argocd/#routing). - **A new GitOps install needs one more secret.** With `global.secrets.create=false`, create `ryax-admin-credentials` (keys `admin-user` and `admin-password`) before the first sync, or set `authorization.adminUsername` and `authorization.adminPassword`. Without it the authorization pod stops with `No initial admin password configured`. An **existing** installation needs nothing — it never seeds again, and both `secretKeyRef`s are `optional`. The service also no longer has a built-in fallback for the initial user name (now `admin` when unset, not `user1`) nor for the JWT signing key: the chart has always injected the latter from `api-jwt-secret-key`, but a deployment that starts the services without the chart must now set `RYAX_JWT_SECRET_KEY`. - **⚠️ Apply the RabbitMQ CRD before upgrading.** The broker moves to the RabbitMQ Cluster Operator, whose `RabbitmqCluster` CRD ships in the chart, and Helm never installs a CRD on an upgrade. Without it the upgrade stops before changing anything, and prints this command: ```sh kubectl apply --server-side -f https://gitlab.com/ryax-tech/ryax/ryax-engine/-/raw/26.10.0/charts/ryax/subcharts/rabbitmq/crds/rabbitmqclusters.rabbitmq.com.yaml ``` Offline, take it from the chart package instead: ```sh tar -xzOf ryax-engine-26.10.0.tgz ryax-engine/charts/rabbitmq/crds/rabbitmqclusters.rabbitmq.com.yaml \ | kubectl apply --server-side -f - ``` `--server-side` is required: the CRD is too large for a client-side apply. ArgoCD applies the CRD itself, with the `ServerSideApply=true` the reference Application already sets. - **The broker is replaced, not upgraded, and starts empty.** Helm removes the Bitnami StatefulSet and the operator starts `ryax-broker-server-0` in its place, about a minute later. The services keep the same address (`ryax-broker:5672`) and the same credentials (`ryax-broker-secret`), and reconnect by themselves. There is nothing to migrate: Ryax publishes its messages as transient, so a broker restart has always dropped whatever was still queued. As for any upgrade, run it while no workflow is running. The services retry on their own during the switch, which took about a minute on a test cluster. Once Ryax is back, delete the old broker volume, if there is one (there is none with `rabbitmq.persistence.enabled: false`, as in `minimal.yaml`): ```sh kubectl -n ryaxns delete pvc data-ryax-broker-0 ``` - **The `rabbitmq:` values now configure the new broker.** `persistence.*`, `resources`, `tolerations`, `nodeSelector`, `affinity`, `priorityClassName` and `metrics.enabled` keep their meaning. Every other Bitnami key (`auth.*`, `image.*`, `clustering`, `plugins`, ...) is ignored: remove them. A `rabbitmq.image.repository` still naming a Bitnami image stops the render. The broker password is always the one in `ryax-broker-secret`, so a `rabbitmq.auth.password` set by the old troubleshooting guide does nothing. The broker also follows `global.tolerations`, `nodeSelector` and `affinity` now. - **Uninstalling now takes one more step.** `helm uninstall` removes the operator at the same time as the broker, so nothing clears the `RabbitmqCluster` finalizer: the broker pod keeps running, and `helm uninstall --wait` times out. Delete the broker first, while the operator still runs: ```sh kubectl -n ryaxns delete rabbitmqcluster ryax-broker helm uninstall ryax -n ryaxns ``` If an uninstall is already stuck, clear the finalizer: `kubectl -n ryaxns patch rabbitmqcluster ryax-broker --type merge -p '{"metadata":{"finalizers":[]}}'`. - **If the cluster already runs a RabbitMQ Cluster Operator** that watches the Ryax namespace, set `rabbitmq.operator.enabled: false`. Two operators would reconcile the same broker. - **Multi-site with Skupper:** the `ryax-broker-ext` connector copied the old broker's pod selector when it was created, so it matches no pod after the upgrade. Recreate it on the main site: ```sh skupper -n ryaxns connector delete ryax-broker-ext skupper -n ryaxns connector create ryax-broker-ext 5672 --workload service/ryax-broker ``` - **GitOps (`global.secrets.create=false`):** `ryax-broker-cookie` is no longer read, delete it whenever you like. `ryax-broker-secret` keeps its keys; its `broker-user` and `rabbitmq-password` now seed the broker's user, so they must match the `broker` URL as before. With automated pruning off, prune the old broker resources in the same sync: the operator cannot create its `ryax-broker` Service while the Bitnami one is still there. - **Grafana 13 (kube-prometheus-stack 89).** The bundled Grafana runs the `-distroless` image with a read-only root filesystem and an `emptyDir` on `/tmp`. `GF_*__FILE` environment variables and `GF_INSTALL_PLUGINS` are no longer supported. The Ryax values set none of them, and the plugins (`grafana-piechart-panel`, `grafana-clock-panel`, `vonage-status-panel`) are now preinstalled synchronously at startup, so no values change is needed, but check that the Grafana pod starts. If your own overrides set `GF_*__FILE`, `GF_INSTALL_PLUGINS` or an `extraVolumeMounts` on `/tmp`, migrate them first. The rollout also replaces the Grafana and MinIO pods instead of rolling them, so both are briefly unavailable during the upgrade. - **Kubernetes worker: the MIG auto-labeler is removed.** The values `config.MIG` and `labeler` are gone, along with the node-labeler DaemonSet. If you relied on it to turn a `gpu-pool-mig-*` node label into `nvidia.com/mig.config`, label the GPU nodes with `nvidia.com/mig.config` yourself (cloud node-pool labels are the usual way), as the NVIDIA GPU Operator's MIG Manager expects. Remove both keys from your worker values. To keep actions off GPU nodes that are not ready yet, see the new opt-in `gpuReadiness` values: enable it **before** adding the startup taint to the node pool, as a tainted pool with the gate off never runs anything. - **GPU node pools: record the card, and mind the renamed field.** Existing pools keep working without it, but record the catalog `gpu_model` on each GPU pool (web interface, or `GET /api/runner/gpu-models` then the node pool API) to get precise matching and per-hardware recommendations. If your clusters are all A30 (4 compute slices), also set `intelliscale.config.algorithm_configs.simple_mig_recommender.total_compute_slices` to 4 for pools with no recorded model. Placement now defaults to `best_fit`; to keep the previous behaviour set `RYAX_SCHEDULER_GPU_FIT_POLICY=first_fit` through `runner.extraEnv`. - **API users:** - `POST /workflows/{id}/deploy` returns **409** (with `deployment_status` in the body) instead of 400 when the workflow is already deploying, deployed or undeploying. A client that matched the 400 text must handle 409 instead. - Node pools take and return `gpu_count`, not `gpu` (the stored value is migrated). A client that still sends `gpu` when creating a node pool gets a 422. - New root endpoints `/api/runner/node-pools` (filters on GPU, site, CPU and memory; the GPU is named by `gpu_config_id`) and `GET /api/runner/gpu-models` returns a flat list. The nested `/sites/{id}/node-pools` routes remain, marked deprecated. - `GET /modules/{id}` and workflow action views now return `gpu_brand`, `gpu_model`, `gpu_memory_gb` and `gpu_compute_fraction` in `resources`. - The `table` input/output type no longer exists. - `Authorization: Bearer <token>` is accepted by all services. - **Registry garbage collection now actually runs.** The first scheduled run after the upgrade removes every untagged image left in the registry, and a failing run now shows as a failed Job. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.10.0 \ -n ryaxns \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.10.0 \ -n ryaxns \ -f worker.yaml ``` -
26.9.0
Release: 26.9.0f605cdd7 · ·Ryax 26.9.0 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ # Ryax 26.9.0 ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ Stability and security updates, plus a GitOps-ready Helm chart and a rebuilt web interface. ## New features - **GitOps installs.** The chart renders without a cluster connection, so ArgoCD and Flux no longer mint fresh credentials on every sync. Set `global.secrets.create=false` to supply the secrets yourself. - **Pick the Ingress controller.** Ryax's Ingresses name their class through `global.ingress.className`, overridable per service or switched off with `<subchart>.ingress.enabled=false`. The bundled Traefik no longer registers itself as the cluster-wide default class. - **Pod placement.** `global.tolerations` and `global.affinity` apply to every Ryax pod, with per-subchart overrides; the worker charts gained `tolerations` and `nodeSelector`. - **Per-site action registry.** A worker whose nodes cannot resolve the registry the Runner recorded overrides it with `internalRegistryOverride`. - **Rebuilt web interface.** Angular 16 → 21, Nx 22, TypeScript 5.9. - **Generated API reference.** <https://docs.ryax.tech/reference/api/> is now built from the service sources on every release, so it cannot drift again — the 26.7.0 document still described 26.2.0. ## Bug fixes and Improvements - Action builds no longer get stuck in "Starting". When the builder reported back before the queue had committed the action's status, the move to "Building" was lost, and the successful build that followed was refused as well — leaving the action in "Starting" for good and, since builds run one at a time, blocking every action queued behind it. The Library now also offers Cancel Build while an action is "Starting" or "Cancelling", so a stalled build can be cleared by hand. - Prometheus keeps its metrics across restarts: the volume request sat one level too high in the values and was silently ignored, leaving it on an `emptyDir`. - Kubernetes worker database upgrades work again — the PostgreSQL service name and the database URL now agree, so the migration init container resolves its host. - IntelliScale no longer restarts on unrelated configuration changes. - The Runner and Repository APIs moved to FastAPI and pydantic; Studio dropped marshmallow; the Authorization service is now part of the core service. - The V1 worker protocol and the legacy worker module are removed. - Container image publishing is reproducible again, fixing intermittent `Digest did not match` failures. - Security fixes across the stack, including `fast-uri` 3.1.6 and the `js-yaml`, `svgo` and `extract-zip` advisories, plus a repaired image CVE scan. - Observability dependencies updated: kube-prometheus-stack 88.x, Loki 7.3, Alloy 1.12, Traefik 41.5. ## Upgrade to this version Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **`--take-ownership` is required when upgrading from 26.7.0.** Three credential secrets used to be created as Helm *hook* resources, which Helm never records as part of the release: the Studio password encryption key in the main chart, and the PostgreSQL credentials in both worker charts. They are now ordinary chart-managed resources, so Helm finds them un-owned and refuses the upgrade with `invalid ownership metadata` unless you let it adopt them. Adoption preserves the existing values -- the templates read the current secret before falling back -- so the encryption key and the database passwords are unchanged. - **Prometheus storage:** if your values set `kube-prometheus-stack.prometheus.storage.volumeClaimTemplate`, move it to `prometheus.prometheusSpec.storageSpec` and add `accessModes: ["ReadWriteOnce"]`. The old key was never read. There is no metric history to preserve, since Prometheus was running on an `emptyDir`. - **Traefik:** the bundled instance no longer claims Ingresses that name no class. Name `<release-name>-traefik` on your own Ingresses, or mark your controller as the cluster default. - **Workers relying on the old `internalRegistryOverride` default:** on the SLURM_SSH chart it changed from `ryax-registry:5000` to empty. If you never set it yourself, set it explicitly to keep pulling through the in-cluster registry. On a Kubernetes worker whose nodes cannot resolve the Runner's address, use `127.0.0.1:30012`. - **API users:** the Repository V1 endpoints `/api/repository/modules` and `/api/repository/modules/{module_id}` are removed; use `/api/repository/v2/`. - **Still on the pre-26.7.0 `ryax-worker` chart:** migrate to `ryax-worker-k8s` or `ryax-worker-slurm-ssh` first, as the V1 worker protocol is gone. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.9.0 \ -n ryaxns \ --take-ownership \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.9.0 \ -n ryaxns \ --take-ownership \ -f worker.yaml ``` -
26.9.0-rc4
3e6a9251 · ·Ryax 26.9.0-rc4 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ # Ryax 26.9.0 ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ Stability and security updates, plus a GitOps-ready Helm chart and a rebuilt web interface. ## New features - **GitOps installs.** The chart renders without a cluster connection, so ArgoCD and Flux no longer mint fresh credentials on every sync. Set `global.secrets.create=false` to supply the secrets yourself. - **Pick the Ingress controller.** Ryax's Ingresses name their class through `global.ingress.className`, overridable per service or switched off with `<subchart>.ingress.enabled=false`. The bundled Traefik no longer registers itself as the cluster-wide default class. - **Pod placement.** `global.tolerations` and `global.affinity` apply to every Ryax pod, with per-subchart overrides; the worker charts gained `tolerations` and `nodeSelector`. - **Per-site action registry.** A worker whose nodes cannot resolve the registry the Runner recorded overrides it with `internalRegistryOverride`. - **Rebuilt web interface.** Angular 16 → 21, Nx 22, TypeScript 5.9. - **Generated API reference.** <https://docs.ryax.tech/reference/api/> is now built from the service sources on every release, so it cannot drift again — the 26.7.0 document still described 26.2.0. ## Bug fixes and Improvements - Action builds no longer get stuck in "Starting". When the builder reported back before the queue had committed the action's status, the move to "Building" was lost, and the successful build that followed was refused as well — leaving the action in "Starting" for good and, since builds run one at a time, blocking every action queued behind it. The Library now also offers Cancel Build while an action is "Starting" or "Cancelling", so a stalled build can be cleared by hand. - Prometheus keeps its metrics across restarts: the volume request sat one level too high in the values and was silently ignored, leaving it on an `emptyDir`. - Kubernetes worker database upgrades work again — the PostgreSQL service name and the database URL now agree, so the migration init container resolves its host. - IntelliScale no longer restarts on unrelated configuration changes. - The Runner and Repository APIs moved to FastAPI and pydantic; Studio dropped marshmallow; the Authorization service is now part of the core service. - The V1 worker protocol and the legacy worker module are removed. - Container image publishing is reproducible again, fixing intermittent `Digest did not match` failures. - Security fixes across the stack, including `fast-uri` 3.1.6 and the `js-yaml`, `svgo` and `extract-zip` advisories, plus a repaired image CVE scan. - Observability dependencies updated: kube-prometheus-stack 88.x, Loki 7.3, Alloy 1.12, Traefik 41.5. ## Upgrade to this version Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **`--take-ownership` is required when upgrading from 26.7.0.** Three credential secrets used to be created as Helm *hook* resources, which Helm never records as part of the release: the Studio password encryption key in the main chart, and the PostgreSQL credentials in both worker charts. They are now ordinary chart-managed resources, so Helm finds them un-owned and refuses the upgrade with `invalid ownership metadata` unless you let it adopt them. Adoption preserves the existing values -- the templates read the current secret before falling back -- so the encryption key and the database passwords are unchanged. - **Prometheus storage:** if your values set `kube-prometheus-stack.prometheus.storage.volumeClaimTemplate`, move it to `prometheus.prometheusSpec.storageSpec` and add `accessModes: ["ReadWriteOnce"]`. The old key was never read. There is no metric history to preserve, since Prometheus was running on an `emptyDir`. - **Traefik:** the bundled instance no longer claims Ingresses that name no class. Name `<release-name>-traefik` on your own Ingresses, or mark your controller as the cluster default. - **Workers relying on the old `internalRegistryOverride` default:** on the SLURM_SSH chart it changed from `ryax-registry:5000` to empty. If you never set it yourself, set it explicitly to keep pulling through the in-cluster registry. On a Kubernetes worker whose nodes cannot resolve the Runner's address, use `127.0.0.1:30012`. - **API users:** the Repository V1 endpoints `/api/repository/modules` and `/api/repository/modules/{module_id}` are removed; use `/api/repository/v2/`. - **Still on the pre-26.7.0 `ryax-worker` chart:** migrate to `ryax-worker-k8s` or `ryax-worker-slurm-ssh` first, as the V1 worker protocol is gone. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.9.0 \ -n ryaxns \ --take-ownership \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.9.0 \ -n ryaxns \ --take-ownership \ -f worker.yaml ``` -
26.9.0-rc3
45fb18c2 · ·Ryax 26.9.0-rc3 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ Stability and security updates, plus a GitOps-ready Helm chart and a rebuilt web interface. - **GitOps installs.** The chart renders without a cluster connection, so ArgoCD and Flux no longer mint fresh credentials on every sync. Set `global.secrets.create=false` to supply the secrets yourself. - **Pick the Ingress controller.** Ryax's Ingresses name their class through `global.ingress.className`, overridable per service or switched off with `<subchart>.ingress.enabled=false`. The bundled Traefik no longer registers itself as the cluster-wide default class. - **Pod placement.** `global.tolerations` and `global.affinity` apply to every Ryax pod, with per-subchart overrides; the worker charts gained `tolerations` and `nodeSelector`. - **Per-site action registry.** A worker whose nodes cannot resolve the registry the Runner recorded overrides it with `internalRegistryOverride`. - **Rebuilt web interface.** Angular 16 → 21, Nx 22, TypeScript 5.9. - **Generated API reference.** <https://docs.ryax.tech/reference/api/> is now built from the service sources on every release, so it cannot drift again — the 26.7.0 document still described 26.2.0. - Prometheus keeps its metrics across restarts: the volume request sat one level too high in the values and was silently ignored, leaving it on an `emptyDir`. - Kubernetes worker database upgrades work again — the PostgreSQL service name and the database URL now agree, so the migration init container resolves its host. - IntelliScale no longer restarts on unrelated configuration changes. - The Runner and Repository APIs moved to FastAPI and pydantic; Studio dropped marshmallow; the Authorization service is now part of the core service. - The V1 worker protocol and the legacy worker module are removed. - Container image publishing is reproducible again, fixing intermittent `Digest did not match` failures. - Security fixes across the stack, including `fast-uri` 3.1.6 and the `js-yaml`, `svgo` and `extract-zip` advisories, plus a repaired image CVE scan. - Observability dependencies updated: kube-prometheus-stack 88.x, Loki 7.3, Alloy 1.12, Traefik 41.5. Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **`--take-ownership` is required when upgrading from 26.7.0.** Three credential secrets used to be created as Helm *hook* resources, which Helm never records as part of the release: the Studio password encryption key in the main chart, and the PostgreSQL credentials in both worker charts. They are now ordinary chart-managed resources, so Helm finds them un-owned and refuses the upgrade with `invalid ownership metadata` unless you let it adopt them. Adoption preserves the existing values -- the templates read the current secret before falling back -- so the encryption key and the database passwords are unchanged. - **Prometheus storage:** if your values set `kube-prometheus-stack.prometheus.storage.volumeClaimTemplate`, move it to `prometheus.prometheusSpec.storageSpec` and add `accessModes: ["ReadWriteOnce"]`. The old key was never read. There is no metric history to preserve, since Prometheus was running on an `emptyDir`. - **Traefik:** the bundled instance no longer claims Ingresses that name no class. Name `<release-name>-traefik` on your own Ingresses, or mark your controller as the cluster default. - **Workers relying on the old `internalRegistryOverride` default:** on the SLURM_SSH chart it changed from `ryax-registry:5000` to empty. If you never set it yourself, set it explicitly to keep pulling through the in-cluster registry. On a Kubernetes worker whose nodes cannot resolve the Runner's address, use `127.0.0.1:30012`. - **API users:** the Repository V1 endpoints `/api/repository/modules` and `/api/repository/modules/{module_id}` are removed; use `/api/repository/v2/`. - **Still on the pre-26.7.0 `ryax-worker` chart:** migrate to `ryax-worker-k8s` or `ryax-worker-slurm-ssh` first, as the V1 worker protocol is gone. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.9.0 \ -n ryaxns \ --take-ownership \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.9.0 \ -n ryaxns \ --take-ownership \ -f worker.yaml ``` -
26.9.0-rc2
f0860125 · ·Ryax 26.9.0-rc2 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ Stability and security updates, plus a GitOps-ready Helm chart and a rebuilt web interface. - **GitOps installs.** The chart renders without a cluster connection, so ArgoCD and Flux no longer mint fresh credentials on every sync. Set `global.secrets.create=false` to supply the secrets yourself. - **Pick the Ingress controller.** Ryax's Ingresses name their class through `global.ingress.className`, overridable per service or switched off with `<subchart>.ingress.enabled=false`. The bundled Traefik no longer registers itself as the cluster-wide default class. - **Pod placement.** `global.tolerations` and `global.affinity` apply to every Ryax pod, with per-subchart overrides; the worker charts gained `tolerations` and `nodeSelector`. - **Per-site action registry.** A worker whose nodes cannot resolve the registry the Runner recorded overrides it with `internalRegistryOverride`. - **Rebuilt web interface.** Angular 16 → 21, Nx 22, TypeScript 5.9. - **Generated API reference.** <https://docs.ryax.tech/reference/api/> is now built from the service sources on every release, so it cannot drift again — the 26.7.0 document still described 26.2.0. - Prometheus keeps its metrics across restarts: the volume request sat one level too high in the values and was silently ignored, leaving it on an `emptyDir`. - Kubernetes worker database upgrades work again — the PostgreSQL service name and the database URL now agree, so the migration init container resolves its host. - IntelliScale no longer restarts on unrelated configuration changes. - The Runner and Repository APIs moved to FastAPI and pydantic; Studio dropped marshmallow; the Authorization service is now part of the core service. - The V1 worker protocol and the legacy worker module are removed. - Container image publishing is reproducible again, fixing intermittent `Digest did not match` failures. - Security fixes across the stack, including `fast-uri` 3.1.6 and the `js-yaml`, `svgo` and `extract-zip` advisories, plus a repaired image CVE scan. - Observability dependencies updated: kube-prometheus-stack 88.x, Loki 7.3, Alloy 1.12, Traefik 41.5. Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **`--take-ownership` is required when upgrading from 26.7.0.** Three credential secrets used to be created as Helm *hook* resources, which Helm never records as part of the release: the Studio password encryption key in the main chart, and the PostgreSQL credentials in both worker charts. They are now ordinary chart-managed resources, so Helm finds them un-owned and refuses the upgrade with `invalid ownership metadata` unless you let it adopt them. Adoption preserves the existing values -- the templates read the current secret before falling back -- so the encryption key and the database passwords are unchanged. - **Prometheus storage:** if your values set `kube-prometheus-stack.prometheus.storage.volumeClaimTemplate`, move it to `prometheus.prometheusSpec.storageSpec` and add `accessModes: ["ReadWriteOnce"]`. The old key was never read. There is no metric history to preserve, since Prometheus was running on an `emptyDir`. - **Traefik:** the bundled instance no longer claims Ingresses that name no class. Name `<release-name>-traefik` on your own Ingresses, or mark your controller as the cluster default. - **Workers relying on the old `internalRegistryOverride` default:** on the SLURM_SSH chart it changed from `ryax-registry:5000` to empty. If you never set it yourself, set it explicitly to keep pulling through the in-cluster registry. On a Kubernetes worker whose nodes cannot resolve the Runner's address, use `127.0.0.1:30012`. - **API users:** the Repository V1 endpoints `/api/repository/modules` and `/api/repository/modules/{module_id}` are removed; use `/api/repository/v2/`. - **Still on the pre-26.7.0 `ryax-worker` chart:** migrate to `ryax-worker-k8s` or `ryax-worker-slurm-ssh` first, as the V1 worker protocol is gone. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.9.0 \ -n ryaxns \ --take-ownership \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.9.0 \ -n ryaxns \ --take-ownership \ -f worker.yaml ``` -
26.9.0-rc1
43efd1b6 · ·Ryax 26.9.0-rc1 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ Stability and security updates, plus a GitOps-ready Helm chart and a rebuilt web interface. - **GitOps installs.** The chart renders without a cluster connection, so ArgoCD and Flux no longer mint fresh credentials on every sync. Set `global.secrets.create=false` to supply the secrets yourself. - **Pick the Ingress controller.** Ryax's Ingresses name their class through `global.ingress.className`, overridable per service or switched off with `<subchart>.ingress.enabled=false`. The bundled Traefik no longer registers itself as the cluster-wide default class. - **Pod placement.** `global.tolerations` and `global.affinity` apply to every Ryax pod, with per-subchart overrides; the worker charts gained `tolerations` and `nodeSelector`. - **Per-site action registry.** A worker whose nodes cannot resolve the registry the Runner recorded overrides it with `internalRegistryOverride`. - **Rebuilt web interface.** Angular 16 → 21, Nx 22, TypeScript 5.9. - **Generated API reference.** <https://docs.ryax.tech/reference/api/> is now built from the service sources on every release, so it cannot drift again — the 26.7.0 document still described 26.2.0. - Prometheus keeps its metrics across restarts: the volume request sat one level too high in the values and was silently ignored, leaving it on an `emptyDir`. - Kubernetes worker database upgrades work again — the PostgreSQL service name and the database URL now agree, so the migration init container resolves its host. - IntelliScale no longer restarts on unrelated configuration changes. - The Runner and Repository APIs moved to FastAPI and pydantic; Studio dropped marshmallow; the Authorization service is now part of the core service. - The V1 worker protocol and the legacy worker module are removed. - Container image publishing is reproducible again, fixing intermittent `Digest did not match` failures. - Security fixes across the stack, including `fast-uri` 3.1.6 and the `js-yaml`, `svgo` and `extract-zip` advisories, plus a repaired image CVE scan. - Observability dependencies updated: kube-prometheus-stack 88.x, Loki 7.3, Alloy 1.12, Traefik 41.5. Restore your values file if you do not have it: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **`--take-ownership` is required when upgrading from 26.7.0.** Three credential secrets used to be created as Helm *hook* resources, which Helm never records as part of the release: the Studio password encryption key in the main chart, and the PostgreSQL credentials in both worker charts. They are now ordinary chart-managed resources, so Helm finds them un-owned and refuses the upgrade with `invalid ownership metadata` unless you let it adopt them. Adoption preserves the existing values -- the templates read the current secret before falling back -- so the encryption key and the database passwords are unchanged. - **Prometheus storage:** if your values set `kube-prometheus-stack.prometheus.storage.volumeClaimTemplate`, move it to `prometheus.prometheusSpec.storageSpec` and add `accessModes: ["ReadWriteOnce"]`. The old key was never read. There is no metric history to preserve, since Prometheus was running on an `emptyDir`. - **Traefik:** the bundled instance no longer claims Ingresses that name no class. Name `<release-name>-traefik` on your own Ingresses, or mark your controller as the cluster default. - **Workers relying on the old `internalRegistryOverride` default:** on the SLURM_SSH chart it changed from `ryax-registry:5000` to empty. If you never set it yourself, set it explicitly to keep pulling through the in-cluster registry. On a Kubernetes worker whose nodes cannot resolve the Runner's address, use `127.0.0.1:30012`. - **API users:** the Repository V1 endpoints `/api/repository/modules` and `/api/repository/modules/{module_id}` are removed; use `/api/repository/v2/`. - **Still on the pre-26.7.0 `ryax-worker` chart:** migrate to `ryax-worker-k8s` or `ryax-worker-slurm-ssh` first, as the V1 worker protocol is gone. Then run the upgrade: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.9.0 \ -n ryaxns \ --take-ownership \ -f values.yaml ``` And each worker with its own values: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.9.0 \ -n ryaxns \ --take-ownership \ -f worker.yaml ``` -
26.9.0-rc0
69c8cde0 · ·Ryax 26.9.0-rc0 We are proud to announce the release of: ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ This release makes Ryax easier to run the way you already run the rest of your platform. The Helm chart is now installable by ArgoCD and other GitOps engines, you choose which Ingress controller serves Ryax instead of having one imposed on the cluster, and every Ryax pod can be steered onto the nodes you want. The web interface has been rebuilt on a current Angular, the last two services moved to FastAPI, the published API reference is now generated from the service sources, and Prometheus finally keeps its metrics across restarts. The chart can now be rendered without a cluster connection, which is what ArgoCD's and Flux's repo servers do. Previously the generated credentials came from `lookup()`, which returns nothing in that situation, so every render minted fresh passwords and rolled them out to the running pods. Set `global.secrets.create=false` to supply every credential yourself --- with sealed-secrets, external-secrets, or a plain `kubectl create` --- under the secret names documented in the values files. The worker charts get the same option, and the bundled PostgreSQL has its own `createSecret` switch. Ryax's Ingresses now name their IngressClass explicitly through `global.ingress.className`, which defaults to the bundled Traefik. Point it at your own controller, override it for a single service with `<subchart>.ingress.className`, or turn an Ingress off entirely with `<subchart>.ingress.enabled=false`. The bundled Traefik also no longer registers itself as the cluster-wide default IngressClass. That setting is cluster-scoped, so it used to claim every Ingress in every namespace that did not name a class --- including ones belonging to applications that have nothing to do with Ryax. `global.tolerations` and `global.affinity` are injected into every Ryax pod, so Ryax can run on tainted or dedicated nodes. Both can be overridden per subchart, and the worker charts gained their own `tolerations` and `nodeSelector`. The registry an action image is pulled from is now a property of the site that runs it, rather than a single global address. A worker whose nodes cannot resolve the address the Runner recorded overrides it with `internalRegistryOverride` in its own values, and the Runner resolves the registry host at deploy time. The front end moved from Angular 16 to Angular 21, with Nx 22 and TypeScript 5.9. The interface behaves as before, on a supported and maintained toolchain. <https://docs.ryax.tech/reference/api/> and the [`ryax-spec.json`](https://docs.ryax.tech/reference/ryax-spec.json) it is built from are now generated from the Authorization, Repository, Studio and Runner sources on every release, so they cannot drift from the running services again. The document that shipped with 26.7.0 still described 26.2.0; the regenerated one covers 89 endpoints, documents the `/api/...` paths as the ingress actually serves them, and includes the Site and Node Pool management endpoints the old document was missing. - Prometheus keeps its metrics across restarts. The 10Gi volume request had been one level too high in the values for a long time, so it was silently ignored and Prometheus ran on an `emptyDir`, losing every metric on each restart, reschedule and chart upgrade. - Kubernetes worker database upgrades work again: the PostgreSQL service name and the database URL now agree, so the migration init container can resolve its host. - IntelliScale no longer restarts when an unrelated part of the configuration changes. - The Runner and Repository APIs migrated to FastAPI and pydantic, joining the Studio, which also dropped marshmallow. The Authorization service is now part of the core service, one less component to track. - The V1 worker protocol and the legacy worker module are removed. - Container image publishing is reproducible again. Image pushes failed intermittently with `Digest did not match` and missing layer tars, because the Python dependency closure was built by a derivation that downloaded from the network: the same store path held different bytes on different runners. - Security fixes across the stack: `fast-uri` 3.1.6 and the `js-yaml`, `svgo` and `extract-zip` advisories in the front end, the image CVE scan pointed back at the right sources after a repository rename, and vulture and bandit now run in the core service's own CI. - Observability dependencies updated: kube-prometheus-stack 88.x, Loki 7.3, Grafana Alloy 1.12 and Traefik 41.5. To upgrade your main cluster, find the values file from your previous install or restore it using: ```sh helm get values -n ryaxns ryax --output yaml > values.yaml ``` Admins should take care of the following elements when upgrading to this version: - **If your values customise Prometheus storage**, move the block from `kube-prometheus-stack.prometheus.storage.volumeClaimTemplate` to `kube-prometheus-stack.prometheus.prometheusSpec.storageSpec`, and include `accessModes: ["ReadWriteOnce"]`. The old key is not one the subchart reads, so it never took effect. Prometheus gets a PersistentVolumeClaim on upgrade; because it was previously running on an `emptyDir` there is no metric history to preserve. - **If you relied on the bundled Traefik serving your own Ingresses** without naming an IngressClass, it no longer claims them. Name the class explicitly on those Ingresses (`<release-name>-traefik`), or mark your own controller as the cluster default. - **If you run a SLURM_SSH worker on a private network**, the `internalRegistryOverride` default changed from `ryax-registry:5000` to empty. Set it explicitly in your worker values to keep pulling action images through the in-cluster registry. - **On a Kubernetes worker whose nodes cannot resolve the registry address the Runner records**, set `internalRegistryOverride` to `127.0.0.1:30012` to reach the bundled registry through its NodePort. - **API users:** the Repository V1 endpoints `/api/repository/modules` and `/api/repository/modules/{module_id}` are removed. Use the `/api/repository/v2/` endpoints; the current set is published at <https://docs.ryax.tech/reference/ryax-spec.json>. - **If you are still on the pre-26.7.0 `ryax-worker` chart**, migrate to `ryax-worker-k8s` or `ryax-worker-slurm-ssh` following the 26.7.0 release note before upgrading: the V1 worker protocol is gone in this release. Then, run the upgrade with: ```sh helm upgrade ryax oci://registry.ryax.org/release-charts/ryax-engine:26.9.0 \ -n ryaxns \ -f values.yaml ``` And upgrade each worker with its own values, for example: ```sh helm upgrade ryax-worker-k8s oci://registry.ryax.org/release-charts/ryax-worker-k8s:26.9.0 \ -n ryaxns \ -f worker.yaml ``` -
26.7.0-rc2
9f2b00a8 · · -
26.7.0-rc1
652ed3f4 · · -
26.7.0-rc0
834b95d3 · · -
26.4.0-rc5
1dd417ea · · -
26.4.0-rc4
d6cc69d5 · · -
26.4.0-rc0
82fb7321 · · -
26.4.0-rc2
82fb7321 · · -
26.4.0-rc1
57dc06c3 · · -