A working, end-to-end demonstration of an Argo CD + Argo Rollouts + Prometheus pipeline.
The canary v2.4 leaks memory until it either gets OOMKilled or burns its error-rate SLO — whichever trips first — and Argo Rollouts aborts the release on its own. A real GLM-4.5 call writes the root-cause analysis from the analyzer findings. This page is a recording of that pipeline running — the pipeline itself is not mocked and the abort is not scripted.
Three clips from one cluster: Argo CD syncs → Argo Rollouts shifts canary traffic → Prometheus trips the error-rate SLO → the analyzers run → the AnalysisRun fails → Argo Rollouts aborts and reverts.
cycle.sh repeats the loop; a cycle runs roughly 2-5 minutes depending on how fast the canary breaches. Every value on screen — traffic split, SLO badge, pod phase, analyzer findings — comes from a real HTTP call to a real backend endpoint reading the cluster. No client-side state machine hiding the work.
GET /api/analyze — seven analyzers over live cluster state — and routes the finding to POST /api/explain, which makes a real GLM-4.5 call via z-ai-web-dev-sdk. In this recording that call fails: the box had no ZAI_API_KEY, so the route answers 503 and the terminal prints the error instead of a diagnosis. Everything above that line is live cluster data. Set the key and the structured response renders in the same panel.
prometheus-slo AnalysisRun fails, the Rollout goes Degraded, and the controller shifts traffic back to the stable ReplicaSet on its own. The GLM-4.5 call explains what happened; it does not trigger it. In the clip the canary image reverts to v2.3 and Argo CD re-syncs.
The response shape from GET /api/analyze, with one real finding. Rule-based analyzers modelled on the k8sgpt approach — no LLM involved at this stage. Findings come from live cluster state.
$ curl -s localhost:3000/api/analyze | python3 -m json.tool analyzers: pod, deployment, service, rollout, pvc, node, log source: live cluster via KUBECONFIG - findings depend on what the cluster is doing when you call it, so run scripts/capture-assets.sh for your own. { "provider": "", "status": "ProblemDetected", "problems": 1, "analyzers": ["pod", "deployment", "service", "rollout", "pvc", "node", "log"], "results": [ { "kind": "Rollout", "name": "payment-prod/payments-api", "analyzer": "rollout", "severity": "warning", "error": [{ "Text": "Rollout payments-api is paused at the analysis step" }], "suggestedFix": "kubectl argo rollouts get rollout payments-api -n payment-prod" } ], "timestamp": "2026-08-22T22:04:50.146Z" }
The shape of a POST /api/explain response, rendered as an SRE incident card. That endpoint makes a real GLM-4.5 call via z-ai-web-dev-sdk and needs a ZAI_API_KEY; without one it returns an error rather than a diagnosis. The wording below is illustrative - the live model writes its own, grounded only in the analyzer findings.
The payments-api-canary v2.4 pods are being terminated due to memory exhaustion, causing the canary deployment to fail SLO checks.
canary-error-rate measurement is over its 1% thresholdThe canary deployment is paused with 50% traffic, preventing the rollout of v2.4 and potentially causing service degradation if the stable version is also affected.
kubectl logs payments-api-<hash> -c api --previouskubectl logs payments-api-<hash> -c api --previous | grep -i memorykubectl argo rollouts promote payments-api -n payment-prodThe v2.4 canary pods are failing due to memory leaks causing heap growth beyond allocated limits, resulting in OOMKilled termination (exit code 137). The canary traffic split is exposing these memory issues under production load, while v2.3 remains healthy with proper memory management.
Five real backend endpoints, every one replaceable. The UI just renders what they return.
┌──────────────────────────────────────────────────────────────────────────────┐
│ Browser (Next.js page.tsx, polls every 1s) │
│ │ │
│ ▼ fetch("/api/cluster-state") │
│ ┌─────────────────────────────────────────────────────────────────────┐ │
│ │ /api/cluster-state (route.ts) │ │
│ │ ├─ returns: phase, argoCdSync, argoRollouts, traffic, metrics, │ │
│ │ │ slo, pods, findings │ │
│ │ ├─ derives phase from Rollout status + pods + PromQL │ │
│ │ └─ reads every field from the cluster per request │ │
│ └─────────────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ fetch("/api/analyze") │
│ ┌─────────────────────────────────────────────────────────────────────┐ │
│ │ /api/analyze (route.ts) │ │
│ │ ├─ reads from: live cluster via KUBECONFIG │ │
│ │ ├─ runs 7 rule-based analyzers (pod, deployment, service, │ │
│ │ │ rollout, pvc, node, log) — no LLM │ │
│ │ └─ returns: findings derived from live cluster state │ │
│ └─────────────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ fetch("/api/explain", {method: POST, body: findings}) │
│ ┌─────────────────────────────────────────────────────────────────────┐ │
│ │ /api/explain (route.ts) ★ REAL LLM │ │
│ │ ├─ builds a structured prompt (findings → Root Cause template) │ │
│ │ ├─ calls ZAI.create() then zai.chat.completions.create({ │ │
│ │ │ model: "glm-4.5", │ │
│ │ │ messages: [...], temperature: 0.3, max_tokens: 800 │ │
│ │ │ }) │ │
│ │ ├─ server-side cache (instant replay on demo re-runs) │ │
│ │ └─ returns: {content, cached, model: "glm-4.5", timestamp} │ │
│ └─────────────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ fetch("/api/prometheus?query=...&range=1") │
│ ┌─────────────────────────────────────────────────────────────────────┐ │
│ │ /api/prometheus (route.ts) │ │
│ │ ├─ accepts real PromQL via ?query= │ │
│ │ ├─ maintains 30-point ring buffers per metric │ │
│ │ └─ returns: standard v1/query envelope (status, resultType, result)│ │
│ └─────────────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ fetch("/api/k8s/api/v1/namespaces/payment-prod/pods") │
│ ┌─────────────────────────────────────────────────────────────────────┐ │
│ │ /api/k8s/[...path] (route.ts) — live kube-apiserver proxy │ │
│ │ ├─ returns standard K8s list envelopes (PodList, DeploymentList…) │ │
│ │ ├─ auth + TLS taken from the kubeconfig │ │
│ │ └─ source: live cluster via KUBECONFIG │ │
│ │ (whatever the cluster reports at request time - │ │
│ │ no fixtures, no canned bodies, no synthetic latency) │ │
│ └─────────────────────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────────────────────────┘
Five components, every one replaceable. Every one reads from the live cluster; nothing is stubbed.
Watches manifests-repo/ and syncs it into the payment-prod namespace. The card shows the Application's real revision, sync status and last operation message.
Drives the canary: 20% → 50% traffic shift, pauses on the analysis step, aborts on failure. The traffic pipe animates in real-time as Argo Rollouts shifts weight between stable and canary ReplicaSets.
Scrapes http_error_rate and http_request_duration_p99 for the canary pods. When the canary breaks, both series move, and the AnalysisRun compares them against the 1% error-rate SLO in the AnalysisTemplate.
Runs 7 structured analyzers over live cluster state (Pod, Deployment, Service, Rollout, PVC, Node, log). Returns findings in a k8sgpt-shaped JSON envelope — no LLM at this stage, just deterministic rule-based detection.
For each finding, the LLM produces a structured root-cause analysis citing specific pod names + kubectl commands. Called via the z-ai-web-dev-sdk. Server-side cached so demo re-runs are instant.
Nothing on this page is asserted by hand. scripts/uat-test.sh
queries the live cluster, Argo CD, Argo Rollouts, Prometheus and every app route, then writes
its verdict to uat-results.json. The table below is that file, rendered.
bash scripts/uat-test.sh| ID | Layer | Check | Status | Observed |
|---|---|---|---|---|
No run recorded yet — run bash scripts/uat-test.sh. | ||||
Two environments, the same three commands. Nothing auto-starts — you bring the cluster up, then the dashboard, then the controller.
The devcontainer provides Docker-in-Docker, Node 22, kubectl, Helm and bun, and
runs bun install. It does not bootstrap the cluster —
that is setup.sh, below. The 2-core machine type will run the cluster but
not the Playwright recording.
Clone the repo and run the same three commands in three terminals. k3d builds the
k3s cluster inside Docker; setup.sh is idempotent, so re-running it against
an existing cluster is cheap.
Every assertion on this page is re-checkable. The script queries the cluster, Argo CD,
Argo Rollouts, Prometheus and all five app routes, writes uat-results.json
(the table in section 06), and exits non-zero if anything fails.