★ RECORDED RUN · REAL KUBERNETES · REAL LLM

GitOps Progressive Delivery,
diagnosed by AI.

A working, end-to-end demonstration of an Argo CD + Argo Rollouts + Prometheus pipeline. The canary v2.4 leaks memory until it either gets OOMKilled or burns its error-rate SLO — whichever trips first — and Argo Rollouts aborts the release on its own. A real GLM-4.5 call writes the root-cause analysis from the analyzer findings. This page is a recording of that pipeline running — the pipeline itself is not mocked and the abort is not scripted.

CI License LLM Next.js
5
Real backend endpoints
7
Rule-based analyzers
--
live checks pass
GLM-4.5
Real LLM diagnosis

01Recorded run

Three clips from one cluster: Argo CD syncs → Argo Rollouts shifts canary traffic → Prometheus trips the error-rate SLO → the analyzers run → the AnalysisRun fails → Argo Rollouts aborts and reverts.

Full pipeline walkthrough (70s)
RECORDED FROM LIVE APP
Pipeline walkthrough
cycle.sh repeats the loop; a cycle runs roughly 2-5 minutes depending on how fast the canary breaches. Every value on screen — traffic split, SLO badge, pod phase, analyzer findings — comes from a real HTTP call to a real backend endpoint reading the cluster. No client-side state machine hiding the work.
Analyzer terminal (45s)
LLM CALL NOT CONFIGURED
Analyzer terminal stream
The terminal calls GET /api/analyze — seven analyzers over live cluster state — and routes the finding to POST /api/explain, which makes a real GLM-4.5 call via z-ai-web-dev-sdk. In this recording that call fails: the box had no ZAI_API_KEY, so the route answers 503 and the terminal prints the error instead of a diagnosis. Everything above that line is live cluster data. Set the key and the structured response renders in the same panel.
Rollout abort + traffic revert (24s)
ARGO ROLLOUTS
Rollback transition
The abort is Argo Rollouts’ own decision, not the LLM’s: the prometheus-slo AnalysisRun fails, the Rollout goes Degraded, and the controller shifts traffic back to the stable ReplicaSet on its own. The GLM-4.5 call explains what happened; it does not trigger it. In the clip the canary image reverts to v2.3 and Argo CD re-syncs.

02Real cluster analyzer output

The response shape from GET /api/analyze, with one real finding. Rule-based analyzers modelled on the k8sgpt approach — no LLM involved at this stage. Findings come from live cluster state.

analyzer — payment-prod — 7 analyzers
$ curl -s localhost:3000/api/analyze | python3 -m json.tool
analyzers: pod, deployment, service, rollout, pvc, node, log
source: live cluster via KUBECONFIG - findings depend on what the cluster
is doing when you call it, so run scripts/capture-assets.sh for your own.

{
  "provider": "",
  "status": "ProblemDetected",
  "problems": 1,
  "analyzers": ["pod", "deployment", "service", "rollout", "pvc", "node", "log"],
  "results": [
    {
      "kind": "Rollout",
      "name": "payment-prod/payments-api",
      "analyzer": "rollout",
      "severity": "warning",
      "error": [{ "Text": "Rollout payments-api is paused at the analysis step" }],
      "suggestedFix": "kubectl argo rollouts get rollout payments-api -n payment-prod"
    }
  ],
  "timestamp": "2026-08-22T22:04:50.146Z"
}

03Real GLM-4.5 root-cause analysis

The shape of a POST /api/explain response, rendered as an SRE incident card. That endpoint makes a real GLM-4.5 call via z-ai-web-dev-sdk and needs a ZAI_API_KEY; without one it returns an error rather than a diagnosis. The wording below is illustrative - the live model writes its own, grounded only in the analyzer findings.

#sre-incidents payments-api → analyzer → GLM-4.5 model=glm-4.5

🔥 Root Cause

The payments-api-canary v2.4 pods are being terminated due to memory exhaustion, causing the canary deployment to fail SLO checks.

Evidence

  • The canary ReplicaSet is not reaching ready replicas
  • The Rollout is paused at its analysis step
  • The canary-error-rate measurement is over its 1% threshold

Impact

The canary deployment is paused with 50% traffic, preventing the rollout of v2.4 and potentially causing service degradation if the stable version is also affected.

Recommended Action

  1. kubectl logs payments-api-<hash> -c api --previous
  2. kubectl logs payments-api-<hash> -c api --previous | grep -i memory
  3. kubectl argo rollouts promote payments-api -n payment-prod

Diagnosis

The v2.4 canary pods are failing due to memory leaks causing heap growth beyond allocated limits, resulting in OOMKilled termination (exit code 137). The canary traffic split is exposing these memory issues under production load, while v2.3 remains healthy with proper memory management.

04Architecture

Five real backend endpoints, every one replaceable. The UI just renders what they return.

request flow — single page load
┌──────────────────────────────────────────────────────────────────────────────┐
│  Browser  (Next.js page.tsx, polls every 1s)                                  │
│     │                                                                         │
│     ▼  fetch("/api/cluster-state")                                            │
│  ┌─────────────────────────────────────────────────────────────────────┐     │
│  │  /api/cluster-state  (route.ts)                                       │     │
│  │   ├─ returns: phase, argoCdSync, argoRollouts, traffic, metrics,    │     │
│  │   │            slo, pods, findings                                    │     │
│  │   ├─ derives phase from Rollout status + pods + PromQL                   │     │
│  │   └─ reads every field from the cluster per request               │         │
│  └─────────────────────────────────────────────────────────────────────┘     │
│     │                                                                         │
│     ▼  fetch("/api/analyze")                                                  │
│  ┌─────────────────────────────────────────────────────────────────────┐     │
│  │  /api/analyze  (route.ts)                                            │     │
│  │   ├─ reads from: live cluster via KUBECONFIG                            │     │
│  │   ├─ runs 7 rule-based analyzers (pod, deployment, service,    │          │
│  │   │   rollout, pvc, node, log) — no LLM                              │     │
│  │   └─ returns: findings derived from live cluster state                    │     │
│  └─────────────────────────────────────────────────────────────────────┘     │
│     │                                                                         │
│     ▼  fetch("/api/explain", {method: POST, body: findings})                  │
│  ┌─────────────────────────────────────────────────────────────────────┐     │
│  │  /api/explain  (route.ts)  ★ REAL LLM                                │     │
│  │   ├─ builds a structured prompt (findings → Root Cause template)     │     │
│  │   ├─ calls ZAI.create() then zai.chat.completions.create({                    │     │
│  │   │     model: "glm-4.5",                                                  │     │
│  │   │     messages: [...], temperature: 0.3, max_tokens: 800                  │     │
│  │   │   })                                                                    │     │
│  │   ├─ server-side cache (instant replay on demo re-runs)              │     │
│  │   └─ returns: {content, cached, model: "glm-4.5", timestamp}         │     │
│  └─────────────────────────────────────────────────────────────────────┘     │
│     │                                                                         │
│     ▼  fetch("/api/prometheus?query=...&range=1")                             │
│  ┌─────────────────────────────────────────────────────────────────────┐     │
│  │  /api/prometheus  (route.ts)                                          │     │
│  │   ├─ accepts real PromQL via ?query=                                  │     │
│  │   ├─ maintains 30-point ring buffers per metric                       │     │
│  │   └─ returns: standard v1/query envelope (status, resultType, result)│     │
│  └─────────────────────────────────────────────────────────────────────┘     │
│     │                                                                         │
│     ▼  fetch("/api/k8s/api/v1/namespaces/payment-prod/pods")                 │
│  ┌─────────────────────────────────────────────────────────────────────┐     │
│  │  /api/k8s/[...path]  (route.ts)  — live kube-apiserver proxy                │     │
│  │   ├─ returns standard K8s list envelopes (PodList, DeploymentList…)  │     │
│  │   ├─ auth + TLS taken from the kubeconfig                                      │     │
│  │   └─ source: live cluster via KUBECONFIG                                │     │
│  │       (whatever the cluster reports at request time -               │     │
│  │        no fixtures, no canned bodies, no synthetic latency)          │     │
│  └─────────────────────────────────────────────────────────────────────┘     │
└──────────────────────────────────────────────────────────────────────────────┘

05The stack

Five components, every one replaceable. Every one reads from the live cluster; nothing is stubbed.

A

Argo CD

GitOps sync engine

Watches manifests-repo/ and syncs it into the payment-prod namespace. The card shows the Application's real revision, sync status and last operation message.

API: /api/cluster-state → argoCdSync
R

Argo Rollouts

Canary + traffic shifting

Drives the canary: 20% → 50% traffic shift, pauses on the analysis step, aborts on failure. The traffic pipe animates in real-time as Argo Rollouts shifts weight between stable and canary ReplicaSets.

API: /api/cluster-state → argoRollouts
P

Prometheus

SLO metrics

Scrapes http_error_rate and http_request_duration_p99 for the canary pods. When the canary breaks, both series move, and the AnalysisRun compares them against the 1% error-rate SLO in the AnalysisTemplate.

API: /api/prometheus?query=...
K

Cluster analyzer

SRE agent (rule-based)

Runs 7 structured analyzers over live cluster state (Pod, Deployment, Service, Rollout, PVC, Node, log). Returns findings in a k8sgpt-shaped JSON envelope — no LLM at this stage, just deterministic rule-based detection.

API: /api/analyze · 7 analyzers, no LLM
G

GLM-4.5 (Z.AI)

LLM brain — REAL API calls

For each finding, the LLM produces a structured root-cause analysis citing specific pod names + kubectl commands. Called via the z-ai-web-dev-sdk. Server-side cached so demo re-runs are instant.

API: /api/explain · model=glm-4.5

06End-to-end verification

Nothing on this page is asserted by hand. scripts/uat-test.sh queries the live cluster, Argo CD, Argo Rollouts, Prometheus and every app route, then writes its verdict to uat-results.json. The table below is that file, rendered.

--/ --
Run it yourself: bash scripts/uat-test.sh
IDLayerCheckStatusObserved
No run recorded yet — run bash scripts/uat-test.sh.

07Run it yourself

Two environments, the same three commands. Nothing auto-starts — you bring the cluster up, then the dashboard, then the controller.

▲

GitHub Codespaces

4-core / 16GB · ~5 min to first cycle

The devcontainer provides Docker-in-Docker, Node 22, kubectl, Helm and bun, and runs bun install. It does not bootstrap the cluster — that is setup.sh, below. The 2-core machine type will run the cluster but not the Playwright recording.

$

Local machine

Docker · 4 CPU · 8GB

Clone the repo and run the same three commands in three terminals. k3d builds the k3s cluster inside Docker; setup.sh is idempotent, so re-running it against an existing cluster is cheap.

Run: bash setup.sh → bash dev-real.sh → bash demo-controller/cycle.sh
✓

Check the claims

17 live checks

Every assertion on this page is re-checkable. The script queries the cluster, Argo CD, Argo Rollouts, Prometheus and all five app routes, writes uat-results.json (the table in section 06), and exits non-zero if anything fails.

Run: bash scripts/uat-test.sh