Skip to content

Eval scorecard

The nightly replay eval, published in full — per scenario, red or green.

Everything above the Latest published run heading is written by hand and lives in the repo; everything below it is generated by the nightly and replaced on every publish.

What a scenario is

One recorded incident: the exact output every tool returned during a real failure, frozen into a YAML file, together with the cause the investigation is expected to reach (must_contain, min_confidence, and for the recall cases expect_recall). The nightly replays those files through the real investigation loop — same prompts, same tools, same recall and verify gates — with no cluster attached.

So the pass-rate scores the model and the loop reasoning over fixed evidence. Nothing else in the run can move.

Read the cases before you read the number: examples/eval/. That directory is the suite as it stands today; the published run below states its own date and how many scenarios it covered. If the directory now holds more, the suite has grown since that run and the next nightly picks up the difference.

What the number does not mean

Not a live-cluster result. Replay isolates reasoning over fixed evidence, so it says nothing about tool flakiness, latency, or the evidence gaps that only appear when a query actually runs against a cluster. Live-fire (lore eval --live) is the harder test, and it is not what this page reports.

Not a leaderboard. The corpus is small and hand-built; it exists to catch regressions in RunLore. That makes the pass-rate meaningful only against its own history — the run-history table the scorecard carries. To weigh models against each other, run the benchmarking harness on your own models rather than reading this one number as a ranking.

Not a field result. ITBench (IBM/ICML 2025) found frontier models identify the root cause < 50% of the time and fully resolve only ~11–14% of real K8s incidents. Treat sub-50% as the baseline and a high replay pass-rate as a ceiling — see Prior art.

A run below the 70% CI gate publishes here exactly like a green one — same schedule, same detail, no editorial pass. A pass-rate that drops is visible on this page before it is fixed; that is what publishing it is for.

Latest published run

Published 2026-09-16T11:00:04Z from the eval-scorecard branch. It covered the 6 scenarios the suite held at that time. The case definitions are the suite as it stands today — more files there than this run covered means it has grown since.

Auto-published by .github/workflows/eval.yaml. Reproduce it yourself:

lore eval -config eval/ci.runlore.yaml -cases examples/eval -n 5 -fail-under 0.7

Latest run: 2026-09-16T11:00:02Z · model openai/glm-4.5-air · 2/6 scenarios reached (33%) · n=5 runs/case, k-of-n bar 70% · est. cost $0.31 (1.2M in / 64.1k out tokens)

Scenarios (latest run)

scenarioresultpass-ratemedian confidencerecallnotes
gitops-broken-kustomization✅ PASS100% (n=5)0.90
harbor-chart-bump❌ MISS0% (n=5)0.90harbor-db
node-eviction-no-commons❌ MISS20% (n=5)0.70fired 0/5 · short-circuit 0/5 (expect: rejected)request
node-eviction-with-commons✅ PASS80% (n=5)0.90fired 0/5 · short-circuit 0/5 (expect: rejected)report-worker, request
poisoned-recall-rejected⚠️ FLAKY60% (n=5)0.90eval-victim, pull, v9.9.9
poisoned-recall-verify⚠️ FLAKY40% (n=5)1.00fired 5/5 · short-circuit 0/5 (expect: withdrawn)eval-victim, v9.9.9

Cost per investigation

Median provider-reported tokens per case on openai/glm-4.5-air, priced at $0.20/MTok in · $1.10/MTok out. Replay evidence, so tool latency and live-cluster variance are excluded.

pathcasesmedian in tokmedian out tokest. cost
full investigation632.7k1.9k$0.009

Confidence calibration

  • Confidently wrong (missed with median confidence ≥ 0.70): 4 — harbor-chart-bump, node-eviction-no-commons, poisoned-recall-rejected, poisoned-recall-verify
  • Underconfident (reached with median confidence < 0.50): none

History

Newest first, last 30 shown — the full log is history.jsonl. A run that reached no answer to score at all is labelled in the pass-rate column instead of scored — it is not a 0%.

datemodelreachedpass-rateest. cost
2026-09-16T11:00:02Zopenai/glm-4.5-air2/633%$0.31
2026-09-15T11:08:25Zopenai/glm-4.5-air2/633%$0.32
2026-09-14T11:49:15Zopenai/glm-4.5-air3/650%$0.27
2026-09-13T11:13:48Zopenai/glm-4.5-air1/617%$0.30
2026-09-12T10:11:24Zopenai/glm-4.5-air2/633%$0.32
2026-09-11T10:42:36Zopenai/glm-4.5-air3/650%$0.29
2026-09-10T10:41:10Zopenai/glm-4.5-air2/633%$0.34
2026-09-09T11:04:57Zopenai/glm-4.5-air2/633%$0.31
2026-09-08T10:47:50Zopenai/glm-4.5-air3/650%$0.34
2026-09-07T11:40:44Zopenai/glm-4.5-air2/633%$0.31
2026-09-06T10:30:32Zopenai/glm-4.5-air2/633%$0.35
2026-09-05T10:11:28Zopenai/glm-4.5-air2/633%$0.33
2026-09-04T10:23:29Zopenai/glm-4.5-air⚠️ errored
2026-09-03T10:33:32Zopenai/glm-4.5-air⚠️ errored
2026-09-02T10:25:45Zopenai/glm-4.5-air⚠️ errored
2026-09-01T11:13:35Zopenai/glm-4.5-air2/633%$0.35
2026-08-31T12:45:35Zopenai/glm-4.5-air2/633%$0.32
2026-08-30T11:21:46Zopenai/glm-4.5-air2/633%$0.30
2026-08-29T12:25:38Zopenai/glm-4.5-air3/650%$0.32
2026-08-28T18:23:22Zopenai/glm-4.5-air1/617%$0.32
2026-08-27T17:34:55Zopenai/glm-4.5-air3/650%$0.34
2026-08-26T07:04:32Zopenai/glm-4.5-air2/633%$0.32
2026-08-25T06:59:22Zopenai/glm-4.5-air1/617%$0.30
2026-08-24T07:07:05Zopenai/glm-4.5-air2/633%$0.35
2026-08-23T06:50:15Zopenai/glm-4.5-air2/633%$0.28
2026-08-22T06:54:44Zopenai/glm-4.5-air2/633%$0.32
2026-08-21T07:04:23Zopenai/glm-4.5-air2/633%$0.28
2026-08-20T07:02:55Zopenai/glm-4.5-air1/617%$0.28
2026-08-19T06:58:02Zopenai/glm-4.5-air3/650%$0.28
2026-08-18T06:58:03Zopenai/glm-4.5-air2/633%$0.31