Skip to content

RunLore vs HolmesGPT

Claims about HolmesGPT below were checked against the project’s own repository and documentation on 2026-08-02. If HolmesGPT has shipped something new since — especially Flux support, which was an open pull request at the time of writing — please open an issue.

HolmesGPT is the strongest open-source investigation agent and the one to beat on breadth: 56 built-in toolsets, a CNCF Sandbox home, and a real contributor base. RunLore does not try to out-toolset it — the difference is what happens after the investigation.

What HolmesGPT does better

  • A much broader native toolset. 56 built-in toolsets as of this check, spanning Datadog, New Relic, Splunk, Sentry, Elasticsearch/OpenSearch, VictoriaMetrics, and more — see the full list. RunLore’s native set is a fraction of that (see below).
  • A CNCF Sandbox project with a real contributor base. HolmesGPT was accepted into the CNCF Sandbox in October 2025 and has 71 contributors on GitHub as of this check. RunLore, today, is maintained by one person. That is a materially different bus factor and support surface — if you need a project that survives its founder losing interest, that matters.
  • Approval-gated remediation, shipped and versioned. Since v0.34 (June 2026), HolmesGPT’s Kubernetes Remediation toolset issues signed (JWT) approval tickets before a write action executes — a real, shipped human-in-the-loop gate for actions, not a design sketch.
  • Backed by a company, with outside contributions. Originally built by Robusta.Dev, with contributions from Microsoft; Apache-2.0 licensed, ~3,000 GitHub stars, and 28 pull requests merged in the 30 days before this check. RunLore has no company behind it.

What RunLore does that HolmesGPT does not

  • HolmesGPT doesn’t learn across incidents; RunLore does. HolmesGPT’s skills (formerly runbooks) are human-authored, static troubleshooting guides — the docs describe no mechanism for the agent to generate or refine them from what it observes. Every incident starts from the same skill catalog. RunLore curates each verified finding into a pull request against a Git repo you own; once merged, the same incident is answered from the catalog in milliseconds next time, and a note’s trust decays if it stops matching real-world outcomes. See the learning loop.
  • Flux support, and a revision-exact “what changed.” HolmesGPT has an ArgoCD toolset that can list deployment history and diff an application’s live state against its desired Git state — but as of this check there is no built-in Flux toolset (an open PR adds one, unmerged), and no tool that renders the diff between two already-reconciled revisions — the “what changed in the last deploy” view. RunLore’s what_changed tool does exactly that, for both Flux and Argo CD, scoped to the failing workload.
  • A published, continuous eval scorecard. HolmesGPT runs a weekly automated regression suite (150+ evals, Sundays) plus manual full benchmarks, with results archived by date. RunLore’s nightly eval workflow is designed to publish a per-scenario scorecard — pass/fail, recall outcomes, confidence calibration, model, and cost — on every run, red or green, not just on a schedule; see Benchmarking.

At a glance

RunLoreHolmesGPT
Investigation breadth7 native providers (data sources) + any MCP server56 built-in toolsets
Learning loopYes — outcome-weighted recall, decays on real resolve-rateNo — static, human-authored skills
Knowledge portabilityPlain markdown in your Git repo, OKF-compatibleNot applicable — no auto-generated knowledge to port
Human review gateEvery KB entry is a human-merged pull requestNo knowledge to review; remediation actions are approval-gated (since v0.34)
GitOps “what changed”Revision-exact diff, Flux + Argo CD (data sources)Argo CD only (live-vs-desired diff + history, no Flux, no between-revisions diff)
Autonomy ceilingRead-only by default; auto executes only suspend/resume/reconcileRead-only investigation; signed-ticket approval-gated remediation since v0.34
Published eval numbersNightly workflow, every run, pass/fail + cost + calibration (details)Weekly regression + manual full benchmarks, 150+ evals, archived
Project maturitySolo-maintained, pre-1.0CNCF Sandbox (Oct 2025), Apache-2.0, 71 contributors, backed by Robusta.Dev
MCP clientYes — extends the agent’s toolbox (MCP); also serves as an MCP server itselfYes — remote MCP servers plus MCP-based built-in toolsets

Toolset breadth, honestly

RunLore’s native tool set is narrower by design — seven pluggable providers (GitOps, metrics, logs, network, cloud, Kubernetes, source repos) rather than a toolset per vendor. That is a real gap against HolmesGPT’s 56. Any specific observability vendor RunLore doesn’t natively speak, an MCP server closes — RunLore is an MCP client as well as a server, so a Datadog or Splunk MCP server extends its investigation toolbox with zero Go code. It is not the same as a maintained, first-party integration, and we’re not pretending otherwise.

Pick HolmesGPT instead if…

  • You need a specific vendor toolset today — Datadog, New Relic, Splunk, and dozens more ship built-in, with no MCP server to stand up first.
  • You want a CNCF-governed project with a large, active contributor base and a company behind it.
  • You don’t want a knowledge-review workflow at all — you just want an agent that investigates and reports, every time, with no PRs to merge.

Sources