Matrix
What it gives you โ findings delivered to any Matrix room, with an optional zero-ingress
๐/๐ feedback loop over reactions, and a ๐ reaction (or a silence: thread command) to silence a
known incident.
Minimal config
notify:
matrix:
homeserver: https://matrix.org
room_id: "!yourroom:matrix.org"
access_token_env: MATRIX_TOKENVerify it locally
kubectl -n runlore create secret generic runlore-secrets \
--from-literal=MATRIX_TOKEN='<matrix-access-token>'Fire a test incident (see Alertmanager or hack/demo.sh) and
confirm delivery:
kubectl -n runlore logs deploy/runlore | grep 'msg=findings'Notes
- Matrix delivers the same content as Slack’s incoming-webhook path: a single message per finding (no threading).
matrix.feedback_reactions(opt-in, defaultfalse) โ react ๐/๐ to a RunLore message and the rating lands in the outcome ledger, with the same per-user dedup and trust weighting as Slack’s feedback buttons.- Nothing is exposed to enable it โ reactions arrive over the client-server
/synclong-poll, an outbound HTTPS request authenticated by the notifier’s own access token. This is the zero-ingress alternative to Slack’s Interactivity Request URL. - The listener runs on the leader only, skips reactions from before startup, ignores every emoji
except ๐/๐/๐, and only counts votes (and silences) on messages the bot itself sent
(attribution is anchored on
/whoami). - Startup fails loud unless
homeserver/room_id/access_token_envandoutcome.ledger_pathare set. - Use an invite-only room โ any room member can vote, and any room member can silence.
feedback_reactions,silence_reactionsandthread_captureare independently-gated capabilities sharing ONE listener. Enabling any of them starts the same/synclong-poll and the same leader-only goroutine โ there is one Matrix listener and one/syncconnection per process, no matter how many of the three are on. What differs per flag is which event types that shared/syncfilter requests and which handler actually acts on them:feedback_reactionsandsilence_reactionsboth ask form.reaction(๐/๐ and ๐ are the same event type on the wire, and are told apart only oncehandleReactionreads which emoji it was โ each still gated on its own flag, sosilence_reactionsoff means a ๐ reaction is read and dropped, never recorded); enabling either alone is therefore enough to receive the event, but recording a silence additionally needssilence_reactionsitself on.thread_capturealone asks form.room.message. Turning one flag on never turns another’s handling on โ each is gated on its own flag in code, not just requested-or-not on the wire โ but it does mean a homeserver hiccup or a leadership change pauses whichever capabilities are enabled together, since they ride the same poll loop.
Silence a recurring incident
notify.matrix.silence_reactions (opt-in, default false) turns on a ๐ reaction; a second
path, the silence: <duration> thread command, needs thread_capture too. Both land in the same
outcome ledger as Slack’s silence_button:
- A ๐ reaction on a RunLore investigation message โ the zero-ingress equivalent of Slack’s
๐ menu. A bare reaction carries no duration, so it always silences for
notify.silence.windows[0](the FIRST configured preset) โ there is no way to pick a longer window from a reaction alone.silence_reactionsalone is enough: the/syncfilter already carriesm.reactionfor feedback voting, so no other flag is required. - A
silence: <duration>thread command โ reply in the investigation thread with, for examplesilence: 4h, and RunLore silences for exactly that duration (up tonotify.silence.max_window). Thesilence:token is matched case-insensitively (SILENCE: 4hworks) and must start your message, once the bot mention is stripped. That is stricter thannote:andreinvestigate:, which are recognised anywhere in a message, and deliberately so: acting on this one writes โ it records a ledger event and switches investigation off โ so a sentence that merely mentions it (note: we agreed on silence: 4h) is answered with a request to rephrase rather than silently muting the incident and discarding the note. A sentence using the word without a trailing colon is not a command at all. This path needs nomodel.chatconfiguration and costs no model call: parsing a duration is deterministic. Unlike the reaction, it also needsnotify.matrix.thread_capture: trueโ a plain message only reaches RunLore at all once the/syncfilter widens tom.room.message, whichsilence_reactionson its own does not request. That in turn means the same forge/registry/replier prerequisites Write knowledge back from a thread below documents forthread_captureapply here too, even though recording a silence never touches the forge itself: with any of those three missing,thread_capturesilently degrades to off and the command is never parsed. Eithersilence_reactionsor Slack’ssilence_buttonsatisfies the sharednotify.silencerequirements (outcome.ledger_path, non-emptywindows, a positivemax_window) the command also depends on.
Either path suppresses re-investigating this exact TriggerKey โ no model call, no notification,
no ledger open โ until the window expires, the incident fires again as CRITICAL, a standing
๐ re-arms it (newest human wins: a ๐ cast after the silence lifts it, a ๐ clicked after the
newest ๐ still suppresses), or the incident resolves. The silence: thread command is
acknowledged with a message naming who silenced it, until when, and restating the escapes that apply
to this deployment. A bare ๐ reaction is recorded and logged but gets no reply โ a reaction has
no thread to answer in, so check runlore_investigations_completed_total{result="silenced"} or the
msg="matrix silence recorded" log line if you need confirmation.
Two of those four are alert-specific, exactly as on the Slack side (see Slack โ Silence a
recurring incident
for the full explanation): a GitOps-sourced failure sets no severity, so the CRITICAL carve-out never
fires, and its fingerprint is synthetic with no resolve channel, so the resolve escape never fires
either. For a GitOps failure, a silence is bounded by its expiry and a ๐ โ and nothing else. The
same loss of the resolve escape applies to an Alertmanager receiver configured with
send_resolved: false.
A third is config-specific. The ๐ escape only exists where a ๐/๐ control is actually enabled
somewhere: with silence_reactions: true and feedback_reactions: false, handleReaction drops a
๐, so nobody can cast one. Combine that with a GitOps trigger and the expiry is the only bound left.
RunLore warns at startup when silencing is on with no ๐/๐ control on any transport; votes and
silences share one outcome ledger, so Slack’s feedback_buttons satisfies it too.
notify:
silence:
windows: [1h, 4h, 24h]
max_window: 24h
matrix:
homeserver: https://matrix.org
room_id: "!yourroom:matrix.org"
access_token_env: MATRIX_TOKEN
silence_reactions: true
outcome:
ledger_path: /var/lib/runlore/outcome.jsonlLike Matrix’s feedback reactions, silencing is deliberately unprivileged โ any member of the
room can silence an incident, with no allowlist โ which is safe for the same reason Slack’s button
is: for an alert-sourced incident with a ๐/๐ control enabled the blast radius is bounded by the
four escapes above (narrower for a GitOps failure, a send_resolved: false receiver, or a deployment
with no ๐/๐ control anywhere, per above), every silence is attributed to the sender’s Matrix id in
the outcome ledger, and the window is capped by notify.silence.max_window.
Use an invite-only room, as for feedback reactions.
silence_reactions is gated independently of feedback_reactions โ enabling it on its own
records ๐ silences without recording ๐/๐ votes, and vice versa. Turning silencing on without
either transport’s ๐/๐ control removes the ๐ escape entirely; startup warns when that is the
configuration.
Write knowledge back from a thread
With thread_capture on, replying inside an investigation thread records what you know into the
knowledge base:
@runlore note: the real cause was the spot-node reclaim at 14:02
RunLore recognises being addressed via MSC3952 m.mentions when your client sends it, or โ as a
fallback โ the bot’s full Matrix ID or its localpart (runlore in @runlore:example.org)
appearing in the message body as a whole word. The word boundary is not decoration: a plain
substring match would read the localpart sre inside “misread” as addressing the bot. A reply is attributed to its thread via the MSC3440
m.thread relation, falling back to m.in_reply_to for clients that only send the legacy
non-threaded reply fallback.
The note lands as a comment on that finding’s knowledge-base PR. When the finding has no PR โ an
instant recall, or a no_action verdict โ RunLore opens a small Concept entry PR instead, so the
knowledge still lands somewhere. A human reviews and merges it, like every other entry.
@runlore reinvestigate: โฆ is reserved and not supported yet; add the reinvestigate label to the
knowledge-base issue to re-run an investigation.
notify:
matrix:
homeserver: https://matrix.org
room_id: "!yourroom:matrix.org"
access_token_env: MATRIX_TOKEN
thread_capture: true
outcome:
ledger_path: /var/lib/runlore/outcome.jsonl # required: the thread registry lives beside itthread_capture: true also requires outcome.ledger_path โ the same registry Slack’s thread
capture uses (one registry, shared by both transports): see Configuration โ
notify for the
full persistence breakdown. Matrix degrades more gracefully than Slack when the registry itself
is unavailable (past its TTL, evicted past the size cap, or orphaned by a restart or leader
failover onto a replica that never saw it): every investigation message RunLore posts also carries
its own thread identity stamped directly on the Matrix event. A reply still resolves against that
stamp โ fetched straight from the homeserver and re-verified as sent by RunLore itself โ so a
registry miss answers the note instead of refusing it with “no context for this thread.” What the
registry alone tracks (the standalone note PR a thread already opened, and how many notes it has
used against the per-thread cap) is unavailable on a fallback resolution, so a note recorded this
way opens a fresh Concept PR rather than commenting on one an earlier note in the same thread
opened. Slack has no per-event stamp to fall back to, so a registry miss there is unrecoverable โ
see Configuration’s restart/leader-failover warning for what that costs.
Unlike Slack, this needs no exposed HTTP endpoint, no ingress change and no new permission.
Mentions arrive over the same outbound /sync long-poll RunLore already runs for
feedback_reactions โ turning thread_capture on only widens what that poll asks the homeserver
for. That widening is real and worth understanding before you flip the flag: see Security model โ
Matrix thread
capture.
Announce a knowledge write
A captured note is reported into the thread it came from, and nowhere else โ so the
knowledge lands, but nobody outside that thread learns it did. announce_kb_updates
changes that: every write that reaches the forge is also announced, with its own
message naming who wrote the note and where.
notify:
matrix:
homeserver: https://matrix.org
room_id: "!yourroom:matrix.org"
access_token_env: MATRIX_TOKEN
thread_capture: true
thread:
announce_kb_updates: channel # default false; also true (= channel), thread, both
outcome:
ledger_path: /var/lib/runlore/outcome.jsonl # required by thread_capture, as aboveThe post names the pull request, the entry that was created, who the note came from
and which chat system they typed it in, and quotes the note itself โ capped at
512 bytes, with the pull request holding the full text. It carries the same
secret-redacted, capped text that reached the forge, and everything a human or the
model wrote in it is neutralised before sending: an @room inside a note is rendered
readable but inert, never a room-wide notification. That holds wherever it lands.
Where it lands is yours to pick. channel (what true has always meant) sends a
plain m.notice with no m.thread relation to the room this notifier delivers
findings to. thread sends it into the thread the note was typed in, as an
m.notice carrying the MSC3440 m.thread relation. both does each.
channelproduces two messages for one write, and with several sinks that is the point. The thread reply answers the person who typed; the room post reaches everyone who was not reading that thread. But with Matrix as your only transport those are the same people โ the thread lives in that very room โ so each write restates in the room what the thread just said. That is whatthreadis for: the announcement goes into the thread instead, keeping the provenance line the reply does not carry, without the echo.threadnever silences your other sinks. A Slack channel, awebhookendpoint โ neither can post into a Matrix thread, so they each receive the announcement in their own channel exactly as before. The echo only exists where the thread and the room are the same place.- The announcement carries note content. A note written in a thread nobody else was
watching is broadcast to every sink you have configured โ this room, a Slack channel
if you run both, any
webhookendpoint. That is your decision to take knowingly, which is why the key is off unless you set it.
Conversational replies
Without model.chat, anything you say in the thread that is not note: gets a fixed reply telling
you so, and nothing is recorded. Adding a model.chat block changes that: RunLore answers the
question instead, from the investigation’s own evidence plus the top knowledge-base hits, and files a
note itself when your message contained a durable fact.
model:
chat:
model: claude-haiku-4-5 # required โ never inherited from model.model
notify:
matrix:
homeserver: https://matrix.org
room_id: "!yourroom:matrix.org"
access_token_env: MATRIX_TOKEN
thread_capture: true # the chat layer has no channel without it
thread:
chat_calls_per_hour: 30 # default 30
chat_tokens_per_hour: 109480 # default 109480 โ derived, not round
outcome:
ledger_path: /var/lib/runlore/outcome.jsonl # required by thread_capture, as aboveThis is a paid path any member of the room can trigger, and on Matrix “addressed” is looser than
you may expect. RunLore treats a message as addressed to it via MSC3952 m.mentions โ but also
when the bot’s full MXID or its bare localpart (runlore in @runlore:example.org) appears in
the body as a whole word. No mention entity is required. So anyone in the room can trigger a
model call by typing the bot’s name in a reply to an investigation thread. Keep the room invite-only.
With model.chat set, every addressed message rooted in one of RunLore’s own messages that is not a
recognised command costs exactly one model call โ one structurally, not on average: the model is
given a single forced tool and no search tool, so there is no agent loop.
@runlore note: โฆ still costs nothing โ it is the same deterministic write it always was โ and a
bare mention with nothing after it costs nothing either.
Spend is bounded by two global hourly ceilings (chat_calls_per_hour, chat_tokens_per_hour) shared
across every thread and both transports. The token default is derived โ what one maxed-out call can
cost, times the call budget โ which is why it is not a round number and why it moves when
max_note_bytes moves. It is not bounded in currency: model.pricing is a reporting table, not
a limit, and nothing compares a cost to a threshold and stops. Read
Conversational replies and what they cost
before enabling it โ it states the full set of bounds and, just as plainly, what has none.
Reference
- Configuration โ
notifyfor the full key reference. - Security model โ the feedback channels for the exposure and vote trust model.