agentsclimarketplace

Memory consolidation

Skill AnthonyAlcaraz/agentic-graph-rag-skills/skills/memory/memory-consolidation

Consolidation pipeline — turn noisy raw episodes into durable knowledge (Agentic GraphRAG Ch4, Example 4-5 + Example 4-13). Four steps: cluster related episodes by topic, summarize each cluster into one consolidated fact, create the consolidated node, and maintain a provenance chain back to the source episodes so "how do you know that?" is answerable. Clusters below a minimum size (default 3) are skipped — not enough examples to generalize. Adds the sleep-time-compute discipline: run consolidation during idle periods, never on the synchronous response path, and pre-compute inferences that anticipate likely queries. Use when an agent accumulates redundant, overlapping experiences that must compress into stable, queryable patterns. NOT for one-shot agents (nothing to consolidate), NOT for the response hot path (consolidation is a background/idle job), NOT for facts that must stay individually addressable (consolidation merges them).From its SKILL.md

Install
npx -y skills add AnthonyAlcaraz/agentic-graph-rag-skills --skill memory-consolidation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 5 stars5 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
  • runs commandsInstructs the agent to run 2 commands, including `python cli.py benchmark` and 1 more.

SKILL.md

8.2 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it

Memory Consolidation

Overview

Your agent accumulates many interactions, but most are redundant, overlapping, or partially inconsistent. Consolidation is the agent's "sleep phase": it compresses short-term experiences into long-term understanding. Ch4 frames this as four steps (Example 4-5):

  1. Cluster related memories (cluster_by_topic) — group conversations about the same project/incident by similarity.
  2. Summarize each cluster (summarize_cluster) — replace "Monday: deadline Friday; Tuesday: confirmed Friday; Wednesday: Friday again" with one fact: "Project deadline: Friday (confirmed 3 times)". Meaning preserved, redundancy gone.
  3. Consolidate into permanent graph nodes (create_consolidated_memory).
  4. Maintain provenance (maintain_provenance_chain) — keep the DERIVED_FROM links so the agent can trace a belief back to the exact interactions that produced it.

Two disciplines from the chapter shape the implementation:

  • Minimum cluster size (Example 4-13: "Need enough examples to generalize"). A pattern derived from a single episode is not a pattern. Default minimum is 3.
  • Sleep-time compute (Letta + UC Berkeley, Ch4): consolidation runs during idle periods, not while a user waits. Shifting it off the response path cuts active inference cost ~5x and lets you pre-compute inferences (which tasks are at risk, given a deadline + dependencies) before anyone asks.

When to Use

  • An agent with accumulating episodic memory that grows noisy over time
  • DevOps incident memory: turn many similar 503-after-deploy incidents into one durable Pattern node with a runbook
  • Conversational assistants that repeat the same fact across sessions
  • Any system that needs to answer "how do you know X?" with a provenance trace

Phrases: "consolidate memory", "summarize episodes", "sleep-time compute", "provenance chain", "compress experience into knowledge", "cluster incidents".

When NOT to Use

  • One-shot / stateless agents — there is nothing to consolidate
  • The synchronous response path — consolidation is a background/idle job
  • Facts that must remain individually addressable (an audit log of distinct events) — consolidation deliberately merges them
  • Clusters that never reach the minimum size — keep the raw episodes; do not fabricate a pattern from one example

Process

StepInputActionOutputVerification
1list of raw episodeslib.cluster_by_topic(episodes, threshold)list of clustersrelated episodes group; unrelated ones split
2one clusterlib.summarize_cluster(cluster)one durable fact stringmulti-episode summary carries a confirmation count
3all episodeslib.consolidate(episodes, min_cluster_size)list of ConsolidatedFactclusters below min size are skipped
4a fact + episodeslib.provenance_of(fact, episodes)source episodesround-trips; dangling links raise
5facts + idle workerlib.precompute_inferences(facts, inference_fn){fact_id: inferences}runs off the response path

CLI: cluster, consolidate, scenario incident-consolidation, benchmark.

Rationalizations

Agent rationalizationDocumented rebuttal
"I'll consolidate on every turn so memory is always fresh."Consolidation is expensive (clustering + summarization). The chapter's sleep-time-compute result is explicit: run it during idle periods, not the response path — ~5x cheaper active inference. On-turn consolidation pays the cost when the user is waiting.
"One vivid episode is enough to make a pattern."Example 4-13 sets if len(cluster) < 3: continue. A pattern from one example is an overfit. The minimum-size gate is part of the design, not a tuning knob.
"I'll drop the provenance links to save space."Then "how do you know the deadline is Friday?" becomes unanswerable and you cannot debug a wrong conclusion. The DERIVED_FROM chain is what makes consolidated knowledge trustworthy and inspectable.
"Sentence-level facts are simpler to extract."Sentence-level extraction fragments cross-turn context (Ch4). HINDSIGHT extracts 2-5 narrative facts per conversation. The summarizer_fn seam is where you swap the extractive default for the narrative LLM form.

Red Flags

  • Every episode lands in its own singleton cluster. The similarity threshold is too high, or the similarity function is mis-tuned for the domain (the default token-overlap may need an embedding swap).
  • One giant cluster swallows everything. Threshold too low — unrelated episodes are being glued together. Tighten it or switch to semantic similarity.
  • Consolidated facts have empty derived_from. The provenance chain was dropped — the fact is now unexplainable and undebuggable.
  • Consolidation appears in the request latency profile. It is running on the response path. Move it to a background/idle worker.

Non-Negotiable Verification

  1. Run the benchmark battery. python cli.py benchmark must report 10/10: related episodes cluster, the min-size gate skips small clusters, full provenance is recorded and round-trips, dangling provenance raises, clustering is deterministic, and sleep-time precompute keys by fact id.
  2. Verify CLI help. python cli.py --help exits 0 and prints the SKILL.md description.

Security Posture

  • Prompt injection. Episodes are untrusted content, and repetition is the attack: planting the same false claim >= min-cluster-size times manufactures a "confirmed 3 times" consolidated fact. Clustering never executes episode text; the provenance chain is the audit path - inspect sources before trusting a consolidated node.
  • Data exfiltration. Consolidation copies episode content into durable nodes: sensitive details survive in summaries and DERIVED_FROM links long after raw episodes age out. Redact before consolidating; the skill makes no network calls and writes no files itself.
  • Privilege escalation. No shell invocation, no eval. Consolidated facts gain durable, trusted standing in the graph, and sleep-time jobs run unattended - bound the background job's write scope and gate what gets promoted from episode to knowledge.

Source Attribution

Distilled from Agentic GraphRAG (O'Reilly) by Anthony Alcaraz and Sam Julien — Ch4: Memory — "Consolidation: From Experience to Knowledge" (Example 4-5) and "Adding Memory to the DevOps Agent" (Example 4-13). Sleep-time compute: Letta + UC Berkeley research cited in Ch4. Narrative-fact extraction: HINDSIGHT (Latimer et al. 2025).

What ships with it: 2 files

18.5 KB alongside SKILL.md, 2 of them executable

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.