{"slug":"agent-memory-systems","title":"agent-memory-systems","summary":"\"Memory is the cornerstone of intelligent agents. Without it, every","platform":"ChatGPT","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-08-16T13:38:13.053522Z","repo":{"url":"https://github.com/sickn33/agentic-awesome-skills","stars":46883,"forks":6831,"license":"MIT","updatedAt":"2026-09-25T05:43:16Z"},"bodyHtml":"<hr>\n<h2>name: agent-memory-systems\ndescription: \"Memory is the cornerstone of intelligent agents. Without it, every\ninteraction starts from zero. This skill covers the architecture of agent\nmemory: short-term (context window), long-term (vector stores), and the\ncognitive architectures that organize them.\"\nrisk: safe\nsource: vibeship-spawner-skills (Apache 2.0)\ndate_added: 2026-02-27</h2>\n<h1>Agent Memory Systems</h1>\n<p>Memory is the cornerstone of intelligent agents. Without it, every interaction\nstarts from zero. This skill covers the architecture of agent memory: short-term\n(context window), long-term (vector stores), and the cognitive architectures\nthat organize them.</p>\n<p>Key insight: Memory isn't just storage - it's retrieval. A million stored facts\nmean nothing if you can't find the right one. Chunking, embedding, and retrieval\nstrategies determine whether your agent remembers or forgets.</p>\n<p>The field is fragmented with inconsistent terminology. We use the CoALA cognitive\narchitecture framework: semantic memory (facts), episodic memory (experiences),\nand procedural memory (how-to knowledge).</p>\n<h2>Principles</h2>\n<ul>\n<li>Memory quality = retrieval quality, not storage quantity</li>\n<li>Chunk for retrieval, not for storage</li>\n<li>Context isolation is the enemy of memory</li>\n<li>Right memory type for right information</li>\n<li>Decay old memories - not everything should be forever</li>\n<li>Test retrieval accuracy before production</li>\n<li>Background memory formation beats real-time</li>\n</ul>\n<h2>Capabilities</h2>\n<ul>\n<li>agent-memory</li>\n<li>long-term-memory</li>\n<li>short-term-memory</li>\n<li>working-memory</li>\n<li>episodic-memory</li>\n<li>semantic-memory</li>\n<li>procedural-memory</li>\n<li>memory-retrieval</li>\n<li>memory-formation</li>\n<li>memory-decay</li>\n</ul>\n<h2>Scope</h2>\n<ul>\n<li>vector-database-operations → data-engineer</li>\n<li>rag-pipeline-architecture → llm-architect</li>\n<li>embedding-model-selection → ml-engineer</li>\n<li>knowledge-graph-design → knowledge-engineer</li>\n</ul>\n<h2>Tooling</h2>\n<h3>Memory_frameworks</h3>\n<ul>\n<li>LangMem (LangChain) - When: LangGraph agents with persistent memory Note: Semantic, episodic, procedural memory types</li>\n<li>MemGPT / Letta - When: Virtual context management, OS-style memory Note: Hierarchical memory tiers, automatic paging</li>\n<li>Mem0 - When: User memory layer for personalization Note: Designed for user preferences and history</li>\n</ul>\n<h3>Vector_stores</h3>\n<ul>\n<li>Pinecone - When: Managed, enterprise-scale (billions of vectors) Note: Best query performance, highest cost</li>\n<li>Qdrant - When: Complex metadata filtering, open-source Note: Rust-based, excellent filtering</li>\n<li>Weaviate - When: Hybrid search, knowledge graph features Note: GraphQL interface, good for relationships</li>\n<li>ChromaDB - When: Prototyping, small/medium apps Note: Developer-friendly, ~20ms p50 at 100K vectors</li>\n<li>pgvector - When: Already using PostgreSQL, simpler setup Note: Good for &lt;1M vectors, familiar tooling</li>\n</ul>\n<h3>Embedding_models</h3>\n<ul>\n<li>OpenAI text-embedding-3-large - When: Best quality, 3072 dimensions Note: $0.13/1M tokens</li>\n<li>OpenAI text-embedding-3-small - When: Good balance, 1536 dimensions Note: $0.02/1M tokens, 5x cheaper</li>\n<li>nomic-embed-text-v1.5 - When: Open-source, local deployment Note: 768 dimensions, good quality</li>\n<li>all-MiniLM-L6-v2 - When: Lightweight, fast local embedding Note: 384 dimensions, lowest latency</li>\n</ul>\n<h2>Patterns</h2>\n<h3>Memory Type Architecture</h3>\n<p>Choosing the right memory type for different information</p>\n<p><strong>When to use</strong>: Designing agent memory system</p>\n<h1>MEMORY TYPE ARCHITECTURE (CoALA Framework):</h1>\n<p>\"\"\"\nThree memory types for different purposes:</p>\n<ol>\n<li><p>Semantic Memory: Facts and knowledge</p>\n<ul>\n<li>What you know about the world</li>\n<li>User preferences, domain knowledge</li>\n<li>Stored in profiles (structured) or collections (unstructured)</li>\n</ul>\n</li>\n<li><p>Episodic Memory: Experiences and events</p>\n<ul>\n<li>What happened (timestamped events)</li>\n<li>Past conversations, task outcomes</li>\n<li>Used for learning from experience</li>\n</ul>\n</li>\n<li><p>Procedural Memory: How to do things</p>\n<ul>\n<li>Rules, skills, workflows</li>\n<li>Often implemented as few-shot examples</li>\n<li>\"How did I solve this before?\"\n\"\"\"</li>\n</ul>\n</li>\n</ol>\n<h2>LangMem Implementation</h2>\n<p>\"\"\"\nfrom langmem import MemoryStore\nfrom langgraph.graph import StateGraph</p>\n<h1>Initialize memory store</h1>\n<p>memory = MemoryStore(\nconnection_string=os.environ[\"POSTGRES_URL\"]\n)</p>\n<h1>Semantic memory: user profile</h1>\n<p>await memory.semantic.upsert(\nnamespace=\"user_profile\",\nkey=user_id,\ncontent={\n\"name\": \"Alice\",\n\"preferences\": [\"dark mode\", \"concise responses\"],\n\"expertise_level\": \"developer\",\n}\n)</p>\n<h1>Episodic memory: past interaction</h1>\n<p>await memory.episodic.add(\nnamespace=\"conversations\",\ncontent={\n\"timestamp\": datetime.now(),\n\"summary\": \"Helped debug authentication issue\",\n\"outcome\": \"resolved\",\n\"key_insights\": [\"Token expiry was root cause\"],\n},\nmetadata={\"user_id\": user_id, \"topic\": \"debugging\"}\n)</p>\n<h1>Procedural memory: learned pattern</h1>\n<p>await memory.procedural.add(\nnamespace=\"skills\",\ncontent={\n\"task_type\": \"debug_auth\",\n\"steps\": [\"Check token expiry\", \"Verify refresh flow\"],\n\"example_interaction\": few_shot_example,\n}\n)\n\"\"\"</p>\n<h2>Memory Retrieval at Runtime</h2>\n<p>\"\"\"\nasync def prepare_context(user_id, query):\n# Get user profile (semantic)\nprofile = await memory.semantic.get(\nnamespace=\"user_profile\",\nkey=user_id\n)</p>\n<pre><code># Find relevant past experiences (episodic)\nsimilar_experiences = await memory.episodic.search(\n    namespace=\"conversations\",\n    query=query,\n    filter={\"user_id\": user_id},\n    limit=3\n)\n\n# Find relevant skills (procedural)\nrelevant_skills = await memory.procedural.search(\n    namespace=\"skills\",\n    query=query,\n    limit=2\n)\n\nreturn {\n    \"profile\": profile,\n    \"past_experiences\": similar_experiences,\n    \"relevant_skills\": relevant_skills,\n}\n</code></pre>\n<p>\"\"\"</p>\n<h3>Vector Store Selection Pattern</h3>\n<p>Choosing the right vector database for your use case</p>\n<p><strong>When to use</strong>: Setting up persistent memory storage</p>\n<h1>VECTOR STORE SELECTION:</h1>\n<p>\"\"\"\nDecision matrix:</p>\n<p>|            | Pinecone | Qdrant | Weaviate | ChromaDB | pgvector |\n|------------|----------|--------|----------|----------|----------|\n| Scale      | Billions | 100M+  | 100M+    | 1M       | 1M       |\n| Managed    | Yes      | Both   | Both     | Self     | Self     |\n| Filtering  | Basic    | Best   | Good     | Basic    | SQL      |\n| Hybrid     | No       | Yes    | Best     | No       | Yes      |\n| Cost       | High     | Medium | Medium   | Free     | Free     |\n| Latency    | 5ms      | 7ms    | 10ms     | 20ms     | 15ms     |\n\"\"\"</p>\n<h2>Pinecone (Enterprise Scale)</h2>\n<p>\"\"\"\nfrom pinecone import Pinecone</p>\n<p>pc = Pinecone(api_key=os.environ[\"PINECONE_API_KEY\"])\nindex = pc.Index(\"agent-memory\")</p>\n<h1>Upsert with metadata</h1>\n<p>index.upsert(\nvectors=[\n{\n\"id\": f\"memory-{uuid4()}\",\n\"values\": embedding,\n\"metadata\": {\n\"user_id\": user_id,\n\"timestamp\": datetime.now().isoformat(),\n\"type\": \"episodic\",\n\"content\": memory_text,\n}\n}\n],\nnamespace=namespace\n)</p>\n<h1>Query with filter</h1>\n<p>results = index.query(\nvector=query_embedding,\nfilter={\"user_id\": user_id, \"type\": \"episodic\"},\ntop_k=5,\ninclude_metadata=True\n)\n\"\"\"</p>\n<h2>Qdrant (Complex Filtering)</h2>\n<p>\"\"\"\nfrom qdrant_client import QdrantClient\nfrom qdrant_client.models import PointStruct, Filter, FieldCondition</p>\n<p>client = QdrantClient(url=\"http://localhost:6333\")</p>\n<h1>Complex filtering with Qdrant</h1>\n<p>results = client.search(\ncollection_name=\"agent_memory\",\nquery_vector=query_embedding,\nquery_filter=Filter(\nmust=[\nFieldCondition(key=\"user_id\", match={\"value\": user_id}),\nFieldCondition(key=\"type\", match={\"value\": \"semantic\"}),\n],\nshould=[\nFieldCondition(key=\"topic\", match={\"any\": [\"auth\", \"security\"]}),\n]\n),\nlimit=5\n)\n\"\"\"</p>\n<h2>ChromaDB (Prototyping)</h2>\n<p>\"\"\"\nimport chromadb</p>\n<p>client = chromadb.PersistentClient(path=\"./memory_db\")\ncollection = client.get_or_create_collection(\"agent_memory\")</p>\n<h1>Simple and fast for prototypes</h1>\n<p>collection.add(\nids=[str(uuid4())],\nembeddings=[embedding],\ndocuments=[memory_text],\nmetadatas=[{\"user_id\": user_id, \"type\": \"episodic\"}]\n)</p>\n<p>results = collection.query(\nquery_embeddings=[query_embedding],\nn_results=5,\nwhere={\"user_id\": user_id}\n)\n\"\"\"</p>\n<h3>Chunking Strategy Pattern</h3>\n<p>Breaking documents into retrievable chunks</p>\n<p><strong>When to use</strong>: Processing documents for memory storage</p>\n<h1>CHUNKING STRATEGIES:</h1>\n<p>\"\"\"\nThe chunking dilemma:</p>\n<ul>\n<li>Too large: Vector loses specificity</li>\n<li>Too small: Loses context</li>\n</ul>\n<p>Optimal chunk size depends on:</p>\n<ul>\n<li>Document type (code vs prose vs data)</li>\n<li>Query patterns (factual vs exploratory)</li>\n<li>Embedding model (each has sweet spot)</li>\n</ul>\n<p>General guidance: 256-512 tokens for most use cases\n\"\"\"</p>\n<h2>Fixed-Size Chunking (Baseline)</h2>\n<p>\"\"\"\nfrom langchain.text_splitter import RecursiveCharacterTextSplitter</p>\n<p>splitter = RecursiveCharacterTextSplitter(\nchunk_size=500,      # Characters\nchunk_overlap=50,    # Overlap prevents cutting sentences\nseparators=[\"\\n\\n\", \"\\n\", \". \", \" \", \"\"]  # Priority order\n)</p>\n<p>chunks = splitter.split_text(document)\n\"\"\"</p>\n<h2>Semantic Chunking (Better Quality)</h2>\n<p>\"\"\"\nfrom langchain_experimental.text_splitter import SemanticChunker\nfrom langchain_openai import OpenAIEmbeddings</p>\n<h1>Splits based on semantic similarity</h1>\n<p>splitter = SemanticChunker(\nembeddings=OpenAIEmbeddings(),\nbreakpoint_threshold_type=\"percentile\",\nbreakpoint_threshold_amount=95\n)</p>\n<p>chunks = splitter.split_text(document)\n\"\"\"</p>\n<h2>Structure-Aware Chunking (Documents with Hierarchy)</h2>\n<p>\"\"\"\nfrom langchain.text_splitter import MarkdownHeaderTextSplitter</p>\n<h1>Respect document structure</h1>\n<p>splitter = MarkdownHeaderTextSplitter(\nheaders_to_split_on=[\n(\"#\", \"Header 1\"),\n(\"##\", \"Header 2\"),\n(\"###\", \"Header 3\"),\n]\n)</p>\n<p>chunks = splitter.split_text(markdown_doc)</p>\n<h1>Each chunk has header metadata for context</h1>\n<p>\"\"\"</p>\n<h2>Contextual Chunking (Anthropic's Approach)</h2>\n<p>\"\"\"</p>\n<h1>Add context to each chunk before embedding</h1>\n<h1>Reduces retrieval failures by 35%</h1>\n<p>def add_context_to_chunk(chunk, document_summary):\ncontext_prompt = f'''\nDocument summary: </p>\n<pre><code>The following is a chunk from this document:\n{chunk}\n'''\nreturn context_prompt\n</code></pre>\n<h1>Embed the contextualized chunk, not raw chunk</h1>\n<p>for chunk in chunks:\ncontextualized = add_context_to_chunk(chunk, summary)\nembedding = embed(contextualized)\nstore(chunk, embedding)  # Store original, embed contextualized\n\"\"\"</p>\n<h2>Code-Specific Chunking</h2>\n<p>\"\"\"\nfrom langchain.text_splitter import Language, RecursiveCharacterTextSplitter</p>\n<h1>Language-aware splitting</h1>\n<p>python_splitter = RecursiveCharacterTextSplitter.from_language(\nlanguage=Language.PYTHON,\nchunk_size=1000,\nchunk_overlap=200\n)</p>\n<h1>Respects function/class boundaries</h1>\n<p>chunks = python_splitter.split_text(python_code)\n\"\"\"</p>\n<h3>Background Memory Formation</h3>\n<p>Processing memories asynchronously for better quality</p>\n<p><strong>When to use</strong>: You want higher recall without slowing interactions</p>\n<h1>BACKGROUND MEMORY FORMATION:</h1>\n<p>\"\"\"\nReal-time memory extraction slows conversations and adds\ncomplexity to agent tool calls. Background processing after\nconversations yields higher quality memories.</p>\n<p>Pattern: Subconscious memory formation\n\"\"\"</p>\n<h2>LangGraph Background Processing</h2>\n<p>\"\"\"\nfrom langgraph.graph import StateGraph\nfrom langgraph.checkpoint.postgres import PostgresSaver</p>\n<p>async def background_memory_processor(thread_id: str):\n# Run after conversation ends or goes idle\nconversation = await load_conversation(thread_id)</p>\n<pre><code># Extract insights without time pressure\ninsights = await llm.invoke('''\n    Analyze this conversation and extract:\n    1. Key facts learned about the user\n    2. User preferences revealed\n    3. Tasks completed or pending\n    4. Patterns in user behavior\n\n    Be thorough - this runs in background.\n\n    Conversation:\n    {conversation}\n''')\n\n# Store to long-term memory\nfor insight in insights:\n    await memory.semantic.upsert(\n        namespace=\"user_insights\",\n        key=generate_key(insight),\n        content=insight,\n        metadata={\"source_thread\": thread_id}\n    )\n</code></pre>\n<h1>Trigger on conversation end or idle timeout</h1>\n<p>@on_conversation_idle(timeout_minutes=5)\nasync def process_conversation(thread_id):\nawait background_memory_processor(thread_id)\n\"\"\"</p>\n<h2>Memory Consolidation (Like Sleep)</h2>\n<p>\"\"\"</p>\n<h1>Periodically consolidate and deduplicate memories</h1>\n<p>async def consolidate_memories(user_id: str):\n# Get all memories for user\nmemories = await memory.semantic.list(\nnamespace=\"user_insights\",\nfilter={\"user_id\": user_id}\n)</p>\n<pre><code># Find similar memories (potential duplicates)\nclusters = cluster_by_similarity(memories, threshold=0.9)\n\n# Merge similar memories\nfor cluster in clusters:\n    if len(cluster) &gt; 1:\n        merged = await llm.invoke(f'''\n            Consolidate these related memories into one:\n            {cluster}\n\n            Preserve all important information.\n        ''')\n        await memory.semantic.upsert(\n            namespace=\"user_insights\",\n            key=generate_key(merged),\n            content=merged\n        )\n        # Delete originals\n        for old in cluster:\n            await memory.semantic.delete(old.id)\n</code></pre>\n<p>\"\"\"</p>\n<h3>Memory Decay Pattern</h3>\n<p>Forgetting old, irrelevant memories</p>\n<p><strong>When to use</strong>: Memory grows large, retrieval slows down</p>\n<h1>MEMORY DECAY:</h1>\n<p>\"\"\"\nNot all memories should live forever:</p>\n<ul>\n<li>Old preferences may be outdated</li>\n<li>Task details lose relevance</li>\n<li>Conflicting memories confuse retrieval</li>\n</ul>\n<p>Implement intelligent decay based on:</p>\n<ul>\n<li>Recency (when was it created/accessed?)</li>\n<li>Frequency (how often is it retrieved?)</li>\n<li>Importance (is it a core fact or detail?)\n\"\"\"</li>\n</ul>\n<h2>Time-Based Decay</h2>\n<p>\"\"\"\nfrom datetime import datetime, timedelta</p>\n<p>async def decay_old_memories(namespace: str, max_age_days: int):\ncutoff = datetime.now() - timedelta(days=max_age_days)</p>\n<pre><code>old_memories = await memory.episodic.list(\n    namespace=namespace,\n    filter={\"last_accessed\": {\"$lt\": cutoff.isoformat()}}\n)\n\nfor mem in old_memories:\n    # Soft delete (mark as archived)\n    await memory.episodic.update(\n        id=mem.id,\n        metadata={\"archived\": True, \"archived_at\": datetime.now()}\n    )\n</code></pre>\n<p>\"\"\"</p>\n<h2>Utility-Based Decay (MIRIX Approach)</h2>\n<p>\"\"\"\ndef calculate_memory_utility(memory):\n'''\nComposite utility score inspired by cognitive science:\n- Recency: When was it last accessed?\n- Frequency: How often is it accessed?\n- Importance: How critical is this information?\n'''\nnow = datetime.now()</p>\n<pre><code># Recency score (exponential decay with 72h half-life)\nhours_since_access = (now - memory.last_accessed).total_seconds() / 3600\nrecency_score = 0.5 ** (hours_since_access / 72)\n\n# Frequency score\nfrequency_score = min(memory.access_count / 10, 1.0)\n\n# Importance (from metadata or heuristic)\nimportance = memory.metadata.get(\"importance\", 0.5)\n\n# Weighted combination\nutility = (\n    0.4 * recency_score +\n    0.3 * frequency_score +\n    0.3 * importance\n)\n\nreturn utility\n</code></pre>\n<p>async def prune_low_utility_memories(threshold=0.2):\nall_memories = await memory.list_all()\nfor mem in all_memories:\nif calculate_memory_utility(mem) &lt; threshold:\nawait memory.archive(mem.id)\n\"\"\"</p>\n<h2>Sharp Edges</h2>\n<h3>Chunking Isolates Information From Its Context</h3>\n<p>Severity: CRITICAL</p>\n<p>Situation: Processing documents for vector storage</p>\n<p>Symptoms:\nRetrieval finds chunks but they don't make sense alone. Agent\nanswers miss the big picture. \"The function returns X\" retrieved\nwithout knowing which function. References to \"this\" without\nknowing what \"this\" refers to.</p>\n<p>Why this breaks:\nWhen we chunk for AI processing, we're breaking connections,\nreducing a holistic narrative to isolated fragments that often\nmiss the big picture. A chunk about \"the configuration\" without\ncontext about what system is being configured is nearly useless.</p>\n<p>Recommended fix:</p>\n<h3>Contextual Chunking (Anthropic's approach)</h3>\n<h1>Add document context to each chunk before embedding</h1>\n<h1>Reduces retrieval failures by 35%</h1>\n<p>def contextualize_chunk(chunk, document):\nsummary = summarize(document)</p>\n<pre><code># LLM generates context for chunk\ncontext = llm.invoke(f'''\n    Document summary: {summary}\n\n    Generate a brief context statement for this chunk\n    that would help someone understand what it refers to:\n\n    {chunk}\n''')\n\nreturn f\"{context}\\n\\n{chunk}\"\n</code></pre>\n<h1>Embed the contextualized version</h1>\n<p>for chunk in chunks:\ncontextualized = contextualize_chunk(chunk, full_doc)\nembedding = embed(contextualized)\n# Store original chunk, embed contextualized\nstore(original=chunk, embedding=embedding)</p>\n<h2>Hierarchical Chunking</h2>\n<h1>Store at multiple granularities</h1>\n<p>chunks_small = split(doc, size=256)\nchunks_medium = split(doc, size=512)\nchunks_large = split(doc, size=1024)</p>\n<h1>Retrieve at appropriate level based on query</h1>\n<h3>Chunk Size Mismatched to Query Patterns</h3>\n<p>Severity: HIGH</p>\n<p>Situation: Configuring chunking for memory storage</p>\n<p>Symptoms:\nHigh-quality documents produce low-quality retrievals. Simple\nquestions miss relevant information. Complex questions get\nfragments instead of complete answers.</p>\n<p>Why this breaks:\nOptimal chunk size depends on query patterns:</p>\n<ul>\n<li>Factual queries need small, specific chunks</li>\n<li>Conceptual queries need larger context</li>\n<li>Code needs function-level boundaries</li>\n</ul>\n<p>The sweet spot varies by document type and embedding model.\nDefault 1000 characters works for nothing specific.</p>\n<p>Recommended fix:</p>\n<h2>Test different sizes</h2>\n<p>from sklearn.metrics import recall_score</p>\n<p>def evaluate_chunk_size(documents, test_queries, chunk_size):\nchunks = split_documents(documents, size=chunk_size)\nindex = build_index(chunks)</p>\n<pre><code>correct_retrievals = 0\nfor query, expected_chunk in test_queries:\n    results = index.search(query, k=5)\n    if expected_chunk in results:\n        correct_retrievals += 1\n\nreturn correct_retrievals / len(test_queries)\n</code></pre>\n<h1>Test multiple sizes</h1>\n<p size=\"\">for size in [256, 512, 768, 1024]:\nrecall = evaluate_chunk_size(docs, test_queries, size)\nprint(f\"Size : Recall@5 = {recall:.2%}\")</p>\n<h2>Size recommendations by content type</h2>\n<p>CHUNK_SIZES = {\n\"documentation\": 512,   # Complete concepts\n\"code\": 1000,          # Function-level\n\"conversation\": 256,   # Turn-level\n\"articles\": 768,       # Paragraph-level\n}</p>\n<h2>Use overlap to prevent boundary issues</h2>\n<p>splitter = RecursiveCharacterTextSplitter(\nchunk_size=512,\nchunk_overlap=50,  # 10% overlap\n)</p>\n<h3>Semantic Search Returns Irrelevant Results</h3>\n<p>Severity: HIGH</p>\n<p>Situation: Querying memory for context</p>\n<p>Symptoms:\nAgent retrieves memories that seem related but aren't useful.\n\"Tell me about the user's preferences\" returns conversation\nabout preferences in general, not this user's. High similarity\nscores for wrong content.</p>\n<p>Why this breaks:\nSemantic similarity isn't the same as relevance. \"The user\nlikes Python\" and \"Python is a programming language\" are\nsemantically similar but very different types of information.\nWithout metadata filtering, retrieval is just word matching.</p>\n<p>Recommended fix:</p>\n<h2>Always filter by metadata first</h2>\n<h1>Don't rely on semantic similarity alone</h1>\n<h1>Bad: Only semantic search</h1>\n<p>results = index.query(\nvector=query_embedding,\ntop_k=5\n)</p>\n<h1>Good: Filter then search</h1>\n<p>results = index.query(\nvector=query_embedding,\nfilter={\n\"user_id\": current_user.id,\n\"type\": \"preference\",\n\"created_after\": cutoff_date,\n},\ntop_k=5\n)</p>\n<h2>Use hybrid search (semantic + keyword)</h2>\n<p>from qdrant_client import QdrantClient</p>\n<p>client = QdrantClient(...)</p>\n<h1>Hybrid search with fusion</h1>\n<p>results = client.search(\ncollection_name=\"memories\",\nquery_vector=semantic_embedding,\nquery_text=query,  # Also keyword match\nfusion={\"method\": \"rrf\"},  # Reciprocal Rank Fusion\n)</p>\n<h2>Rerank results with cross-encoder</h2>\n<p>from sentence_transformers import CrossEncoder</p>\n<p>reranker = CrossEncoder(\"cross-encoder/ms-marco-MiniLM-L-6-v2\")</p>\n<h1>Initial retrieval (recall-oriented)</h1>\n<p>candidates = index.query(query_embedding, top_k=20)</p>\n<h1>Rerank (precision-oriented)</h1>\n<p>pairs = [(query, c.text) for c in candidates]\nscores = reranker.predict(pairs)\nreranked = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True)</p>\n<h3>Old Memories Override Current Information</h3>\n<p>Severity: HIGH</p>\n<p>Situation: User preferences or facts change over time</p>\n<p>Symptoms:\nAgent uses outdated preferences. \"User prefers dark mode\" from\n6 months ago overrides recent \"switch to light mode\" request.\nAgent confidently uses stale data.</p>\n<p>Why this breaks:\nVector stores don't have temporal awareness by default. A memory\nfrom a year ago has the same retrieval weight as one from today.\nRecent information should generally override old information\nfor preferences and mutable facts.</p>\n<p>Recommended fix:</p>\n<h2>Add temporal scoring</h2>\n<p>from datetime import datetime, timedelta</p>\n<p>def time_decay_score(memory, half_life_days=30):\nage = (datetime.now() - memory.created_at).days\ndecay = 0.5 ** (age / half_life_days)\nreturn decay</p>\n<p>def retrieve_with_recency(query, user_id):\n# Get candidates\ncandidates = index.query(\nvector=embed(query),\nfilter={\"user_id\": user_id},\ntop_k=20\n)</p>\n<pre><code># Apply time decay\nfor candidate in candidates:\n    time_score = time_decay_score(candidate)\n    candidate.final_score = candidate.similarity * 0.7 + time_score * 0.3\n\n# Re-sort by final score\nreturn sorted(candidates, key=lambda x: x.final_score, reverse=True)[:5]\n</code></pre>\n<h2>Update instead of append for preferences</h2>\n<p>async def update_preference(user_id, category, value):\n# Delete old preference\nawait memory.delete(\nfilter={\"user_id\": user_id, \"type\": \"preference\", \"category\": category}\n)</p>\n<pre><code># Store new preference\nawait memory.upsert(\n    id=f\"pref-{user_id}-{category}\",\n    content={\"category\": category, \"value\": value},\n    metadata={\"updated_at\": datetime.now()}\n)\n</code></pre>\n<h2>Explicit versioning for facts</h2>\n<p>await memory.upsert(\nid=f\"fact--v\",\ncontent=new_fact,\nmetadata={\n\"version\": version,\n\"supersedes\": previous_id,\n\"valid_from\": datetime.now()\n}\n)</p>\n<h3>Contradictory Memories Retrieved Together</h3>\n<p>Severity: MEDIUM</p>\n<p>Situation: User has changed preferences or provided conflicting info</p>\n<p>Symptoms:\nAgent retrieves \"user prefers dark mode\" and \"user prefers light\nmode\" in same context. Gives inconsistent answers. Seems confused\nor forgetful to user.</p>\n<p>Why this breaks:\nWithout conflict resolution, both old and new information coexist.\nSemantic search might return both because they're both about the\nsame topic (preferences). Agent has no way to know which is current.</p>\n<p>Recommended fix:</p>\n<h2>Detect conflicts on storage</h2>\n<p>async def store_with_conflict_check(memory, user_id):\n# Find potentially conflicting memories\nsimilar = await index.query(\nvector=embed(memory.content),\nfilter={\"user_id\": user_id, \"type\": memory.type},\nthreshold=0.9,  # Very similar\ntop_k=5\n)</p>\n<pre><code>for existing in similar:\n    if is_contradictory(memory.content, existing.content):\n        # Ask for resolution\n        resolution = await resolve_conflict(memory, existing)\n        if resolution == \"replace\":\n            await index.delete(existing.id)\n        elif resolution == \"version\":\n            await mark_superseded(existing.id, memory.id)\n\nawait index.upsert(memory)\n</code></pre>\n<h2>Conflict detection heuristic</h2>\n<p>def is_contradictory(new_content, old_content):\n# Use LLM to detect contradiction\nresult = llm.invoke(f'''\nDo these two statements contradict each other?</p>\n<pre><code>    Statement 1: {old_content}\n    Statement 2: {new_content}\n\n    Respond with just YES or NO.\n''')\nreturn result.strip().upper() == \"YES\"\n</code></pre>\n<h2>Periodic consolidation</h2>\n<p>async def consolidate_memories(user_id):\nall_memories = await index.list(filter={\"user_id\": user_id})\nclusters = cluster_by_topic(all_memories)</p>\n<pre><code>for cluster in clusters:\n    if has_conflicts(cluster):\n        resolved = await llm.invoke(f'''\n            These memories may conflict. Create one consolidated\n            memory that represents the current truth:\n            {cluster}\n        ''')\n        await replace_cluster(cluster, resolved)\n</code></pre>\n<h3>Retrieved Memories Exceed Context Window</h3>\n<p>Severity: MEDIUM</p>\n<p>Situation: Retrieving too many memories at once</p>\n<p>Symptoms:\nToken limit errors. Agent truncates important information.\nSystem prompt gets cut off. Retrieved memories compete with\nuser query for space.</p>\n<p>Why this breaks:\nRetrieval typically returns top-k results. If k is too high or\nchunks are too large, retrieved context overwhelms the window.\nCritical information (system prompt, recent messages) gets pushed\nout.</p>\n<p>Recommended fix:</p>\n<h2>Budget tokens for different memory types</h2>\n<p>TOKEN_BUDGET = {\n\"system_prompt\": 500,\n\"user_profile\": 200,\n\"recent_messages\": 2000,\n\"retrieved_memories\": 1000,\n\"current_query\": 500,\n\"buffer\": 300,  # Safety margin\n}</p>\n<p>def budget_aware_retrieval(query, context_limit=4000):\nremaining = context_limit - TOKEN_BUDGET[\"system_prompt\"] - TOKEN_BUDGET[\"buffer\"]</p>\n<pre><code># Prioritize recent messages\nrecent = get_recent_messages(limit=TOKEN_BUDGET[\"recent_messages\"])\nremaining -= count_tokens(recent)\n\n# Then user profile\nprofile = get_user_profile(limit=TOKEN_BUDGET[\"user_profile\"])\nremaining -= count_tokens(profile)\n\n# Finally retrieved memories with remaining budget\nmemories = retrieve_memories(query, max_tokens=remaining)\n\nreturn build_context(profile, recent, memories)\n</code></pre>\n<h2>Dynamic k based on chunk size</h2>\n<p>def retrieve_with_budget(query, max_tokens=1000):\navg_chunk_tokens = 150  # From your data\nmax_k = max_tokens // avg_chunk_tokens</p>\n<pre><code>results = index.query(query, top_k=max_k)\n\n# Trim if still over budget\ntotal_tokens = 0\nfiltered = []\nfor result in results:\n    tokens = count_tokens(result.text)\n    if total_tokens + tokens &lt;= max_tokens:\n        filtered.append(result)\n        total_tokens += tokens\n    else:\n        break\n\nreturn filtered\n</code></pre>\n<h3>Query and Document Embeddings From Different Models</h3>\n<p>Severity: MEDIUM</p>\n<p>Situation: Upgrading embedding model or mixing providers</p>\n<p>Symptoms:\nRetrieval quality suddenly drops. Relevant documents not found.\nRandom results returned. Works for new documents, fails for old.</p>\n<p>Why this breaks:\nEmbedding models produce different vector spaces. A query embedded\nwith text-embedding-3 won't match documents embedded with text-ada-002.\nMixing models creates garbage similarity scores.</p>\n<p>Recommended fix:</p>\n<h2>Track embedding model in metadata</h2>\n<p>await index.upsert(\nid=doc_id,\nvector=embedding,\nmetadata={\n\"embedding_model\": \"text-embedding-3-small\",\n\"embedding_version\": \"2024-01\",\n\"content\": content\n}\n)</p>\n<h2>Filter by model version on retrieval</h2>\n<p>results = index.query(\nvector=query_embedding,\nfilter={\"embedding_model\": current_model},\ntop_k=10\n)</p>\n<h2>Migration strategy for model upgrades</h2>\n<p>async def migrate_embeddings(old_model, new_model):\n# Get all documents with old model\nold_docs = await index.list(filter={\"embedding_model\": old_model})</p>\n<pre><code>for doc in old_docs:\n    # Re-embed with new model\n    new_embedding = await embed(doc.content, model=new_model)\n\n    # Update in place\n    await index.update(\n        id=doc.id,\n        vector=new_embedding,\n        metadata={\"embedding_model\": new_model}\n    )\n</code></pre>\n<h2>Use separate collections during migration</h2>\n<h1>Old collection: production queries</h1>\n<h1>New collection: re-embedding in progress</h1>\n<h1>Switch over when complete</h1>\n<h2>Validation Checks</h2>\n<h3>In-Memory Store in Production Code</h3>\n<p>Severity: ERROR</p>\n<p>In-memory stores lose data on restart</p>\n<p>Message: In-memory store detected. Use persistent storage (Postgres, Qdrant, Pinecone) for production.</p>\n<h3>Vector Upsert Without Metadata</h3>\n<p>Severity: WARNING</p>\n<p>Vectors should have metadata for filtering</p>\n<p>Message: Vector upsert without metadata. Add user_id, type, timestamp for proper filtering.</p>\n<h3>Query Without User Filtering</h3>\n<p>Severity: ERROR</p>\n<p>Queries should filter by user to prevent data leakage</p>\n<p>Message: Vector query without user filtering. Always filter by user_id to prevent data leakage.</p>\n<h3>Hardcoded Chunk Size Without Justification</h3>\n<p>Severity: INFO</p>\n<p>Chunk size should be tested and justified</p>\n<p>Message: Hardcoded chunk size. Test different sizes for your content type and measure retrieval accuracy.</p>\n<h3>Chunking Without Overlap</h3>\n<p>Severity: WARNING</p>\n<p>Chunk overlap prevents boundary issues</p>\n<p>Message: Text splitting without overlap. Add chunk_overlap (10-20%) to prevent boundary issues.</p>\n<h3>Semantic Search Without Filters</h3>\n<p>Severity: WARNING</p>\n<p>Pure semantic search often returns irrelevant results</p>\n<p>Message: Pure semantic search. Add metadata filters (user, type, time) for better relevance.</p>\n<h3>Retrieval Without Result Limit</h3>\n<p>Severity: WARNING</p>\n<p>Unbounded retrieval can overflow context</p>\n<p>Message: Retrieval without limit. Set top_k to prevent context overflow.</p>\n<h3>Embeddings Without Model Version Tracking</h3>\n<p>Severity: WARNING</p>\n<p>Track embedding model to handle migrations</p>\n<p>Message: Store embedding model version in metadata to handle model migrations.</p>\n<h3>Different Models for Document and Query Embedding</h3>\n<p>Severity: ERROR</p>\n<p>Documents and queries must use same embedding model</p>\n<p>Message: Ensure same embedding model for indexing and querying.</p>\n<h2>Collaboration</h2>\n<h3>Delegation Triggers</h3>\n<ul>\n<li>user needs vector database at scale -&gt; data-engineer (Production vector store operations)</li>\n<li>user needs embedding model optimization -&gt; ml-engineer (Custom embeddings, fine-tuning)</li>\n<li>user needs knowledge graph -&gt; knowledge-engineer (Graph-based memory structures)</li>\n<li>user needs RAG pipeline -&gt; llm-architect (End-to-end retrieval augmented generation)</li>\n<li>user needs multi-agent shared memory -&gt; multi-agent-orchestration (Memory sharing between agents)</li>\n</ul>\n<h2>Related Skills</h2>\n<p>Works well with: <code>autonomous-agents</code>, <code>multi-agent-orchestration</code>, <code>llm-architect</code>, <code>agent-tool-builder</code></p>\n<h2>When to Use</h2>\n<ul>\n<li>User mentions or implies: agent memory</li>\n<li>User mentions or implies: long-term memory</li>\n<li>User mentions or implies: memory systems</li>\n<li>User mentions or implies: remember across sessions</li>\n<li>User mentions or implies: memory retrieval</li>\n<li>User mentions or implies: episodic memory</li>\n<li>User mentions or implies: semantic memory</li>\n<li>User mentions or implies: vector store</li>\n<li>User mentions or implies: rag</li>\n<li>User mentions or implies: langmem</li>\n<li>User mentions or implies: memgpt</li>\n<li>User mentions or implies: conversation history</li>\n</ul>\n<h2>Limitations</h2>\n<ul>\n<li>Use this skill only when the task clearly matches the scope described above.</li>\n<li>Do not treat the output as a substitute for environment-specific validation, testing, or expert review.</li>\n<li>Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.</li>\n</ul>\n","files":[{"path":"SKILL.md","sizeBytes":31187,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"notes-only","suspicious":0,"notes":2,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-08-16T13:39:13.486184Z","sha256":"BF37C5E9C53EF6C4E96EAF4F507B294DC621F8C195576597A3D62CD9C21FAE4A","sizeBytes":10957},"review":null,"source":{"repositoryUrl":"https://github.com/sickn33/agentic-awesome-skills","path":"skills/agent-memory-systems","license":"MIT","commit":"f2bba339de74414b0771234cbe4f6a15258e32a3","subtreeSha":"1999C34BFEFFE4A22DDEE0107E7845E9273BCCF5EDFD310FF2AEB1295E1BC4EE","lastSyncedAt":"2026-09-25T06:48:39.853703Z"},"reviewedAt":"2026-08-16T13:39:28.033838Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/agent-memory-systems"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install sickn33-agentic-awesome-skills@llmmart"},{"target":"git","command":"git clone https://github.com/sickn33/agentic-awesome-skills.git"}]}