Memory · Data store
Working Memory Buffer
Data storeMemoryCognition & MemoryVariation point (abstract)arc:WorkingMemoryBuffer
Short-term, in-context storage of the current task's reasoning traces, tool invocations, observations, and reflection insights, bounded by the model's context window.
Responsibility. Holds current-task context within the token budget.
Also known as: Short-term memory, Scratchpad, Conversation buffer, Context window, Agent state, Workflow state, Execution state, Shared state, Conversation history, Task-scoped working memory, Chat history buffer, Agent scratchpad, Conversation message history, Messages array, Multi-hop reasoning state, Short-term working memory, Agent trajectory, System non-parametric short-term memory, Reasoning scratch pad, System non-parametric short-term memory (Quadrant V), Active desk space, Current-session context, Working memory, Mental scratchpad, Active workspace, Short-term conversational memory
Variants
| Variant | When to choose |
|---|---|
| Graph Reasoning State Store | Choose when thoughts need multiple parents so independent reasoning chains can merge. |
| MCTS Search Tree Store | — |
| Thought Tree Store | Choose when reasoning branches diverge but never need to reconverge. |
Relationships
is configured by structural
is read by dependency
- Agent Controller abstract Ch1.4 Ch1.7B +1
- Context Window Manager abstract Ch1.2 Ch1.6 +1
- Episode Encoder abstract Ch5.7
- Failure Analyzer Ch2.2
- Memory Consolidator Ch1.6 Ch5.1 +1
- Memory Retriever Ch10.1
- Parameter Provenance Validator Ch3.7
- Plan Executor Ch1.5A Ch1.6
- Prompt Context Builder Ch1.6
- Query Rewriter Ch3.3
- ReAct Agent Controller Ch1.2 Ch2.3
- Reasoning Engine Ch1.5A Ch2.2 +2
- Replanner abstract Ch1.5A Ch1.6
- Rule-Based Transition Router Ch2.2
- Sequential Tool Dispatcher Ch2.6
- State Cycle Detector Ch1.6
- State-Graph Orchestrator Ch2.2
- Thought Exploration Controller abstract Ch5.2
- Tool Precondition Checker Ch3.7
- Trajectory Pruner Ch3.10
- Transition Router abstract Ch1.5A Ch1.5B +1
- Worker Agent abstract Ch1.5A Ch1.5B +1
- Workflow Orchestrator abstract Ch1.5A Ch1.5B
is written by dependency
- Agent Controller abstract Ch1.5A Ch1.6 +1
- Context Window Manager abstract Ch1.4 Ch1.6 +1
- Dialogue Flow Manager Ch10.1
- Escalation Handler Ch1.5B
- Failure Analyzer Ch2.2
- Intent Router Ch1.5B
- Multi-Hop Retrieval Controller Ch3.3
- Output Verifier Ch2.2
- Plan Executor Ch1.5A Ch1.6
- Procedural Skill Executor Ch5.9
- Prompt Context Builder Ch5.9
- ReAct Agent Controller Ch1.2 Ch2.3
- Reasoning Engine Ch5.1 Ch5.9
- Reasoning Path Sampler Ch5.3
- Reflection Critic abstract Ch1.2
- Replanner abstract Ch1.5A Ch1.6
- State-Graph Orchestrator Ch2.2 Ch2.6
- Task Planner abstract Ch1.5A Ch1.6
- Tool Executor Ch2.6
- Trajectory Pruner Ch3.10
- Web Navigation Agent Ch3.3
- Worker Agent abstract Ch1.5A Ch1.5B +1
- Workflow Orchestrator abstract Ch1.5A Ch1.5B
is guarded by control
is monitored by assurance
Design guidance
- SHOULD keep recent conversation in full detail while older history enters only as summaries.
- MUST NOT be relied on for cross-session questions; those require long-term memory.
- SHOULD serve as the shared context through which specialised agents read inputs and write outputs for downstream agents.
- SHOULD be updated immutably (copy-and-return) so transitions are explicit and before/after states can be compared.
- SHOULD be preferred over message passing or publish-subscribe for coordinating specialised agents, as it captures all artifacts in one observable structure and simplifies recovery.
- SHOULD manage tool-chain state explicitly; accumulated history and tool results can overflow the context window.
- MUST retain within-session working context; removing short-term memory hurts complex multi-turn reasoning more than removing long-term memory.
- SHOULD NOT keep every CoT trace in the prompt context; move traces to external stores with explicit indexing, summarisation and decay decisions.
- MUST budget the context window as capacity shared by system prompt, history, retrieved documents, reasoning traces, tool outputs and generated output, not by history or input alone.
- SHOULD design for effective reasoning capacity rather than nominal window size, reserving margin for attention limitations and cognitive-load effects.
- SHOULD minimise extraneous load even when the token budget permits inclusion, because irrelevant content actively degrades reasoning.
- MUST NOT treat working memory as persistent across sessions; information to be kept must be explicitly transferred to long-term memory.
Quantitative guidance
As stated by the sources; verify before use.
- Example: last 10 messages kept in the buffer (Ch1.7B).
- Large LLM context windows max out at 100K-200K tokens (Ch5.7).
- 256k-token window example: 500 system + 50,000 history + 100,000 retrieved + 1,000 query leaves ~104,500-105,000 tokens (~41%) for reasoning and output (Ch5.9).
- Effective reasoning capacity of a 256k-token window estimated at ~100,000-150,000 tokens (Ch5.9).
- Performance drops sharply once extraneous content exceeds ~40-60% of context; three-hop reasoning accuracy 95% (low extraneous load) vs 72% (high) vs 45% (extremely saturated) (Ch5.9).
- Paris-hotels example: 5,000 tokens of targeted reviews outperform 50,000 tokens of general Paris history in a 100k window (Ch5.9).
- Long-output tasks can claim 30-50% of total capacity for output (Ch5.9).
- Short-term memory stores the most recent 5-10 turns verbatim; Erica keeps 8-10 turns verbatim (Ch10.1).
Classification
- Patterns
- Memory integration (reflection)Shared-state coordinationObservation accumulationTiered memoryTracking already-verified conditions to skip redundant checksFour-phase working-memory lifecycle (context assembly, processing/reasoning, generation, disposal)Token accounting: total = system prompt + input + retrieved context + history + outputSelective transfer to long-term memory before disposal
- Technologies
- LangGraphLangChain
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Maintainability (ISO/IEC 25010)Cost efficiency
- Risks mitigated
- Lost context across step boundariesLost intermediate resultsStale data passed downstream
Sources
- Ch1.2: T. Nguyen, "Core Agent Patterns," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.2. ISBN: 9798244538229.
- Ch1.4: T. Nguyen, "Memory and Perception Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.4. ISBN: 9798244538229.
- Ch1.5A: T. Nguyen, "Stateful Orchestration - Introduction and Core Concepts," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.5A. ISBN: 9798244538229.
- Ch1.5B: T. Nguyen, "Stateful Orchestration - Worked Examples," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.5B. ISBN: 9798244538229.
- Ch1.6: T. Nguyen, "Stateful Orchestration - Pitfalls, Integration, and Synthesis," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.6. ISBN: 9798244538229.
- Ch1.7B: T. Nguyen, "Relational Reasoning with Knowledge Graphs - Hybrid RAG+KG Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.7B. ISBN: 9798244538229.
- Ch2.3: T. Nguyen, "LangChain Sequential Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.3. ISBN: 9798244538229.
- Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.7: T. Nguyen, "Tool Usage Auditing," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.7. ISBN: 9798244538229.
- Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch5.1: T. Nguyen, "Chain-of-Thought (CoT) Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.1. ISBN: 9798244538229.
- Ch5.2: T. Nguyen, "Tree-of-Thought (ToT) Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.2. ISBN: 9798244538229.
- Ch5.3: T. Nguyen, "Self-Consistency Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.3. ISBN: 9798244538229.
- Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
- Ch5.7: T. Nguyen, "Episodic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.7. ISBN: 9798244538229.
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ref5.04: C. Stryker, "What is AI agent memory?," IBM Think. Accessed: Sep. 27, 2026. [Online]. Available: https://www.ibm.com/think/topics/ai-agent-memory