82.2 ms
Mean Search Latency
H800 · 10 QPS
An In-Process Semantic Memory Runtime
for Agent Retrieval Systems
Current Mandol exposes MemoryUnit, SemanticMap, SemanticGraph, MultiRetriever, and three-tower retrieval components from the src/mandol package.
Mandol is an in-process semantic memory package for agent retrieval experiments. Its indexes and graph topology remain resident, with optional RocksDB-backed paging for cold MemoryUnit payloads.
Retrieval is handled through MultiRetriever for BM25, SPLADE, cosine search, graph expansion, score fusion, and reranker orchestration. Current development is on main; exact reproduction of the published experiments uses the frozen paper-repro branch.

The maintained runtime is built from explicit, inspectable memory primitives
MemoryUnit stores the payload, MemorySpace stores logical membership, and MemorySpaceRegistry defines canonical three-tower spaces.
SemanticMap owns embeddings, FAISS indexing, sparse vectors and space-filtered search; SemanticGraph adds rustworkx relationships and graph persistence.
MultiRetriever and TripleTowerRetriever provide retrieval-facing orchestration for dense, sparse, graph, episodic and entity-relation paths.
SOTA-level accuracy with lower token consumption on long-term conversational memory benchmarks
| GPT-4o-mini | Avg. Tok | Single | Multi | Temp. | Open | Overall |
|---|---|---|---|---|---|---|
| Mem0 | 1.0k | 66.71 | 58.16 | 55.45 | 40.62 | 61.00 |
| MemU | 4.0k | 72.77 | 62.41 | 33.96 | 46.88 | 61.15 |
| MemOS | 2.5k | 81.45 | 69.15 | 72.27 | 60.42 | 75.87 |
| Zep | 1.4k | 88.11 | 71.99 | 74.45 | 66.67 | 81.06 |
| EverMemOS | 2.5k | 91.68 | 82.74 | 79.34 | 70.14 | 86.13 |
| Mandol | 2.0k | 93.82 | 85.11 | 89.10 | 65.63 | 89.48 |
| GPT-4.1-mini | Avg. Tok | Single | Multi | Temp. | Open | Overall |
|---|---|---|---|---|---|---|
| Mem0 | 1.0k | 68.97 | 61.70 | 58.26 | 50.00 | 64.20 |
| MemU | 4.0k | 74.91 | 72.34 | 43.61 | 54.17 | 66.67 |
| MemOS | 2.5k | 85.37 | 79.43 | 75.08 | 64.58 | 80.76 |
| Zep | 1.4k | 90.84 | 81.91 | 77.26 | 75.00 | 85.22 |
| EverMemOS | 2.3k | 95.32 | 89.01 | 90.13 | 77.43 | 91.97 |
| Mandol | 1.9k | 95.36 | 92.20 | 87.85 | 79.17 | 92.21 |
Mandol achieves 92.21% overall on LoCoMo with only 1.9k tokens. Paper rows use generated three-tower graphs with router + quantification; Qwen/DeepSeek build and deduplicate memories, while GPT-4.1-mini/GPT-4o-mini are used for task evaluation and judging.
| GPT-4o-mini | Avg. Tok | SS-Pref | SS-Asst | Temporal | Multi-S | Know. Upd. | SS-User | Overall |
|---|---|---|---|---|---|---|---|---|
| MemU | 0.5k | 76.70 | 19.60 | 17.30 | 42.10 | 41.00 | 67.10 | 38.40 |
| Mem0 | 1.1k | 90.00 | 26.78 | 72.18 | 63.15 | 66.67 | 82.86 | 66.40 |
| Zep | 1.6k | 53.30 | 75.00 | 54.10 | 47.40 | 74.40 | 92.90 | 63.80 |
| MemOS | 1.4k | 96.67 | 67.86 | 77.44 | 70.67 | 74.26 | 95.71 | 77.80 |
| Mandol | 2.1k | 96.67 | 98.21 | 78.95 | 74.44 | 88.46 | 97.14 | 85.00 |
| GPT-4.1-mini | Avg. Tok | SS-Pref | SS-Asst | Temporal | Multi-S | Know. Upd. | SS-User | Overall |
|---|---|---|---|---|---|---|---|---|
| EverMemOS | 2.8k | 93.33 | 85.71 | 77.44 | 73.68 | 89.74 | 97.14 | 83.00 |
| Mandol | 2.3k | 96.67 | 98.21 | 87.22 | 77.44 | 89.74 | 98.57 | 88.40 |
Mandol achieves 88.40% overall on LongMemEval with 2.3k tokens. Reproduction requires the cleaned dataset plus generated hierarchical, episodic, and entity-relation graph artifacts before running task_eval.
Complete tail and mean latency results, with server and local deployments reported separately
82.2 ms
H800 · 10 QPS
39.7 ms
H800 · 10 QPS
5.4×
server mean · vs. best non-Mandol baseline
4.8×
server mean · vs. best non-Mandol baseline
NVIDIA H800 80GB · Search 10 QPS · Add 10 QPS
| System | Search | Add | ||||
|---|---|---|---|---|---|---|
| P99 | P90 | Mean | P99 | P90 | Mean | |
| MemU | 63000.7 | 60539.5 | 47554.5 | 12070.6 | 7273.1 | 5077.9 |
| EverMemOS† | 37192.4 | 35220.4 | 20092.1 | 790.2 | 555.5 | 317.7 |
| Mem0 | 4637.0 | 1397.0 | 1089.0 | 2841.0 | 1650.0 | 888.0 |
| Zep | 5348.7 | 614.8 | 571.7 | 375.1 | 254.5 | 239.0 |
| MemOS | 777.1 | 528.4 | 440.5 | 376.4 | 211.6 | 191.9 |
| Mandol (Ours) | 94.8 | 88.5 | 82.2 | 67.3 | 46.9 | 39.7 |
All values are milliseconds. Server measurements use an NVIDIA H800 80GB at 10 QPS for Search and Add. Local measurements use an NVIDIA RTX 5090 Laptop 24GB at 5 QPS for Search and 10 QPS for Add.
† EverMemOS was reproduced using its official implementation.
These results correspond to the frozen paper reproduction artifact.
Install the package or source environment, add MemoryUnit records, search with MultiRetriever, then persist the graph
If this work is helpful to your research, please cite our paper
@misc{zhang2026mandol,
title={Mandol: An Agglomerative Agent Memory System for Long-Term Conversations},
author={Yuhan Zhang and Zhiyuan Guo and Ziheng Zeng and Wei Wang and Wentao Wu and Lijie Xu},
year={2026},
eprint={2606.29778},
archivePrefix={arXiv},
primaryClass={cs.DB},
doi={10.48550/arXiv.2606.29778},
url={https://arxiv.org/abs/2606.29778}
}The paper is available as arXiv:2606.29778.