Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
53 changes: 51 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,12 +4,14 @@

### Agents remember,Humans innovate.

<a href="https://trendshift.io/repositories/29310?utm_source=repository-badge&amp;utm_medium=badge&amp;utm_campaign=badge-repository-29310" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/repositories/29310" alt="TencentCloud%2FTencentDB-Agent-Memory | Trendshift" width="250" height="55"/></a>

[![npm](https://img.shields.io/npm/v/@tencentdb-agent-memory/memory-tencentdb?color=blue)](https://www.npmjs.com/package/@tencentdb-agent-memory/memory-tencentdb)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](./LICENSE)
[![Node](https://img.shields.io/badge/node-%3E=22.16-brightgreen)](https://nodejs.org/)
[![OpenClaw](https://img.shields.io/badge/OpenClaw-%3E=2026.3.13-orange)](https://github.com/openclaw/openclaw)
[![Hermes](https://img.shields.io/badge/Hermes-Gateway-7B61FF)](https://hermes-agent.nousresearch.com/docs/)
[![Discord](https://img.shields.io/badge/Discord-Join-5865F2?logo=discord&logoColor=white)](https://discord.gg/kDtHb5RW2)
[![Discord](https://img.shields.io/badge/Discord-Join-5865F2?logo=discord&logoColor=white)](https://discord.gg/dJQM6mKMF)

[Highlights](#-highlights) · [Overview](#overview) · [Core Technology](#core-technology-reject-flat-storage-embrace-layering-and-symbolization) · [Features](#-features) · [Quick Start](#quick-start)

Expand Down Expand Up @@ -351,6 +353,36 @@ curl http://127.0.0.1:8420/health

---

### 3. Hermes (Windows native)

For a Windows-native Hermes install, run the bundled batch script from the
repository root in Command Prompt or PowerShell:

```powershell
$env:TDAI_LLM_API_KEY="your-api-key"
$env:TDAI_LLM_BASE_URL="https://api.openai.com/v1"
$env:TDAI_LLM_MODEL="gpt-4o"
.\scripts\setup-hermes-memory-tencentdb.bat
```

The script checks `node`, `npm`, Python, and Hermes, requires Node.js
`>=22.16.0`, runs `npm install --omit=dev` when Gateway dependencies are
missing, creates `%USERPROFILE%\.memory-tencentdb\memory-tdai`, copies the
provider to `%USERPROFILE%\.hermes\plugins\memory_tencentdb`, writes Gateway
environment variables to `%USERPROFILE%\.hermes\.env`, and starts the Gateway
before polling:

```powershell
curl.exe http://127.0.0.1:8420/health
```

If `%USERPROFILE%\.hermes\config.yaml` already exists, make sure it contains:

```yaml
memory:
provider: memory_tencentdb
```


## 🔒 Gateway Security (optional)

Expand Down Expand Up @@ -403,6 +435,8 @@ If `MEMORY_TENCENTDB_GATEWAY_API_KEY` is unset, the plugin also looks at `TDAI_G
| `timezone` | `"system"` | Timezone for user/LLM-facing timestamps: `"system"` (follow process tz) / IANA name (`Asia/Shanghai`) / offset string (`+08:00`) |
| `storeBackend` | `"sqlite"` | Storage backend: `sqlite` |
| `recall.strategy` | `"hybrid"` | Recall strategy: `keyword` / `embedding` / `hybrid` (RRF fusion, recommended) |
| `recall.injectionMode` | `"prepend"` | Dynamic L1 recall placement: `prepend` keeps legacy behavior; `append` uses OpenClaw `appendContext` to avoid changing the user prompt prefix on compatible hosts |
| `recall.showInjected` | `false` | Preserve injected `<relevant-memories>` in persisted history. Keep `false` to avoid replaying stale dynamic recall and growing the prompt prefix |
| `recall.maxResults` | `5` | Number of items returned per recall |
| `recall.maxCharsPerMemory` | `0` | Max characters injected for one recalled L1 memory; `0` disables this guard |
| `recall.maxTotalRecallChars` | `0` | Total character budget for auto-recalled L1 memories; `0` disables this guard |
Expand All @@ -411,6 +445,9 @@ If `MEMORY_TENCENTDB_GATEWAY_API_KEY` is unset, the plugin also looks at `TDAI_G
| `persona.triggerEveryN` | `50` | Generate the user persona every N new memories |
| `offload.enabled` | `false` | Whether to enable short-term compression |

See [Prompt-cache mitigation and verification](docs/prompt-cache-mitigation.md)
for the context layout, multi-turn growth model, and provider A/B protocol.

</details>

<details>
Expand Down Expand Up @@ -517,7 +554,7 @@ We welcome every kind of contribution — bug reports, feature ideas, doc fixes,
- 🐞 **Found a bug or have a question?** Open an issue at [GitHub Issues](https://github.com/Tencent/TencentDB-Agent-Memory/issues) — we respond within 24 hours.
- 💡 **Have an idea to share?** Start a thread in [GitHub Discussions](https://github.com/Tencent/TencentDB-Agent-Memory/discussions).
- 🛠️ **Want to contribute code?** Please read [CONTRIBUTING.md](./CONTRIBUTING.md) first.
- 💬 **Want to chat with us?** Join our [Discord community](https://discord.gg/kDtHb5RW2) and talk to the early developers directly.
- 💬 **Want to chat with us?** Join our [Discord community](https://discord.gg/dJQM6mKMF) and talk to the early developers directly.

---

Expand Down Expand Up @@ -545,4 +582,16 @@ We welcome every kind of contribution — bug reports, feature ideas, doc fixes,
</tr>
</table>

---

## Star History

<p align="center">
<a href="https://www.star-history.com/#Tencent/TencentDB-Agent-Memory&Date">
<img src="https://github.com/user-attachments/assets/16753a90-8bc9-471b-819e-311947ed94f7" alt="Star History Chart" width="600" />
</a>
</p>

---

[MIT](./LICENSE) © TencentDB Agent Memory Team
56 changes: 52 additions & 4 deletions README_CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,13 +5,14 @@

### 让 Agent 沉淀经验,让人专注创造。

<a href="https://trendshift.io/repositories/29310?utm_source=repository-badge&amp;utm_medium=badge&amp;utm_campaign=badge-repository-29310" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/repositories/29310" alt="TencentCloud%2FTencentDB-Agent-Memory | Trendshift" width="250" height="55"/></a>

[![npm](https://img.shields.io/npm/v/@tencentdb-agent-memory/memory-tencentdb?color=blue)](https://www.npmjs.com/package/@tencentdb-agent-memory/memory-tencentdb)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](./LICENSE)
[![Node](https://img.shields.io/badge/node-%3E=22.16-brightgreen)](https://nodejs.org/)
[![OpenClaw](https://img.shields.io/badge/OpenClaw-%3E=2026.3.13-orange)](https://github.com/openclaw/openclaw)
[![Hermes](https://img.shields.io/badge/Hermes-Gateway-7B61FF)](https://hermes-agent.nousresearch.com/docs/)
[![Discord](https://img.shields.io/badge/Discord-Join-5865F2?logo=discord&logoColor=white)](https://discord.gg/kDtHb5RW2)
[![Discord](https://img.shields.io/badge/Discord-Join-5865F2?logo=discord&logoColor=white)](https://discord.gg/dJQM6mKMF)

[效果亮点](#-效果亮点) · [项目简介](#项目简介) · [核心技术](#核心技术拒绝平铺走向分层与符号化) · [方案特点](#-方案特点) · [快速开始](#快速开始)

Expand Down Expand Up @@ -55,7 +56,8 @@ TencentDB Agent Memory 帮助 Agent 学会你的流程、保留任务上下文

> **让 Agent 记住该记的,让人把注意力留给判断、创造和真正有价值的工作。**
<p align="center">
<img src="https://github.com/user-attachments/assets/6e9dd306-f462-4cae-adc4-3194be25c514" width="360" alt="Agent Memory 微信社群二维码" />
<img src="https://github.com/user-attachments/assets/9f77b819-c85b-4135-8b91-a612e42580a6" width="360" alt="Agent Memory 微信社群二维码" />


<br/>
<sub>📱 扫码加入 <b>Agent Memory 微信社群</b>,与早期开发者直接对话</sub>
Expand Down Expand Up @@ -353,9 +355,38 @@ curl http://127.0.0.1:8420/health

> Provider 的完整参考(环境变量、故障排查、LLM 工具 schema、supervisor 行为)见 [`hermes-plugin/memory/memory_tencentdb/README.md`](./hermes-plugin/memory/memory_tencentdb/README.md),调整 supervisor / circuit-breaker 默认值之前请先读它。


---

### 3. Hermes(Windows 原生安装)

Windows 原生 Hermes 环境下,在仓库根目录用 Command Prompt 或 PowerShell
运行内置批处理脚本:

```powershell
$env:TDAI_LLM_API_KEY="your-api-key"
$env:TDAI_LLM_BASE_URL="https://api.openai.com/v1"
$env:TDAI_LLM_MODEL="gpt-4o"
.\scripts\setup-hermes-memory-tencentdb.bat
```

脚本会检查 `node`、`npm`、Python 和 Hermes,要求 Node.js `>=22.16.0`;
当 Gateway 依赖缺失时执行 `npm install --omit=dev`,创建
`%USERPROFILE%\.memory-tencentdb\memory-tdai`,复制插件到
`%USERPROFILE%\.hermes\plugins\memory_tencentdb`,把 Gateway 环境变量写入
`%USERPROFILE%\.hermes\.env`,随后启动 Gateway 并轮询:

```powershell
curl.exe http://127.0.0.1:8420/health
```

如果 `%USERPROFILE%\.hermes\config.yaml` 已存在,请确认包含:

```yaml
memory:
provider: memory_tencentdb
```


## 🔒 Gateway 安全配置(可选)

Hermes Gateway 监听 `:8420`,对外提供 capture / search / recall 的 HTTP 接口。新增两个开关,可以把它从“开放的本地 sidecar”切换为“需要鉴权的网络服务”。**两个开关默认都关闭,已有部署的行为不变。**
Expand Down Expand Up @@ -406,6 +437,8 @@ export MEMORY_TENCENTDB_GATEWAY_API_KEY="<与 Gateway 同一份密钥>"
| `timezone` | `"system"` | 时区:`"system"`(跟随系统)/ IANA 名(`Asia/Shanghai`)/ offset 串(`+08:00`) |
| `storeBackend` | `"sqlite"` | 存储后端:`sqlite` |
| `recall.strategy` | `"hybrid"` | 召回策略:`keyword` / `embedding` / `hybrid`(RRF 融合,推荐) |
| `recall.injectionMode` | `"prepend"` | 动态 L1 召回注入位置:`prepend` 保持旧行为;`append` 在兼容 OpenClaw 中使用 `appendContext`,减少对用户消息前缀缓存的影响 |
| `recall.showInjected` | `false` | 是否把注入的 `<relevant-memories>` 保留到持久化历史;建议保持 `false`,避免重复回放动态召回内容导致历史膨胀 |
| `recall.maxResults` | `5` | 每次召回条数 |
| `recall.maxCharsPerMemory` | `0` | 单条 L1 记忆注入的最大字符数;`0` 表示不限制 |
| `recall.maxTotalRecallChars` | `0` | 每轮 auto-recall 注入的 L1 记忆总字符预算;`0` 表示不限制 |
Expand All @@ -414,6 +447,9 @@ export MEMORY_TENCENTDB_GATEWAY_API_KEY="<与 Gateway 同一份密钥>"
| `persona.triggerEveryN` | `50` | 每 N 条新记忆触发用户画像生成 |
| `offload.enabled` | `false` | 是否启用短期记忆压缩 |

上下文结构、多轮膨胀模型及 provider A/B 实测方法见
[Prompt-cache mitigation and verification](docs/prompt-cache-mitigation.md)。

</details>

<details>
Expand Down Expand Up @@ -520,7 +556,7 @@ export MEMORY_TENCENTDB_GATEWAY_API_KEY="<与 Gateway 同一份密钥>"
- 💡 **有想法想交流?** 欢迎在 [GitHub Discussions](https://github.com/Tencent/TencentDB-Agent-Memory/discussions) 发起讨论。
- 🛠️ **想贡献代码?** 请先阅读 [CONTRIBUTING.md](./CONTRIBUTING_CN.md)。
- 💬 **想加入交流群?** 扫码加入 **Agent Memory 微信社群**,与早期开发者直接对话。
<p align="center"><img src="https://github.com/user-attachments/assets/6e9dd306-f462-4cae-adc4-3194be25c514" width="200" alt="Agent Memory 微信社群二维码" />
<p align="center"><img src="https://github.com/user-attachments/assets/9f77b819-c85b-4135-8b91-a612e42580a6" width="200" alt="Agent Memory 微信社群二维码" />

---

Expand Down Expand Up @@ -548,4 +584,16 @@ export MEMORY_TENCENTDB_GATEWAY_API_KEY="<与 Gateway 同一份密钥>"
</tr>
</table>

---

## Star 趋势

<p align="center">
<a href="https://www.star-history.com/#Tencent/TencentDB-Agent-Memory&Date">
<img src="https://github.com/user-attachments/assets/413dc899-9eca-48b6-aff6-9a09c2dea79c" alt="Star History Chart" width="600" />
</a>
</p>

---

[MIT](./LICENSE) © TencentDB Agent Memory Team
129 changes: 129 additions & 0 deletions docs/prompt-cache-mitigation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,129 @@
# Prompt-cache mitigation and verification

This note records the acceptance evidence for issue #120 without treating the
original provider-rate correlation as proof of causation. The reporter later
isolated a separate OpenClaw webchat regression, so provider A/B runs must hold
the OpenClaw version and channel constant.

## Context layout

`auto-recall` produces two classes of context:

- stable persona and scene context in `appendSystemContext`;
- per-turn L1 recall in `prependContext`.

The OpenClaw adapter added by #375 can move only the dynamic L1 block to
`appendContext`. It deliberately leaves the core result unchanged for other
hosts.

### Legacy prepend mode

```mermaid
flowchart LR
S["Stable system prompt"] --> P["Stable persona / scene context"]
P --> H["Persisted conversation history"]
H --> D["Dynamic L1 recall for this turn"]
D --> U["Current user message"]
```

The cache-sensitive prefix ends as soon as a dynamic segment differs. With
`recall.showInjected=true`, every dynamic recall block is also persisted in
`H`, so later requests replay stale blocks and reach the differing portion
earlier.

### Cache-friendly OpenClaw mode

```mermaid
flowchart LR
S["Stable system prompt"] --> P["Stable persona / scene context"]
P --> H["History without injected recall blocks"]
H --> U["Current user message"]
U --> D["Dynamic L1 recall in appendContext"]
```

Configure:

```json
{
"recall": {
"injectionMode": "append",
"showInjected": false
}
}
```

This preserves the longest stable prefix available to an OpenClaw host that
supports `appendContext`. It is an opt-in placement change; `prepend` remains
the backward-compatible default. The `before_message_write` hook removes
`<relevant-memories>` before persistence when `showInjected` is false.

## History-growth analysis

Let `R` be the average serialized recall size per turn and `N` the number of
turns. Preserving injected recall adds approximately `N * R` characters to the
stored transcript. Because request `k` re-sends the previous injected blocks,
the additional characters sent over the whole session grow approximately as
`R * N * (N - 1) / 2`.

The deterministic regression fixture uses 100 turns and roughly 1,000 dynamic
recall characters per turn. It verifies that:

- all 100 recall blocks are removed before persistence when `showInjected` is
false;
- at least 100,000 injected characters are removed;
- the cleaned transcript is less than 5% of the `showInjected=true` transcript;
- original user text and non-text multipart content remain unchanged;
- `showInjected=true` remains an explicit diagnostic opt-in.

Run it with:

```bash
COREPACK_ENABLE_AUTO_PIN=0 pnpm vitest run \
src/adapters/openclaw/recall-injection.test.ts \
src/adapters/openclaw/recall-injection.multiturn.test.ts
```

This is a persisted-history proxy, not a provider cache-hit percentage.

## Provider verification protocol

Use a separate A/B session for each provider and repeat the same scripted turn
sequence in both configurations:

1. baseline: `injectionMode=prepend`, `showInjected=true`;
2. mitigation: `injectionMode=append`, `showInjected=false`.

Keep provider account, model, OpenClaw build, channel, tools, system prompt,
recall corpus, and turn sequence identical. Do not compare one channel (for
example webchat) with another (for example Feishu). Warm each configuration
with the same first request, then record per-request input, cache-hit, and
cache-miss tokens. Report medians and aggregate token-weighted hit rate over at
least 30 measured turns; retain raw response usage fields for audit.

| Provider | Official response fields | Token-weighted hit rate |
| --- | --- | --- |
| DeepSeek OpenAI API | `usage.prompt_cache_hit_tokens`, `usage.prompt_cache_miss_tokens` | `sum(hit) / (sum(hit) + sum(miss))` |
| Xiaomi MiMo OpenAI API | `usage.prompt_tokens_details.cached_tokens`, `usage.prompt_tokens` | `sum(cached) / sum(prompt)` |

DeepSeek documents automatic prefix caching and exposes explicit hit/miss token
counts. MiMo exposes cached prompt tokens in its OpenAI-compatible response;
its matching policy should be treated as provider-controlled and verified by
the A/B data rather than assumed.

Official references:

- DeepSeek context caching: <https://api-docs.deepseek.com/guides/kv_cache>
- Xiaomi MiMo OpenAI-compatible chat API:
<https://mimo.mi.com/docs/en-US/api/chat/openai-api>

## Interpretation

An improvement is attributable to this mitigation only when the stable-prefix
A/B changes while the host/channel control stays fixed. Publish the raw token
counts and confidence interval alongside the percentage. If usage fields are
missing or null, mark that request unavailable instead of treating it as a
zero-token cache hit.

The runtime mitigation and deterministic growth test are complete. Live
DeepSeek/MiMo percentages remain environment evidence and require credentials;
they must not be fabricated in repository tests.
Loading