12 KiB
Memory & Tool Architecture Plan v3
Cập nhật sau review: Approach C vẫn là hướng đúng, nhưng bản production nên là Approach C+. Trọng tâm không chỉ là dynamic tool loading, mà là policy catalog rõ ràng, router multi-label, DTO nhỏ, memory budget theo model, và state có TTL/idempotency.
1. Kết luận ngắn
Không nên load toàn bộ tools cho mọi request. Khi domain Loyalty mở rộng từ prototype 14 tools lên 80+ tools, tool schema sẽ ăn phần lớn context window, nhất là các create tool dùng DTO sâu như CampaignRule.
Hướng nên làm: Dynamic Tool Loading + Domain Knowledge Registry, nhưng không filter bằng substring contains. Cần chuyển thành capability policy rõ ràng:
- Router phân loại
Action + Set<Domain> + confidence + workflow phase - Tool policy catalog khai báo domain, operation, risk, phase, aliases
- Per-request capability bundle chỉ expose tools cần thiết
- Create/update tools dùng DTO nhỏ theo từng bước, không expose full domain model
- Memory budget theo model context window, không dùng một con số 8K chung
- Workflow/draft state có TTL, cleanup, idempotency, optimistic locking nếu chạy production
2. Evidence hiện tại trong repo
| Evidence | Ý nghĩa |
|---|---|
loyalty-agent/src/main/java/dev/sonpx/loyalty/agent/core/classifier/IntentClassifier.java đang là rule-based, single-label |
Dễ drop tool cần thiết khi user hỏi vừa campaign vừa rule, hoặc query phụ trong workflow |
loyalty-agent/src/main/java/dev/sonpx/loyalty/agent/core/executor/AgentToolScope.java đang suy capability từ tool name |
Mở rộng bằng contains("campaign")/contains("rule") sẽ dễ false positive/false negative |
loyalty-agent/src/main/java/dev/sonpx/loyalty/agent/core/executor/SimpleAgentExecutor.java hiện gọi tools bằng READ_ONLY |
Chưa có routing scope theo domain/action ở runtime |
loyalty-agent/src/main/java/dev/sonpx/loyalty/agent/core/memory/InMemoryChatMemory.java dùng estimate token và remove từng message |
Có thể cắt lệch một turn user/assistant/tool, không phải whole-turn eviction |
loyalty-agent/src/main/java/dev/sonpx/loyalty/agent/core/config/AgentProperties.java mặc định maxMemoryTokens = 8000 |
Không an toàn cho mọi agent vì chat model có context 8K, query/creator model 32K |
loyalty-agent/src/main/java/dev/sonpx/loyalty/agent/core/workflow/WorkflowExecutor.java dùng ConcurrentHashMap active workflows |
Có cleanup khi success/cancel, nhưng chưa có TTL nền và chưa bền khi restart |
Không thấy loyalty-mcp-server/.../draft/DraftSessionManager.java |
Ghi chú cũ trong context.md/plan về DraftSessionManager 1h TTL là stale, không được coi là fact |
loyalty-mcp-server/src/main/java/dev/sonpx/loyalty/mcp/reward/model/CampaignRule.java có nhiều field và nested DTO |
Full createRule(CampaignRule dto) làm tool schema rất lớn |
3. Problem statement đúng
Prototype hiện tại
14 MCP tools vẫn fit context khá thoải mái. Vấn đề memory chưa phải bottleneck lớn nhất.
Khi scale full domain
Nếu expose 80+ tools, đặc biệt có nhiều create/update tool với nested DTO:
32K context
- System prompt: ~400 tokens
- 80 tool definitions: ~20K-30K tokens
- Conversation memory: ~4K-8K tokens
- Tool results/response: ~3K-5K tokens
=> Dễ overflow hoặc làm model 4B chọn sai tool
Vì vậy bài toán chính là tool selection + schema size, không chỉ conversation memory.
4. Target Architecture: Approach C+
flowchart TB
U["User message"] --> R["Intent router"]
R --> D["Action + domains + confidence + phase"]
D --> P["ToolPolicyCatalog"]
P --> B["Capability bundle"]
B --> L["LLM call with selected tools"]
D --> K["DomainKnowledgeRegistry"]
K --> L
L --> T["MCP tools"]
L --> M["Chat memory"]
W["Workflow state"] --> D
W --> B
Core rule
Mỗi request chỉ được thấy tools đúng domain/action/phase. Unknown intent không được load ALL; chỉ load discovery/read-only tools có rủi ro thấp.
5. Implementation Plan
Phase 1: Policy catalog + safer routing
-
Tạo
ToolPolicyCatalog- Mỗi tool khai báo:
name,domains,operation,risk,workflowPhases,aliases,requiresConfirmation - Default deny nếu tool chưa có policy
- Không suy domain bằng substring tool name
- Mỗi tool khai báo:
-
Nâng
IntentClassifier- Output không chỉ là một enum intent
- Dùng structure kiểu:
record RoutingDecision(
Action action,
Set<Domain> domains,
double confidence,
WorkflowPhase phase,
boolean sideQuery
) {}
-
Update
SimpleAgentExecutor- Thay
AgentToolScope.READ_ONLYcố định bằng capability bundle sinh từRoutingDecision - Nếu confidence thấp: load discovery/read-only tools, không load create/update/delete tools
- Thay
-
Test cần có
- Contract test: mọi
@McpToolphải có policy hoặc bị reject có chủ đích - Routing test: campaign query, rule query, campaign+rule query, unknown query, create workflow side-query
- Safety test: unknown/create-like prompt không được expose create tool khi confidence thấp
- Contract test: mọi
Phase 2: Shrink tool schemas
-
Không expose full model như
CampaignRulecho create/update- Tạo command DTO nhỏ:
CreateRuleDraftCommand,SetRuleFormulaCommand,SetRuleCriteriaCommand,SubmitRuleCommand - Chia theo workflow phase để mỗi step chỉ cần schema nhỏ
- Tạo command DTO nhỏ:
-
Formula nên chọn theo type
- Không load cả 8 formula schemas cùng lúc nếu user đang chọn một formula type
- Dùng discriminated union hoặc staged tool: chọn formula type trước, load schema cụ thể sau
-
Test cần có
- Đo token/schema size trước-sau cho selected tools
- Golden tests cho create rule flow
- Test validation: thiếu required field phải hỏi lại, không tự bịa value
Phase 3: Memory budget theo model
-
Tách budget theo agent/model
- Chat 8K context không nên giữ
maxMemoryTokens = 8000 - Query/creator 32K có thể giữ lớn hơn, nhưng phải trừ tool schema + result budget
- Chat 8K context không nên giữ
-
Sửa eviction theo whole turn
- Không remove lẻ từng message nếu nó làm mất cặp user/assistant/tool result
- Giữ system message, prune theo turn hoặc summary
-
Test cần có
- Memory không vượt budget sau nhiều turn
- Tool-call/result message không bị cắt lệch khiến transcript invalid
Phase 4: Workflow/draft persistence
-
Không dựa vào
ConcurrentHashMapcho production- Thêm TTL nền và cleanup định kỳ nếu vẫn single-node
- Nếu multi-node: dùng DB/Redis/shared store
-
State cần có
conversationId,workflowId,version,expiresAt,owner,idempotencyKey- Optimistic locking/CAS để tránh double submit hoặc update đè
-
Test cần có
- Expired workflow bị cleanup
- Duplicate submit dùng cùng idempotency key không tạo hai record
- Hai request song song không overwrite state sai
Phase 5: DomainKnowledgeRegistry
-
Bắt đầu bằng static registry có version/source
- Mỗi snippet ghi rõ domain, source file/spec, owner, last reviewed date
- Inject theo domain/action, không inject toàn bộ
-
Khi nào mới cần RAG
- Tổng domain knowledge thực sự vượt khả năng maintain bằng static snippets
- Knowledge thay đổi thường xuyên
- Có nhu cầu search tài liệu dài/không cấu trúc
RAG không giải quyết tool bloat. RAG chỉ giải quyết domain knowledge retrieval.
6. Edge Cases cần cover
| Edge case | Risk | Guardrail |
|---|---|---|
| User hỏi nhiều domain trong một câu: "campaign X có rule nào?" | Single-label router chỉ load campaign hoặc rule | Router trả Set<Domain> |
| User đang create flow nhưng hỏi phụ: "có campaign nào active?" | Scope create thiếu read-only tools cần thiết | sideQuery=true cho phép bounded read-only fallback |
Tool name trùng hoặc ambiguous: campaignRule chứa cả campaign và rule |
Substring filter load sai tools | ToolPolicyCatalog explicit |
| Unknown intent nhưng prompt có từ "tạo" | Lỡ expose create tool | Unknown chỉ discovery/read-only; create cần confidence + workflow phase |
| Create tool schema quá lớn | Context overflow, model chọn sai field | Staged command DTOs, schema measurement gate |
| Memory 8K trên model context 8K | Không còn chỗ cho prompt/tools/response | Per-model budget formula |
| Prune lẻ message | Mất tool result hoặc assistant turn, transcript invalid | Whole-turn eviction |
| Restart service giữa workflow | Mất active workflow | Persist state trước production |
| User submit hai lần | Double create | Idempotency key + CAS/version |
| Tool result bị truncate giữa JSON | Model nhận JSON invalid | Truncate structured/result-aware, không substring raw JSON |
| DomainKnowledgeRegistry stale | Bot trả sai nghiệp vụ | Version/source/owner/review date + tests |
7. Notes for implementation
- Ưu tiên Spring AI built-in tool filtering/resolver nếu đang dùng được trong version hiện tại; chỉ tự build filter layer khi cần thêm policy/risk/phase.
- Không dùng
ALLtrong production path trừ admin/debug mode có guard riêng. - Tool policy phải fail closed: tool mới thêm mà chưa khai báo policy thì test fail.
- Với destructive hoặc state-changing tools, cần confirmation/risk flag riêng, không chỉ dựa vào intent.
- Nên log
routingDecision, selected tool names, schema token estimate, latency, tool-call success/failure để benchmark. - Quyết định RAG/multi-agent phải dựa trên số liệu: selected tool count, schema token size, tool accuracy, latency, prompt overflow rate.
- Tất cả con số token hiện tại là estimate. Trước khi implement rộng, cần đo schema thực tế bằng tokenizer/model target.
8. Roadmap cập nhật
gantt
title Memory & Tool Architecture Roadmap v3
dateFormat YYYY-MM-DD
section Phase 1: Routing Safety
ToolPolicyCatalog + fail-closed tests :p1a, 2026-07-23, 2d
RoutingDecision multi-domain/action :p1b, after p1a, 2d
Runtime capability bundle :p1c, after p1b, 1d
section Phase 2: Schema Reduction
Create/update command DTOs :p2a, after p1c, 3d
Formula staged loading :p2b, after p2a, 2d
Schema token measurement gate :p2c, after p2b, 1d
section Phase 3: Memory
Per-model memory budgets :p3a, after p1c, 1d
Whole-turn eviction :p3b, after p3a, 2d
section Phase 4: Workflow Production
TTL cleanup :p4a, after p3b, 1d
Persistent workflow store :p4b, after p4a, 3d
Idempotency + optimistic locking :p4c, after p4b, 2d
section Phase 5: Knowledge
Versioned DomainKnowledgeRegistry :p5a, after p2c, 2d
Benchmark RAG/multi-agent necessity :p5b, after p5a, 2d
9. Final verdict
Approach C là đúng cho prototype và giai đoạn scale gần, nhưng chưa đủ để gọi là production best practice nếu chỉ implement bằng contains() và single-intent routing.
Best practice nên chốt là Approach C+:
- Dynamic tool loading theo policy explicit
- Router multi-label theo action/domain/phase
- Small command DTOs thay vì full domain DTO
- Memory budget theo model + whole-turn eviction
- Workflow state có TTL/persistence/idempotency
- Domain knowledge có version/source, chưa cần RAG cho tới khi có số liệu chứng minh