Files
loyalty-agent-service/ARCHITECTURE.md

171 lines
6.6 KiB
Markdown

# Architecture — Loyalty Agent Service
## System Overview
AI-powered assistant for managing a Loyalty Platform, built on the **Model Context Protocol (MCP)**.
```text
┌──────────────────┐
│ React/Vite UI │
│ :9333 │
└────────┬─────────┘
│ WebSocket / STOMP
▼
┌──────────────────┐
│ Loyalty Agent │
│ :9332 │
│ │
│ LLM Orchestrator │
│ MCP Client │
│ Conversation │
│ PostgreSQL │
└────────┬─────────┘
│ MCP SSE (/sse, /message)
▼
┌──────────────────┐
│ Loyalty MCP │
│ Server :9331 │
│ │
│ 25+ MCP Tools │
└────────┬─────────┘
│ REST + OAuth2 (Client Credentials)
▼
┌──────────────────┐
│ Loyalty Core API │
│ :8081 │
└──────────────────┘
Keycloak ────── OAuth2 Token Provider
PostgreSQL ──── Conversation Persistence
Ollama ──────── Local LLM (Qwen3.5:4b)
```
## Technology Stack
| Technology | Version | Purpose |
|----------------|---------|----------------------------|
| Java | 25 | Runtime |
| Spring Boot | 4.0.0 | Framework |
| Spring AI | 2.0.0 | LLM & MCP integration |
| Ollama | latest | Local LLM hosting |
| Qwen | 3.5:4b | Primary LLM model |
| PostgreSQL | 15+ | Conversation persistence |
| Keycloak | latest | OAuth2/OIDC provider |
| React + Vite | 19/6 | Frontend UI |
## Module Architecture
### loyalty-mcp-server (`:9331`)
MCP Server exposing Loyalty Platform capabilities as MCP tools via SSE transport.
**Responsibilities:**
- Register 25+ MCP tools (campaign, rule, customer, transaction, etc.)
- Validate tool inputs
- Proxy requests to Loyalty Core API with OAuth2 client credentials
- Return structured `Result<T>` with agent instructions for LLM formatting
- Handle errors (sanitize stack traces, structured error responses)
**Key Classes:**
- `CampaignTools` / `CampaignRuleTools` — Primary tool registration
- `CampaignService` / `CampaignRuleService` — Business logic
- `RestClientConfig` — OAuth2 REST client factory
- `Result<T>` — Standard response wrapper with `_agent_instruction`
### loyalty-agent (`:9332`)
AI Agent / BFF — Orchestrates LLM, MCP, conversations, and WebSocket.
**Responsibilities:**
- Accept user messages via WebSocket/STOMP
- Classify intent (regex-based, zero LLM latency)
- Route to appropriate executor (SimpleAgent, Workflow, DirectChat)
- Stream LLM responses back to frontend
- Persist conversations to PostgreSQL
**Key Architecture:**
```text
AgentController (@MessageMapping /chat)
↓
LoyaltyAgentService
↓
AgentOrchestrator
├── Active Workflow? → WorkflowExecutor (multi-turn creation)
├── IntentClassifier → QUERY → SimpleAgentExecutor (ReAct + tools)
├── IntentClassifier → CONVERSATION → SimpleAgentExecutor (direct chat)
└── IntentClassifier → CREATE_* → WorkflowExecutor
```
**Intent Classification (regex-based):**
| Intent | Trigger Examples |
|--------|-----------------|
| `QUERY` | "tìm chiến dịch", "danh sách rule" |
| `CONVERSATION` | "xin chào", "cảm ơn" |
| `CREATE_CAMPAIGN` | "tạo chiến dịch mới" |
| `CREATE_RULE` | "thêm thể lệ" |
**Tool Scope Enforcement:**
| Scope | Allowed Operations |
|-------|-------------------|
| `READ_ONLY` | search, get, find, count, check |
| `CREATE_CAMPAIGN` | READ_ONLY + createCampaign + generateCampaignId |
| `CREATE_RULE` | READ_ONLY + createRule + generateRuleId |
### Frontend (`:9333`)
React/Vite conversational UI with STOMP WebSocket.
**Key Features:**
- STOMP over WebSocket for real-time messaging
- Token streaming (character-by-character LLM output)
- Tool execution status indicators
- Conversation management (create, list, delete, rename)
- Offline message queue with exponential backoff reconnect
## Security Architecture
```text
Frontend ──(WebSocket)──> Agent ──(MCP SSE)──> MCP Server ──(OAuth2)──> Core API
│
▼
Keycloak
```
**Principles:**
1. **LLM is never an authorization boundary** — tool scope is enforced by `AgentToolScope` before tools reach the LLM
2. **OAuth2 Client Credentials** — MCP Server authenticates to Core API via Keycloak
3. **Tool result sanitization** — `ToolInterceptor` strips Java stack traces
4. **Input validation** — All tool parameters validated before downstream calls
## Data Flow
### Chat Request Flow
```text
1. User types message in React UI
2. STOMP message → /app/chat → AgentController
3. AgentController → LoyaltyAgentService.chat()
4. AgentOrchestrator classifies intent
5. Executor starts LLM streaming with tools
6. LLM selects tool → ToolInterceptor wraps & calls MCP
7. MCP Server executes tool → calls Core API
8. Tool result → LLM reasons → generates response
9. Response streamed back via STOMP → /user/queue/chat-events
10. React UI renders tokens in real-time
```
## Resilience Patterns
| Pattern | Implementation |
|---------|---------------|
| MCP auto-reconnect | `ToolInterceptor.attemptReconnect()` |
| WebSocket reconnect | `StompService` exponential backoff + jitter |
| Request timeout | `SimpleAgentExecutor` configurable via `agent.request-timeout-seconds` |
| Tool result truncation | `ToolInterceptor` caps results > 4000 chars |
| Error sanitization | `ToolInterceptor` strips stack traces |
| Graceful shutdown | `server.shutdown: graceful` on both modules |
| LLM search noise filtering | `SearchUtils.sanitizeSearch()` |