Files
gori-agent/README.md
T
2026-08-15 16:56:58 +08:00

255 lines
9.5 KiB
Markdown

# gori-agent-gateway
Hermes-gateway-style multi-IM Agent Gateway for routing instant-message webhooks to configurable CLI agents. The primary product flow is `gori-agent setup`: discover local agents first, pick one, then choose the IM platform to connect.
## Quick start
Requires Node.js 20+.
```bash
npm install
npm run build
./install.sh
gori-agent setup
```
`install.sh` installs the command wrapper under `~/.gori-agent/bin`, stores runtime PID/log files under `~/.gori-agent/state` and `~/.gori-agent/logs`, and adds the command to `~/.bashrc`. Open a new terminal or run `source ~/.bashrc` before using `gori-agent` directly.
If the package bin has not been linked, use the no-link fallback:
```bash
npm run gori-agent -- setup
```
The setup wizard writes `config.json` only after confirmation. `config.json` is ignored by git and should hold local secrets.
Start the gateway after setup:
```bash
gori-agent start --config ./config.json
```
Fallback:
```bash
npm run gori-agent -- start --config ./config.json
```
Convenience executable:
```bash
gori-agent setup
gori-agent start
gori-agent start-daemon
gori-agent status
gori-agent logs
gori-agent stop
```
If the command is not linked into PATH, run the project-local wrapper:
```bash
./gori-agent.sh status
```
Use `--config path` after the command to override the config path.
## CLI commands
```bash
gori-agent setup
gori-agent discover-agents [--json]
gori-agent start [--config path]
gori-agent status [--config path]
gori-agent doctor [--config path]
gori-agent print feishu [--config path]
gori-agent --help
```
Package scripts mirror common commands:
```bash
npm run setup
npm run discover-agents
npm run doctor
```
## Agent-first setup flow
`gori-agent setup` does the following:
1. Discovers local CLI agents on `PATH` and always includes the built-in `echo` fallback.
2. Lets you pick one agent.
3. Lets you choose an IM platform: Feishu/Lark, WeChat external webhook, WeCom scaffold, QQ Bot webhook, or Generic webhook.
4. Builds a focused `config.json` containing the chosen default agent, the chosen agent plus `echo`, and the chosen platform enabled.
5. Asks before writing.
Detected ready agents with known safe prompt modes:
| Agent | Command | Args | Input mode | Permission mode |
| --- | --- | --- | --- | --- |
| Kimi | `kimi` | `-p` | final argv | `kimi -p` already runs non-interactively under Kimi Code's auto permission policy. `--yolo` cannot be combined with `-p`. |
| OpenCode | `opencode` | `run` | final argv | setup can append `--auto`. |
| Codex | `codex` | `exec` | final argv | default; add manual args if needed. |
| Claude | `claude` | `-p` | final argv | default; add manual args if needed. |
| Built-in echo | `node` | `scripts/echo-agent.js` | stdin | n/a |
When you pick Kimi, setup asks which Kimi Code model to use. It reads your local Kimi model aliases at setup time with `kimi provider list --json`, so custom providers from your actual Kimi config appear in the menu. Choosing the default keeps Kimi Code's configured `default_model`; choosing a model writes `-m <model> -p` into that agent's args. You can also enter a custom model alias.
Gemini, Qwen/Qwen Code, Copilot, Pi, and Hermes are probed with `--version` then `--help` only. If found, they are shown as `needs-config` unless a safe prompt mode is known; the wizard asks you to confirm command, args, input mode, extra auto/yolo permission args, and working directory before enabling them.
Discovery never sends an actual prompt to an agent. Probes use `child_process.spawn(..., { shell: false })`, a 2s timeout, and a 4KB output cap.
## Platform status
| Platform | Inbound | Outbound | Notes |
| --- | --- | --- | --- |
| Feishu/Lark | Implemented | Implemented | Replies to `im.message.receive_v1` via `/im/v1/messages/{message_id}/reply`. |
| WeCom | 501 scaffold | Implemented | Uses `gettoken` and `message/send`; inbound callback verification/encryption is not in v1. |
| personal WeChat | External webhook scaffold | Synchronous webhook response | Native iLink/personal WeChat integration is not included in v1. |
| QQ | Implemented | Implemented | Supports QQ official Bot WebSocket gateway mode and optional HTTP callback mode. |
| Generic webhook | Implemented | Synchronous JSON | HMAC-SHA256 signed JSON endpoint for local bridges and tests. |
## Hermes reference
This project follows the Hermes gateway pattern: platform adapters normalize inbound messages, a central gateway applies policy/session/concurrency handling, then replies are sent through the originating adapter. The Feishu adapter ports the key Hermes behavior for tenant token caching, URL challenge handling, token verification, `im.message.receive_v1` parsing, mention cleanup, and message replies.
## Endpoints
- `GET /health`
- `GET /platforms`
- `POST /webhook/feishu`
- `POST /webhook/wecom`
- `POST /webhook/qq`
- `POST /webhook/generic`
- `POST /webhook/weixin`
## Configuration
Server execution still supports the existing environment variable:
```bash
GORI_GATEWAY_CONFIG=./config.json npm start
```
The CLI resolves config in this order:
1. `--config path`
2. `GORI_GATEWAY_CONFIG`
3. `./config.json`
4. seed from `config.example.json` and write to `./config.json` if confirmed
Important sections:
- `server.host`: bind address, default `0.0.0.0`.
- `server.port`: gateway port, default `3000`.
- `server.publicBaseUrl`: public HTTPS base URL used by print/setup hints, for example `https://agent.example.com`.
- `policy.allowedUsers`: allow only listed normalized user IDs when non-empty.
- `policy.allowedChats`: allow only listed normalized chat IDs when non-empty.
- `policy.requireMentionInGroup`: if true, group messages are ignored unless the adapter reports a bot mention.
- `defaultAgent`: selected when a chat has not chosen an agent.
- `agents[]`: named CLI agents using `child_process.spawn(command, args)` with `shell: false`.
CLI agent options:
- `inputMode: "stdin"`: sends the message text to stdin.
- `inputMode: "arg"`: appends the message text as the final argv item.
- `timeoutMs`: kills slow processes.
- `outputMaxBytes`: caps captured stdout/stderr.
- `cwd`: optional working directory for the agent process.
## Chat commands
The gateway handles these commands per chat before invoking an agent:
- `/help`
- `/status`
- `/agents`
- `/agent <name>`
- `/new`
## Feishu setup
The wizard can collect Feishu values and `gori-agent print feishu --config ./config.json` prints the webhook URL and checklist.
Manual steps:
1. Create a Feishu/Lark custom app and enable bot messaging.
2. Configure event subscription for `im.message.receive_v1`.
3. Set the request URL to `https://<host>/webhook/feishu`.
4. Put `appId`, `appSecret`, and `verificationToken` into your config.
5. Add bot display names to `platforms.feishu.botNames` so group mention detection can also work from flattened text.
The adapter accepts Feishu URL verification challenges and returns `{ "challenge": "..." }`.
## QQ setup
QQ has two connection modes:
- `websocket` (default/recommended): the gateway actively connects to QQ with OAuth access token and WebSocket. No public domain or callback URL is needed.
- `webhook`: QQ posts events to a public HTTPS callback URL.
Manual WebSocket steps:
1. Create a QQ official Bot.
2. Put `appId` and `clientSecret` into your config.
3. Keep `platforms.qq.connectionMode` as `"websocket"`.
4. Keep the default `intents` value `33554432` (`1 << 25`, `GROUP_AND_C2C_EVENT`) to receive `GROUP_AT_MESSAGE_CREATE` and `C2C_MESSAGE_CREATE`.
5. Run `gori-agent start --config ./config.json`; the process should log `QQ websocket ready` after successful authentication.
Manual HTTP callback steps:
1. Set `platforms.qq.connectionMode` to `"webhook"`.
2. Configure the callback URL to `https://<host>/webhook/qq`. QQ callback URLs must use an allowed public HTTPS port such as 443, 8443, 8080, or 80.
3. Put `botSecret` into your config and keep `verifySignature` enabled so `X-Signature-Ed25519` callbacks are verified.
The webhook adapter handles QQ `op: 13` callback URL validation and returns `{ "plain_token": "...", "signature": "..." }`. `GROUP_AT_MESSAGE_CREATE` and `C2C_MESSAGE_CREATE` callbacks return QQ HTTP callback ACK `{ "op": 12 }` immediately, then reply through the QQ Bot group/C2C message APIs.
## Generic webhook
Payload:
```json
{
"chat_id": "demo-chat",
"user_id": "demo-user",
"text": "hello",
"message_id": "optional-message-id",
"is_group": false,
"mentions_bot": true
}
```
Signature header:
- Header name: `X-Gori-Signature`
- Format: `sha256=<hex>` or just `<hex>`
- Input: exact JSON request body bytes as sent by the client
- Algorithm: HMAC-SHA256 using `platforms.webhook.secret`
Example:
```bash
body='{"chat_id":"demo-chat","user_id":"demo-user","text":"/status"}'
sig=$(printf '%s' "$body" | openssl dgst -sha256 -hmac 'replace-me' -hex | awk '{print $2}')
curl -sS http://localhost:3000/webhook/generic \
-H 'Content-Type: application/json' \
-H "X-Gori-Signature: sha256=$sig" \
-d "$body"
```
Response shape:
```json
{ "ok": true, "output": "...agent or command output..." }
```
## Security notes
- Do not expose the gateway publicly without HTTPS and upstream authentication/rate limits.
- Keep platform secrets outside source control; use a copied config file or secret manager.
- CLI agents are spawned without a shell, but they can still execute arbitrary local code. Only configure trusted commands.
- Use `allowedUsers` and `allowedChats` in production.
- Keep `requireMentionInGroup` enabled for group chats to avoid accidental agent invocation.
- Set conservative `timeoutMs` and `outputMaxBytes` values for external-facing agents.