feat: run confirmed proposals continuously until a real decision point

A proposal still needs one confirmation before it starts, but once it is
working the worker now treats that as a single grant to carry out
routine low-risk execution without step-by-step reconfirmation. It only
returns pending early for high-risk actions, key business decisions,
external blockers, or dirty/unexpected targets. Assistant/worker
bootstrap copy and docs now align with that execution model.
This commit is contained in:
zenord
2026-08-22 00:05:40 +08:00
parent 5229059fe3
commit 9ae2639660
5 changed files with 10 additions and 5 deletions
+1 -1
View File
@@ -209,7 +209,7 @@ Kimi Code agent `args` 必须严格为 `["acp"]`,不能添加可能绕过 assi
运行链路是三层:每 conversation(chat + user)一个无工具 Assistant 会话负责对话与创建 Proposal;唯一活跃 Worker 在 `bot.workspace` 执行已确认 Proposal;Proposal 是两者之间的持久工作单,owner 是 chat + 发起用户。`confirm` / `adjust` / `follow_up` / `start_next` 维持 owner-only(runtime 强校验);为了避免单人把单 worker 队列卡死,`finish` / `stop` / `cancel` 是全板共享动作,任何用户都可对可见 proposal 执行。Proposal 板全局共享可见:Assistant prompt 的全板列表包含所有 chat/user 的未完成条目(`scope=own` / `scope=other` + chat 类型,不暴露 openid),Assistant 可向任何用户如实描述全板,并可按规则对共享动作发起 action。
Proposal 状态流:`proposed → queued → working → pending → finished`。`proposed` 必须用户确认才进入 `queued`;`pending` 就是「等用户决定」,不区分 success/failure;只有 `finish` 把 pending 落定为 `finished(done)`,`cancel`(proposed/queued/pending)落定为 `finished(cancelled)`;没有 working、owner 自己没有 pending、全局没有 workspaceDirty pending 时,最早确认的 queued Proposal 才可通过 confirm 或 start_next 开始。
Proposal 状态流:`proposed → queued → working → pending → finished`。`proposed` 必须用户确认才进入 `queued`;proposal 一旦进入 `working`,worker 默认连续执行普通低风险步骤,不为工作区内读写、搜索、本地构建或测试逐步回问,只有遇到高风险动作、关键业务选择、外部阻塞或 dirty/异常目标时才返回 `pending question`;`pending` 就是「等用户决定」,不区分 success/failure;只有 `finish` 把 pending 落定为 `finished(done)`,`cancel`(proposed/queued/pending)落定为 `finished(cancelled)`;没有 working、owner 自己没有 pending、全局没有 workspaceDirty pending 时,最早确认的 queued Proposal 才可通过 confirm 或 start_next 开始。
Assistant 每条回复以隐藏 `GORI_ASSISTANT_ACTION_V2` envelope 结尾(`reply` + `actions`),action 仅 `create_proposal`、`confirm`、`adjust_proposal`(仅 proposed/queued)、`follow_up`(pending → working,优先 resume 原 worker native session,失败则带 Proposal 上下文新起 session,用户图片附件随 prompt 给 Worker)、`finish`、`send_image`(把 Worker 报告过的 workspace 内图片发给用户)、`start_next`、`cancel`、`stop`,格式错误只修复一次。Worker 每轮以隐藏 `GORI_WORKER_RESULT_V2` envelope 收尾:仅 `PENDING`(`summary` 必填,可选 `question`、`workspaceDirty`、`attachments`),同样只修复一次。envelope 缺闭合标签(模型截断)时先按花括号配平 salvage(字符串/转义感知,配平点必须在文本末尾),失败才进入修复流程。
+1 -1
View File
@@ -196,7 +196,7 @@ QQ 出站图片:Worker/Assistant 报告的 workspace 内图片(png/jpg、≤
运行链路是三层:
- **Assistant**:每个 conversation(chat + user)一个无工具 ACP 会话,只与用户对话;同群不同用户的会话互相隔离。它把用户意图整理成 Proposal(title、goal、steps),并解释 Worker 的反馈。Assistant 每条回复必须以隐藏 `GORI_ASSISTANT_ACTION_V2` envelope 结尾(`reply` + `actions`),action 只有 `create_proposal`、`confirm`、`adjust_proposal`、`follow_up`、`finish`、`send_image`、`start_next`、`cancel`、`stop`;格式错误只修复一次(envelope 缺闭合标签时先按花括号配平 salvage,失败才修复)。`send_image { path }` 把 Worker 报告过的 workspace 内图片发给用户,与 Worker 附件同样的路径/类型/大小校验。
- **Proposal**:一份工作单,owner 是 chat + 发起用户;`confirm` / `adjust` / `follow_up` / `start_next` 仍只有发起人本人可操作(runtime 强校验),但为了避免单人把单 worker 队列卡死,`finish` / `stop` / `cancel` 是全板共享动作:任何用户都可对可见 proposal 执行它们。Proposal 板全局共享可见:Assistant prompt 的全板列表包含所有 chat/user 的未完成条目(自己的标 `scope=own`,他人的标 `scope=other` 并附 chat 类型,不暴露 openid 明文),Assistant 可如实向任何用户描述全板状态。状态流为 `proposed → queued → working → pending → finished`。`proposed` 只有用户确认后才进入 `queued`;`pending` 就是「等用户决定」,不再区分 success/failure;只有 `finish` 把 pending 落定为 `finished(done)`,`cancel` 落定为 `finished(cancelled)`。
- **Proposal**:一份工作单,owner 是 chat + 发起用户;`confirm` / `adjust` / `follow_up` / `start_next` 仍只有发起人本人可操作(runtime 强校验),但为了避免单人把单 worker 队列卡死,`finish` / `stop` / `cancel` 是全板共享动作:任何用户都可对可见 proposal 执行它们。Proposal 板全局共享可见:Assistant prompt 的全板列表包含所有 chat/user 的未完成条目(自己的标 `scope=own`,他人的标 `scope=other` 并附 chat 类型,不暴露 openid 明文),Assistant 可如实向任何用户描述全板状态。状态流为 `proposed → queued → working → pending → finished`。`proposed` 只有用户确认后才进入 `queued`;proposal 一旦进入 `working`,worker 默认连续执行普通低风险步骤,不再逐步回问,只有碰到高风险动作、关键业务选择、外部阻塞或 dirty/异常目标时才返回 `pending question`;`pending` 就是「等用户决定」,不再区分 success/failure;只有 `finish` 把 pending 落定为 `finished(done)`,`cancel` 落定为 `finished(cancelled)`。
- **Worker**:同一时刻全实例只有一个,在 `bot.workspace` 用 `bot.permissions` policy 执行一个已确认 Proposal。每轮必须以隐藏 `GORI_WORKER_RESULT_V2` envelope 收尾:`PENDING`(`summary` 必填,可带 `question`、`workspaceDirty`),不区分成功/失败,只把结果交给用户。Worker 给用户看的图片(png/jpg)必须保存在 workspace 内(建议 `.gori-outbox/`),并通过 `attachments: [{ path, mimeType? }]`(最多 3 个)上报;runtime 校验路径必须在 workspace 内、magic bytes 为 png/jpg、单张 ≤10MB,违规的丢弃并在事件文本里说明。
确认语义是刻意的:
+1 -1
View File
@@ -1091,7 +1091,7 @@ function workerTaskPrompt(proposal: Proposal, resumeContext?: string): string {
...proposal.steps.map((step, index) => `${index + 1}. ${step}`),
...(resumeContext ? ["", "Additional context:", resumeContext] : []),
"",
"If you encounter a dirty or unexpected target, or anything that requires the user's decision, stop and report PENDING with a clear question instead of forcing the change."
"Keep executing routine low-risk steps inside the confirmed proposal without pausing for step-by-step confirmation. Only stop and report PENDING with a clear question when you hit a high-risk action, a key business/product choice, an external blocker, or a dirty/unexpected target that genuinely needs the user's decision."
].join("\n");
}
+3 -2
View File
@@ -2,7 +2,7 @@ import crypto from "node:crypto";
import type { AppConfig, BotConfig } from "../config.js";
import { SkillLoader, type LoadedSkill } from "./skill-loader.js";
const BOOTSTRAP_SCHEMA_VERSION = 11;
const BOOTSTRAP_SCHEMA_VERSION = 12;
export interface ResolvedBot extends BotConfig {
loadedSkills: LoadedSkill[];
@@ -47,7 +47,7 @@ function buildAssistantBootstrap(bot: BotConfig): string {
"Speaking style: talk like a reliable colleague, not a console. Lead with the conclusion, then the reason, then the next step. Avoid protocol jargon and field names; never expose internal words like scheduler, pending, envelope, or action types to the user. Default to 2-4 sentences. Do not repeat proposal IDs unless the user asks. If you are unsure, say so plainly. When something is blocked, always give the user an actionable next step.",
"Every reply must end with exactly one hidden action envelope: <GORI_ASSISTANT_ACTION_V2>{\"reply\":\"...\",\"actions\":[...]}</GORI_ASSISTANT_ACTION_V2>. Put the user-facing text in the JSON \"reply\" field, not outside the envelope.",
"Supported actions: {\"type\":\"create_proposal\",\"title\":\"...\",\"goal\":\"...\",\"steps\":[\"...\"]}; {\"type\":\"confirm\",\"id\":\"optional\"}; {\"type\":\"adjust_proposal\",\"id\":\"optional\",\"title\":\"optional\",\"goal\":\"optional\",\"steps\":\"optional\"}; {\"type\":\"follow_up\",\"id\":\"optional\",\"instruction\":\"...\"}; {\"type\":\"finish\",\"id\":\"optional\",\"note\":\"optional\"}; {\"type\":\"send_image\",\"path\":\"...\"}; {\"type\":\"start_next\"}; {\"type\":\"cancel\",\"id\":\"optional\"}; {\"type\":\"stop\"}. Use an empty actions array when no state change is needed.",
"A proposal only starts after the user confirms it. When a worker finishes a turn, the proposal becomes pending: it waits for the user's decision with a summary (and maybe a question). The worker never reports success or failure; treat every result as information for the user. \"finish\" closes a pending proposal as done; \"follow_up\" sends the user's new instruction to the same proposal and resumes its worker; \"cancel\" drops a proposed, queued, or pending proposal; \"stop\" aborts the running worker and leaves the proposal pending; \"start_next\" starts the oldest confirmed queued proposal. \"adjust_proposal\" edits a proposal that has not started yet. Never invent other actions or statuses.",
"A proposal only starts after the user confirms it. Once a proposal is working, do not ask the user to re-confirm ordinary low-risk next steps. The worker should keep executing routine low-risk work until it reaches a real decision point and then return pending with a question. When a worker finishes a turn, the proposal becomes pending: it waits for the user's decision with a summary (and maybe a question). The worker never reports success or failure; treat every result as information for the user. \"finish\" closes a pending proposal as done; \"follow_up\" sends the user's new instruction to the same proposal and resumes its worker; \"cancel\" drops a proposed, queued, or pending proposal; \"stop\" aborts the running worker and leaves the proposal pending; \"start_next\" starts the oldest confirmed queued proposal. \"adjust_proposal\" edits a proposal that has not started yet. Never invent other actions or statuses.",
"When a worker's summary says it saved image files inside the workspace (for example under .gori-outbox/) and the user asks to see one, emit \"send_image\" with that exact path. Only send paths a worker actually reported; never invent paths.",
"Whenever the [Pending proposals] section in a prompt lists one of the user's proposals, your reply MUST acknowledge it: remind the user what is waiting and that they can say finish to close it or just keep talking to continue it.",
"Never claim an action has already taken effect. The runtime executes your actions after your reply and appends a correction to your message when something could not be done (for example when start_next is blocked by another proposal). Treat that correction as the truth and use the [Proposal states], [Pending proposals], [Scheduler state] and [Worker state] sections in each prompt as the only reliable state.",
@@ -65,6 +65,7 @@ function buildWorkerBootstrap(bot: BotConfig, skills: LoadedSkill[]): string {
bot.persona ? `Persona:\n${bot.persona}` : "Persona: general coding assistant",
`Permission policy enforced by the ACP client: ${JSON.stringify(bot.permissions)}`,
"You are the Worker. You execute exactly one confirmed proposal in the configured workspace and never talk to the user directly.",
"A confirmed proposal is a single permission grant to carry out routine low-risk execution inside the proposal scope. Do not stop for step-by-step confirmation during ordinary low-risk work such as reading files, searching, editing inside the workspace, running local builds, tests, lint, or typecheck, and following the proposal's stated steps. Only return control early when you hit a real decision point: a high-risk action, a key product choice, an external blocker, or an unexpected/dirty target that needs the user's call.",
"For every turn, end your response with exactly one hidden result envelope: <GORI_WORKER_RESULT_V2>{\"status\":\"PENDING\",\"summary\":\"...\"}</GORI_WORKER_RESULT_V2>. PENDING is the only status: it hands the result back to the user. Always include a short user-readable \"summary\" of what you did or what is blocking you. Add a \"question\" when you need the user's decision before continuing. Set \"workspaceDirty\": true when you left the workspace modified or are unsure about its state. When you produced image files the user should see (png/jpg only), save them inside the workspace (prefer .gori-outbox/) and report up to 3 as \"attachments\": [{\"path\":\"relative/or/absolute/path\",\"mimeType\":\"optional\"}]; never report paths outside the workspace such as /tmp. Never emit any other status or text after the envelope.",
"Keep every command and tool process attached to this ACP worker. Never daemonize, call setsid, use nohup, create a detached process, or leave a background process running after the turn.",
...skills.map((skill) => `Skill ${skill.id} (${skill.file}):\n${skill.content}`)
+4
View File
@@ -38,6 +38,10 @@ test("assistant bootstrap carries human speaking-style rules, worker bootstrap d
assert.match(bot.assistantBootstrap, /conclusion, then the reason, then the next step/);
assert.match(bot.assistantBootstrap, /2-4 sentences/);
assert.match(bot.assistantBootstrap, /actionable next step/);
assert.match(bot.assistantBootstrap, /proposal only starts after the user confirms/i);
assert.match(bot.assistantBootstrap, /do not ask the user to re-confirm ordinary low-risk next steps/i);
assert.match(bot.workerBootstrap, /single permission grant to carry out routine low-risk execution/i);
assert.match(bot.workerBootstrap, /Do not stop for step-by-step confirmation during ordinary low-risk work/i);
assert.doesNotMatch(bot.workerBootstrap, /reliable colleague/);
});