The global BMAD rules call for relevant verification at each stage; a historical project plan went further by requiring physical verification after every change. One log shows package installation before tests, but does not establish that each run downloaded 198.1 MiB.
Published byImages generated with GPT Image 2
Research answer
![[GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the](https://d15bya8fi85z1v.cloudfront.net/chat/QuPfET7ftYd6An8JH1fI8Jq25Ow2/thumbnails/4264E25748C8AF3AAF71/37E559A30B81BD3A5EDF-medium-640.webp)
Create a landscape editorial hero image for this Studio Global article: [GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the. Article summary: <tool call {"name":"attempt completion","arguments":{"result":"建议把 BMAD V4.2 的“物理背压”改成“隔离强度不降级、验证范围分级、环境准备独立、证据按需进入上下文”。当前证据指向的不是“测试太严格”,而是阶段粒度被细化、测试入口携带环境安装、详细输出直接回灌,以及执行记录反复对账共同放大了成本。\n\n本次仅交付只读审计和替换条款:不修改规则文件,不运行测试或部署. Topic tags: general web, workflow, code, data. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake
BMAD V4.2 validation can be made less costly without relaxing its isolation or safety requirements. The central recommendation is to separate environment preparation from test execution, match verification to the scope of a change, and keep detailed logs in an artifact channel rather than sending them wholesale into the working context.
This is a read-only audit and a set of proposed replacement rules. It does not modify rule files, run tests or deployments, or approve or archive a plan.
The global BMAD rule requires relevant verification at each stage (02-bmad-core.md, section 3). The engineering-mode rule also allows either the current host or an authorized container, and says verification should reflect the scope of the change (01-bmad-engineer-core.md).
A stronger requirement appears in a historical project plan: “physical verification after every change” (plan clause). So it would be inaccurate to say that the global BMAD rules themselves require repeated environment installation or a complete test matrix for every patch.
A more plausible explanation is that the global rules leave stage size and preparation costs undefined; a project plan then tightens the cadence to every change; and the same execution path handles both environment preparation and testing. Repeated setup, output and evidence bookkeeping can compound from there.
The aurora.log records 14 packages being installed before test events begin. At the end of the installation section, it reports 198.1 MiB for 31 packages. That total does not establish how much was downloaded—or newly unpacked—in this particular run (log).
The operation began at about 14:04:04, with test events starting at 14:04:09.330, leaving roughly 5.3 seconds for the preceding steps (timestamps). Package-level tests took about 2.72 seconds; the whole operation took about eight seconds (completion record). The available timestamps do not let us separate installation, startup, compilation or cache checks within that preceding time.
The evidence supports a narrower conclusion: the test entry point had noticeable preparation overhead. It does not quantify installation’s exact share of that overhead.
The installation section takes about 16 lines. The test output, by contrast, repeats start and pass information in both text and event form (example). The supplied log also contains an ellipsis marker representing roughly 76,000 characters of omitted output (truncation point).
Suppressing package-install messages alone would not address a flood of repeated test events. Nor does a chat export or interface display prove that all of the displayed text entered a model request or allow actual token usage to be calculated. The evidence is insufficient for either claim.
The exported history includes separate compilation and test calls, including calls associated with different container identities (fixed-state verification record). Yet the recorded test entry point still runs a Go test command that can involve build preparation (container command).
The isolation script itself was not available for this audit. It is therefore not possible to identify whether installation occurred in the script, a container entry point or another wrapper—or to conclude that every historical call installed packages again. A careful audit should distinguish three things: an installation observed in one run, multiple container calls in the history, and a claim that installation recurred on every call. Only the first two are supported here.
The proposed fix starts by separating concerns that are easy to blur together:
The historical plan requires rechecking an updated document binding and repeatedly saving execution-history summaries (binding history; recheck after recording). Those steps can protect against tampering. But if contract inputs and execution results are not distinguished, they also create a structural risk of a governance loop: record a result, change a document, rerun verification, then record again. The audit identifies that risk; it does not establish that an infinite loop occurred.
The key distinction is simple: reusing an immutable toolchain does not mean reusing a dirty test workspace. Rebuilding a temporary test workspace does not require reinstalling the toolchain.
The audit also recommends replacing claims of “absolute safety” with verifiable boundaries: block external network access and unauthorized persistent writes, while allowing bounded temporary space and designated evidence artifacts. A race-detector pass covers only the paths actually exercised; it cannot prove that a program has no data races. 9
This is a proposed contract, not a description of an existing implementation.
| Tier | Purpose | When it runs—and what should not trigger it |
|---|---|---|
| L0: Immutable execution environment | Provides approved toolchains, system libraries and dependencies; records image identity, toolchain versions and security settings. | Prepare a new version when environment inputs change. A business-source edit should not trigger system-package installation. |
| L1: Coding feedback loop | Runs relevant unit and module checks, plus targeted race or regression checks when warranted. | Run once per semantically complete change batch, using an approved isolated environment by default. Adjust test scope, not the isolation boundary. |
| L2: Stage delivery gate | Runs the required matrix against a frozen candidate, including independent, uninstrumented RSS checks and evidence reconciliation. | Run at stage completion or before an authorized delivery—not again for an ordinary status update. |
The historical authorization supplied for this audit specifies offline verification in an isolated container (authorization scope). A host-based quick test therefore should not become the default substitute without separate, explicit authorization and defined conditions.
The recommended default is a lightweight feedback loop under the same security constraints. Saving time alone is not grounds for silently switching to host execution.
The audit recommends rules based on the kind of change, not the number of edited files.
| Change or event | Required action | Should not automatically trigger |
|---|---|---|
| Implementation and assertions for one behavior are complete | One relevant L1 verification run | Rebuilding L0 or running the full L2 gate |
| Locks, concurrent access, cancellation or resource release change | Add targeted race and lifecycle checks to the current batch | Deferring race checks until final delivery |
| A cross-module interface or shared dependency changes | Expand checks to affected call paths | Testing only the edited file without justification |
| Toolchain, system dependency or isolation configuration changes | Reconfirm L0 and reassess affected evidence | Installing packages temporarily in the test entry point |
| A stage-delivery candidate is frozen | Run the stage’s required L2 matrix | Repeating the full matrix after every intermediate patch |
| Only execution records or progress text change | Check completeness and document contracts | Rerunning business tests and RSS checks |
A semantic batch means one behavior change that can be verified on its own, together with its tests. It may involve several precise edits. It must pass L1 before work that depends on it proceeds; changes should not be accumulated indefinitely, but neither should batches be defined mechanically by tool-call count.
For 02-bmad-core.md, section 3, the audit proposes language along these lines:
- Physical verification means executing in an authorized environment and producing a checkable result. It does not mean rebuilding the base environment after every edit.
- Before implementation, define semantic change batches, stage boundaries, relevant verification sets and escalation conditions. Run L1 for each semantic batch and L2 for stage acceptance.
- Define separate inputs, budgets and results for L0 preparation, project compilation, assertion execution and cleanup. A test entry point must not implicitly install system packages, pull images or download dependencies.
- Classify environment-not-ready, compile failure, assertion failure, timeout, zero test hits, damaged evidence and cleanup failure separately. Any required item that does not pass blocks progression.
- A failed check blocks stage progression, not diagnosis and repair within the current scope. Do not obtain a green result by repeating the same failing command, removing assertions or weakening security constraints.
- Bind each verification result to its actual inputs and coverage. Reuse it only when the inputs and conditions of applicability are shown to be unchanged. Resuming historical work still requires reconfirming authorization and current evidence.
For 01-bmad-engineer-core.md, section 3, add or replace guidance to say that engineering mode:
Use one shared global policy, referenced by engineering mode, rather than duplicating long rules in several places:
Output reduction needs to happen before tool results enter the model’s history. Asking the model to ignore logs—or folding them in the interface—does not meet that requirement.
The verification matrix in implementation-plan.template.md should state, for each check:
The template should also separate contract inputs from execution results. Recording a result must not recursively trigger the same functional verification. Full file fingerprints remain necessary to detect external substitution; integrity protection should not be removed just to avoid reruns.
A revised audit prompt should start with the constraints, so it does not assume the cause before checking the evidence:
This is a read-only methodology audit. Do not write files, run tests or deployments, or resume a historical plan.
Check the global rules, mode rules, project plan, execution entry point and raw results separately. Distinguish observed facts from historical records and inference. Do not assume every run downloads 198 MiB, every change runs the full matrix, or all interface-visible logs entered a model request.
Quantify environment preparation, compilation, assertions, log output and governance reconciliation separately. Where the evidence does not support a measurement, state the gap.
Prefer an L1 design that preserves the same isolation strength. Do not assume permission to test on the host or prepare environments online. For every proposed rule, define its triggers, invalidation conditions, failure classifications, output budget and acceptance criteria.
Stop at the smallest workable proposal. Do not introduce a persistent container platform, generic workflow engine or unrelated business refactor.
A minimal rollout would proceed in three steps:
The accepted governance ADR already assigns responsibilities to the host platform. The proposed optimization should not reintroduce an agent-driven, open-ended permissions-proof loop, or claim that updating static rules has verified client behavior.
The order matters: remove repeated preparation and unbounded output first, then tune verification cadence. Do not weaken isolation or drop important race coverage to compensate for defects in the executor or the process contract.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
The global BMAD rules call for relevant verification at each stage; a historical project plan went further by requiring physical verification after every change.
The global BMAD rules call for relevant verification at each stage; a historical project plan went further by requiring physical verification after every change. One log shows package installation before tests, but does not establish that each run downloaded 198.1 MiB.
A proposed three tier model separates environment preparation (L0), change focused checks (L1) and stage delivery gates (L2), while preserving the authorized isolation boundary.
The global BMAD rules call for relevant verification at each stage; a historical project plan went further by requiring physical verification after every change. One log shows package installation before tests, but does not establish that each run downloaded 198.1 MiB.
Published byImages generated with GPT Image 2
Research answer
![[GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the](https://d15bya8fi85z1v.cloudfront.net/chat/QuPfET7ftYd6An8JH1fI8Jq25Ow2/thumbnails/4264E25748C8AF3AAF71/37E559A30B81BD3A5EDF-medium-640.webp)
Create a landscape editorial hero image for this Studio Global article: [GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the. Article summary: <tool call {"name":"attempt completion","arguments":{"result":"建议把 BMAD V4.2 的“物理背压”改成“隔离强度不降级、验证范围分级、环境准备独立、证据按需进入上下文”。当前证据指向的不是“测试太严格”,而是阶段粒度被细化、测试入口携带环境安装、详细输出直接回灌,以及执行记录反复对账共同放大了成本。\n\n本次仅交付只读审计和替换条款:不修改规则文件,不运行测试或部署. Topic tags: general web, workflow, code, data. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake
BMAD V4.2 validation can be made less costly without relaxing its isolation or safety requirements. The central recommendation is to separate environment preparation from test execution, match verification to the scope of a change, and keep detailed logs in an artifact channel rather than sending them wholesale into the working context.
This is a read-only audit and a set of proposed replacement rules. It does not modify rule files, run tests or deployments, or approve or archive a plan.
The global BMAD rule requires relevant verification at each stage (02-bmad-core.md, section 3). The engineering-mode rule also allows either the current host or an authorized container, and says verification should reflect the scope of the change (01-bmad-engineer-core.md).
A stronger requirement appears in a historical project plan: “physical verification after every change” (plan clause). So it would be inaccurate to say that the global BMAD rules themselves require repeated environment installation or a complete test matrix for every patch.
A more plausible explanation is that the global rules leave stage size and preparation costs undefined; a project plan then tightens the cadence to every change; and the same execution path handles both environment preparation and testing. Repeated setup, output and evidence bookkeeping can compound from there.
The aurora.log records 14 packages being installed before test events begin. At the end of the installation section, it reports 198.1 MiB for 31 packages. That total does not establish how much was downloaded—or newly unpacked—in this particular run (log).
The operation began at about 14:04:04, with test events starting at 14:04:09.330, leaving roughly 5.3 seconds for the preceding steps (timestamps). Package-level tests took about 2.72 seconds; the whole operation took about eight seconds (completion record). The available timestamps do not let us separate installation, startup, compilation or cache checks within that preceding time.
The evidence supports a narrower conclusion: the test entry point had noticeable preparation overhead. It does not quantify installation’s exact share of that overhead.
The installation section takes about 16 lines. The test output, by contrast, repeats start and pass information in both text and event form (example). The supplied log also contains an ellipsis marker representing roughly 76,000 characters of omitted output (truncation point).
Suppressing package-install messages alone would not address a flood of repeated test events. Nor does a chat export or interface display prove that all of the displayed text entered a model request or allow actual token usage to be calculated. The evidence is insufficient for either claim.
The exported history includes separate compilation and test calls, including calls associated with different container identities (fixed-state verification record). Yet the recorded test entry point still runs a Go test command that can involve build preparation (container command).
The isolation script itself was not available for this audit. It is therefore not possible to identify whether installation occurred in the script, a container entry point or another wrapper—or to conclude that every historical call installed packages again. A careful audit should distinguish three things: an installation observed in one run, multiple container calls in the history, and a claim that installation recurred on every call. Only the first two are supported here.
The proposed fix starts by separating concerns that are easy to blur together:
The historical plan requires rechecking an updated document binding and repeatedly saving execution-history summaries (binding history; recheck after recording). Those steps can protect against tampering. But if contract inputs and execution results are not distinguished, they also create a structural risk of a governance loop: record a result, change a document, rerun verification, then record again. The audit identifies that risk; it does not establish that an infinite loop occurred.
The key distinction is simple: reusing an immutable toolchain does not mean reusing a dirty test workspace. Rebuilding a temporary test workspace does not require reinstalling the toolchain.
The audit also recommends replacing claims of “absolute safety” with verifiable boundaries: block external network access and unauthorized persistent writes, while allowing bounded temporary space and designated evidence artifacts. A race-detector pass covers only the paths actually exercised; it cannot prove that a program has no data races. 9
This is a proposed contract, not a description of an existing implementation.
| Tier | Purpose | When it runs—and what should not trigger it |
|---|---|---|
| L0: Immutable execution environment | Provides approved toolchains, system libraries and dependencies; records image identity, toolchain versions and security settings. | Prepare a new version when environment inputs change. A business-source edit should not trigger system-package installation. |
| L1: Coding feedback loop | Runs relevant unit and module checks, plus targeted race or regression checks when warranted. | Run once per semantically complete change batch, using an approved isolated environment by default. Adjust test scope, not the isolation boundary. |
| L2: Stage delivery gate | Runs the required matrix against a frozen candidate, including independent, uninstrumented RSS checks and evidence reconciliation. | Run at stage completion or before an authorized delivery—not again for an ordinary status update. |
The historical authorization supplied for this audit specifies offline verification in an isolated container (authorization scope). A host-based quick test therefore should not become the default substitute without separate, explicit authorization and defined conditions.
The recommended default is a lightweight feedback loop under the same security constraints. Saving time alone is not grounds for silently switching to host execution.
The audit recommends rules based on the kind of change, not the number of edited files.
| Change or event | Required action | Should not automatically trigger |
|---|---|---|
| Implementation and assertions for one behavior are complete | One relevant L1 verification run | Rebuilding L0 or running the full L2 gate |
| Locks, concurrent access, cancellation or resource release change | Add targeted race and lifecycle checks to the current batch | Deferring race checks until final delivery |
| A cross-module interface or shared dependency changes | Expand checks to affected call paths | Testing only the edited file without justification |
| Toolchain, system dependency or isolation configuration changes | Reconfirm L0 and reassess affected evidence | Installing packages temporarily in the test entry point |
| A stage-delivery candidate is frozen | Run the stage’s required L2 matrix | Repeating the full matrix after every intermediate patch |
| Only execution records or progress text change | Check completeness and document contracts | Rerunning business tests and RSS checks |
A semantic batch means one behavior change that can be verified on its own, together with its tests. It may involve several precise edits. It must pass L1 before work that depends on it proceeds; changes should not be accumulated indefinitely, but neither should batches be defined mechanically by tool-call count.
For 02-bmad-core.md, section 3, the audit proposes language along these lines:
- Physical verification means executing in an authorized environment and producing a checkable result. It does not mean rebuilding the base environment after every edit.
- Before implementation, define semantic change batches, stage boundaries, relevant verification sets and escalation conditions. Run L1 for each semantic batch and L2 for stage acceptance.
- Define separate inputs, budgets and results for L0 preparation, project compilation, assertion execution and cleanup. A test entry point must not implicitly install system packages, pull images or download dependencies.
- Classify environment-not-ready, compile failure, assertion failure, timeout, zero test hits, damaged evidence and cleanup failure separately. Any required item that does not pass blocks progression.
- A failed check blocks stage progression, not diagnosis and repair within the current scope. Do not obtain a green result by repeating the same failing command, removing assertions or weakening security constraints.
- Bind each verification result to its actual inputs and coverage. Reuse it only when the inputs and conditions of applicability are shown to be unchanged. Resuming historical work still requires reconfirming authorization and current evidence.
For 01-bmad-engineer-core.md, section 3, add or replace guidance to say that engineering mode:
Use one shared global policy, referenced by engineering mode, rather than duplicating long rules in several places:
Output reduction needs to happen before tool results enter the model’s history. Asking the model to ignore logs—or folding them in the interface—does not meet that requirement.
The verification matrix in implementation-plan.template.md should state, for each check:
The template should also separate contract inputs from execution results. Recording a result must not recursively trigger the same functional verification. Full file fingerprints remain necessary to detect external substitution; integrity protection should not be removed just to avoid reruns.
A revised audit prompt should start with the constraints, so it does not assume the cause before checking the evidence:
This is a read-only methodology audit. Do not write files, run tests or deployments, or resume a historical plan.
Check the global rules, mode rules, project plan, execution entry point and raw results separately. Distinguish observed facts from historical records and inference. Do not assume every run downloads 198 MiB, every change runs the full matrix, or all interface-visible logs entered a model request.
Quantify environment preparation, compilation, assertions, log output and governance reconciliation separately. Where the evidence does not support a measurement, state the gap.
Prefer an L1 design that preserves the same isolation strength. Do not assume permission to test on the host or prepare environments online. For every proposed rule, define its triggers, invalidation conditions, failure classifications, output budget and acceptance criteria.
Stop at the smallest workable proposal. Do not introduce a persistent container platform, generic workflow engine or unrelated business refactor.
A minimal rollout would proceed in three steps:
The accepted governance ADR already assigns responsibilities to the host platform. The proposed optimization should not reintroduce an agent-driven, open-ended permissions-proof loop, or claim that updating static rules has verified client behavior.
The order matters: remove repeated preparation and unbounded output first, then tune verification cadence. Do not weaken isolation or drop important race coverage to compensate for defects in the executor or the process contract.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
The global BMAD rules call for relevant verification at each stage; a historical project plan went further by requiring physical verification after every change.
The global BMAD rules call for relevant verification at each stage; a historical project plan went further by requiring physical verification after every change. One log shows package installation before tests, but does not establish that each run downloaded 198.1 MiB.
A proposed three tier model separates environment preparation (L0), change focused checks (L1) and stage delivery gates (L2), while preserving the authorized isolation boundary.