Study completed, 17 September 2026. See REPORT.md for findings, REPRODUCE.md for reproduction, and results for paired outcomes and costs. All four retention policies passed 8/8 fresh requests; ready source and archived source needed zero future model calls, lessons-only needed eight, and reconstruction with caching needed one. The commissioning brief below is preserved as the pre-experiment record.
Prepared and commissioned for bounded investigation, 17 September 2026. No experiment or investigator session was started by preparation. Begin a fresh ancillary session with this file and AGENTS.md. This study owns workload discovery, methods, execution, diagnosis, evidence and publication.
When does retaining an acquired implementation improve complete future behavior or reduce actual work compared with reconstructing executable behavior from the same acquired procedural experience?
Construct-2 asks how agents accumulate useful experience across sessions and where it should persist. This study examines executable experience within that broader program. It does not require a code advantage, a neural comparison or a universal ranking of memory formats.
Public work already demonstrates tool acquisition, reusable abstractions, execution of disposable scripts, online library maintenance and repair. The remaining comparison concerns the value of the retained implementation beyond competent use of the available experience. Treat the investigation as a bounded allocation comparison, not a generic demonstration that agents can write tools.
Develop one motivated, locally reproducible recurring task and the smallest comparison that distinguishes the consequential explanations. Establish functioning acquired routines and a competent reconstruction alternative. Workload development and bounded acquisition diagnosis are part of this commission; neither a full web benchmark nor a large fixed experimental grid is required.
Separate these effects when interpreting a result:
| Comparison | What it can establish |
|---|---|
| Primitive actions versus generated code | Value of programmatic execution and data flow |
| Retained implementation versus reconstruction | Avoided work and possible differences in preserved procedural information |
| Implementation versus lessons/examples built from shared acquisition | Value of the chosen retained representation, including its construction and omissions |
If both alternatives retain executable source, allow loading or copying it. Forcing a model to rewrite usable source manufactures reconstruction work. Code in trajectories or workspace logs still counts as available implementation. If it lives in a recoverable archive, assess ready access versus actual recovery; if only lessons and examples survive, disclose what their construction preserves or loses. A fair comparison need not impose equal byte counts or forbid useful existing tools.
Both alternatives should receive current observations, execution, useful acquired tests or equivalent checking, and opportunities for purposeful repair. Develop concise contextual lessons and examples, selective loading and actual caching where supported. A library may use parameters, conditional observations and local repair. Distinguish supplied primitives, schemas, authority and evaluation answers from what the agent acquires; benchmark gold is not free runtime feedback.
First assess stable reuse on fresh requests. If a consequential revision would resolve the question within scope, distinguish changed input values, changed interfaces and changed requirements. A valid routine may already accept new values. Superseded obligations and tests must not remain authoritative merely because they are historical. Repair is a conditional extension, not a required second phase regardless of the first result.
These conjectures precede local findings. They guide interpretation without prescribing every contrast:
- EX1: With repeated stable operations, retaining functioning code should reduce reconstruction work; the complete-quality advantage may be small when regeneration is already reliable.
- EX2: Local interface changes should often favor local repair over complete regeneration when the retained abstraction still matches the task. Broadly changed assumptions may reverse that advantage. Mere exposure to a new case is not evidence of either kind of change.
- EX3: An apparent code-memory advantage will shrink when extra sampling, execution and observation access are accounted for. Residual savings or gains would be stronger evidence for persistence itself.
These are method donors and overlap evidence, not mandatory implementations. The root inspected the listed paper passages and code statically; no public experimental outcome was reproduced. Code snapshots are not verified original experimental revisions.
| Source | Use and relevant limit |
|---|---|
| SkillCraft, 2603.00718v2, §§2–4, 5.1, Appendices B/D.2; release 0a9ba880 | Programmatic tool chains and an existing generated-script route. The release saves script logs; library execution adds a quality warning absent from Direct Exec. Paper descriptions of the direct arm conflict, and efficiency averages condition on common successes. |
| SkillWeaver, 2504.07079v1, §§2–4; release f2a63d65 | Acquired browser routines and a reference-only route. Default reference content, primitive permissions and missing bundled verification metadata prevent treating its switch as a ready matched comparison. |
| TroVE re-evaluation, 2507.22069v2, §§2–4 | Generation budget and selection can explain an apparent library gain. Equal call counts alone do not establish equal work. |
| CodeMem, 2512.15813v1, §§4–7 | Explicit reconstruction-avoidance architecture with reported executions; not an isolated matched-experience persistence result. |
| PANDO, 2605.24785v2, §§3–6 | Online routines, checking, demotion and induction accounting are established precedents; seed capability and bundled changes matter. |
| Skill Blocks, 2608.14943v1, §§3–7, Appendix C | Competent contextual delivery need not resend an uncached full lesson. Cost sensitivities do not establish measured billing or latency. |
The root's comparison and reading ledger provide optional background. The brief above is sufficient to begin independently. Further discovery should follow the chosen workload and uncertainty, not repeat a generic donor inventory.
Measure complete outcomes and paired failures across the same request sequence. Count failed acquisition, generation, invocation, checking, observation and repair where incurred. Report tokens, cache use, primitive actions and elapsed time in their native units where available. Success-conditioned costs alone cannot establish deployment savings. Separate full experimental search from the cost of one selected deployment; an inherited library with unknown acquisition costs supports only a marginal-use claim.
Preserve unsuccessful attempts and development decisions. Use fresh evaluation material for claims developed during diagnosis. A functioning learner, a useful explicit alternative, a tie, or an explanation of failed acquisition can each advance the question. Stop on explanatory progress, a demonstrated limitation or a concrete resource constraint, without tuning indefinitely for a winning representation. Publish findings, limits, costs and reproduction evidence here; the root will assess their implications separately.