Can person-grounded skills remain auditable without hidden prompt state?
Explores whether treating extracted expertise as versioned files—rather than persona prompts—enables meaningful accountability over person-grounded knowledge. Matters because audit trails determine whether captured skills can be corrected, rolled back, or safely withheld.
The interesting move in COLLEAGUE.SKILL is not that it distills a person's review judgment, decision heuristics, and interaction style from heterogeneous traces — plenty of memory and persona systems already grab fragments of that. The move is that it refuses to treat the result as a persona prompt and instead treats it as a versioned file subject to a full lifecycle: creation, inspection, invocation, correction, rollback, deletion, install, and optional distribution. Two coordinated tracks — a capability track (practices, mental models, heuristics) and a bounded behavior track (communication style, interaction rules, correction history) — keep the "what they know" and "how they act" separable, so each can be audited independently.
Why this matters: a single prompt can mimic surface behavior, but it makes the extracted knowledge unaccountable — you cannot point to where a claim came from, repair it, or refuse to ship it. The generation effect is the same critique I make of vault notes: passive transfer is not understanding. Here it becomes a governance argument. Person-grounded knowledge becomes auditable only when it lives in work.md and persona.md files that can be diffed, not hidden in prompt state.
This is the human-expertise end of the harness/skill lifecycle. Where Can skill documents be optimized like neural network weights? treats a skill file as trainable external state, COLLEAGUE.SKILL treats it as auditable external state — the discipline is provenance and correctability rather than optimization. The strongest counterargument is that file-level governance does not constrain how the loaded skill actually behaves at inference; a clean manifest can still front a skill that drifts from the person it claims to ground. Inspectability of the artifact is necessary but not sufficient for behavioral fidelity, which the paper itself flags as an open frontier.
Inquiring lines that read this note 31
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can code harness improvements rival direct model scaling for capability? What external process records should verify agent behavior and benchmark claims?- Does inspectable skill artifacts guarantee the behavior matches the person it claims to ground?
- How do signed logs compare to externally anchored records for audit?
- What architectural controls secure capture authenticity beyond signing?
- Can pinned artifacts prevent audit agents from making inconsistent judgments?
- Where else in the vault are recovery and rollback mechanisms already specified?
- Can an auditor verify environment state without trusting the executor's self-report?
- Why do credentials need evidence standards beyond permission categories?
- How can post-training research become reproducible without releasing full interfaces?
- Should platforms downweight or relabel credentials after retiring the format that issued them?
- How do organizations safely retain and control access to committed content?
- What one-time human costs does building a hidden partition require?
- Does held-out validation prevent skill document edits from drifting or accumulating harm?
- How should skills be trusted and installed on sharing platforms?
- What makes a distilled skill verifiable and ready for agent execution?
- What process evidence should assessment systems require alongside finished work?
- What makes a credential robust when the tools for earning it change?
- What counts as evidence that a credential still certifies after GenAI?
- What metadata properties make code-derived skills auditable and comparable to their original source?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can skill documents be optimized like neural network weights?
Explores whether natural-language skill artifacts—packaging procedures, heuristics, and policies—can be systematically improved through iterative editing and validation, similar to how gradient descent refines model parameters.
convergent-with: both treat a skill as an editable external-state file with explicit gating, here governance/audit rather than optimization
-
Can codified expertise let non-experts match specialist output?
When domain knowledge is captured as explicit rules and principles in an AI agent's scaffolding, can non-experts produce work at expert quality levels without consuming scarce specialist time? This explores whether structured knowledge codification dissolves organizational bottlenecks.
exemplifies the same expertise-externalization claim, here grounded in a specific person rather than abstract design rules
-
Can agents learn new skills without forgetting old ones?
Explores whether externalized skill libraries—storing learned behaviors as retrievable code rather than parameter updates—can solve the catastrophic forgetting problem that plagues continual learning systems.
extends the skill-lifecycle framing toward person-grounded artifacts with rollback and withholding
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
- Automation, AI, and the Intergenerational Transmission of Knowledge
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight
- Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
- Putting AI on the Org Chart: Evidence on Delegation and Accountability
Original note title
person-grounded skills demand the same file-level lifecycle as any other artifact — inspect, correct, rollback, and withhold, not just generate