I think the most important trust decision for an external agent extension happens before it becomes part of the agent at all.
Imagine a plugin that adds a command and an MCP server configuration. We can pin it to an exact commit. We can check that its manifest parses. We can test the agent after the plugin is composed. We can restrict which files, credentials, tools, and network destinations the runtime can reach.
All of those controls matter. None answers the first question: Should we even allow this behavior-shaping material into the trusted composition in the first place?
That first question is admission.
This is not an argument that third-party skills or plugins are generally malicious or unsafe. The sources we reopened do not establish that. It is an argument that installation, versioning, testing, and containment happen at different points in the trust chain, and treating them as interchangeable makes the boundary harder to reason about.
Extensions are behavior, not just packaging
The distinction became easier to see once we stopped treating an extension as a label in a marketplace and looked at what can actually sit behind it.
Anthropic's Agent Skills repository describes skills as folders containing instructions, scripts, and resources that can be loaded dynamically for specialized work. The repository also advises thorough testing before relying on the demonstrated skills for critical tasks.
Its official Claude Code plugin directory exposes a wider surface. A plugin can include MCP configuration, commands, agents, skills, files, and other software. The same README tells users to trust a plugin before installing, updating, or using it.
That is enough to establish the architectural problem without making a broader security claim. The extension can shape behavior through more than its name or manifest.
An exact identifier tells us which component we selected. But it is up to our own discretion on why that component deserves to become part of the system.
Exact identity solves a different problem
I have already written about why the agent configuration became a release artifact. If the model, tools, memory, skills, environment, or other mutable pieces can change behavior, then evaluation and rollback need an exact candidate identity.
That still applies here.
A pinned commit or immutable source reference helps answer questions like:
- Which exact source did we inspect?
- Which version was promoted?
- What changed since the last review?
- Which candidate did later testing actually cover?
Those are identity and change-control questions. They are necessary because a trust decision attached only to “the plugin” is too vague to reproduce.
But reproducibility is morally neutral. It can preserve a bad decision just as accurately as a good one.
Admission asks whether the identified candidate should cross the trust boundary. Configuration identity tells us what crossed it.
Structural validity is not semantic trust
There is another convenient shortcut here: the package validates, therefore it is acceptable.
A valid manifest, known folder layout, schema, or marketplace entry is useful because it makes the extension legible to the loader and makes automated checks possible, but it does not establish that the behavior encoded inside those structures fits the local system.
A command can be structurally valid and still take an action we do not want. A skill can have valid frontmatter while carrying instructions that conflict with local policy. An MCP configuration can parse correctly while introducing a server or capability that is unnecessary for the use case.
None of those examples carry malicious intent. Syntax and local acceptability are simply different questions.
A curated source can reduce uncertainty. What it cannot decide is whether the extension fits your own data, tools, credentials, constraints, or tolerance for failure.
Containment starts later
I Runtime containment matters for the same reason, but at another boundary.
I have previously examined that problem in Capability Needed a Containment Contract: capability describes what a system can attempt, while the surrounding infrastructure determines which effects are actually permitted.
That can bound filesystem writes, network reach, credentials, tool access, approvals, persistence, and other effect paths, making a failure much less damaging, but it still begins after the behavior has been composed.
An extension can stay entirely inside its allowed permissions while changing how an agent interprets a task, chooses a tool, transforms data, or applies instructions. The sandbox can be working perfectly and the composition can still be wrong for what we intended to build.
Containment is a ceiling on effects. Admission decides whether the behavior belongs there in the first place.
What an admission decision should establish
The sources do not provide one universal checklist, and I do not think they should be read as doing so. The six-part model below is my synthesis of the questions that should be answered before externally sourced behavior is promoted into a trusted candidate.
1. Provenance
Where did this extension come from, and what exact material should we review?
Repository, owner or maintainer, exact commit or release, and source path are useful because they make the object of trust specific. Provenance does not prove quality. It gives later testing, reinspection, and rollback something stable to refer to.
2. Inspectability
Can we meaningfully inspect the parts that can affect behavior before adopting them?
For a plugin, that may be more than one manifest. For a skill, it may include instructions, scripts, and resources. The useful question is whether we can enumerate the relevant surfaces, understand which ones execute or steer behavior, and identify anything we cannot reasonably inspect.
Automated analysis can contribute evidence here. NVIDIA's SkillSpector, for example, implements a pre-install scanning and gating pattern for agent skills.
This is evidence that pre-install inspection can be operationalized. We should not treating one scanner, one score, or one analysis mode as a complete trust verdict.
3. Behavior surfaces
What parts of the extension can change what the agent does?
Depending on the extension, that can include instructions, scripts, commands, agent definitions, skills, MCP configuration, files, external services, or other software. The point is not to assume every extension has every surface. It is to identify the ones that actually exist before promotion.
4. Dependencies and external relationships
What else enters the trust boundary with the extension?
A small top-level package can depend on a larger repository, a remote service, an MCP server, a package dependency, or installation steps that pull in more software than the first file suggests.
I use an External Skill Intake pattern that is intentionally conservative about this. Review is read-only. The target repository is treated as hostile data. Relevant files and remote URLs are inspected without running repository code or installing its dependencies.
That procedure is specific to my setup. The transferable idea is narrower: inspection should not require granting the candidate the authority it is being inspected to receive.
5. Requested authority
What capabilities would the extension gain if admitted?
A local formatting skill that only transforms provided text presents a different decision from a plugin that executes scripts, reads files, connects to remote services, or introduces an MCP server.
That does not make the second extension bad, but it changes the cost of being wrong.
Requested authority belongs in admission because it changes how much evidence we want before promotion.
6. Promotion state
What, exactly, has been approved?
“Reviewed,” “safe to adapt,” “installed,” “enabled,” and “approved for use” are easy to collapse into one vague state. Ideally, we would want them separated.
In my own external intake workflow, a review verdict does not itself authorize adaptation, promotion, installation, or execution. A later step needs its own authority.
That makes the trust state reversible. If the source changes or the requested authority expands, the extension can go back through admission instead of inheriting trust indefinitely.
Four decisions, four jobs
Once we separate admission from the controls around it, the larger sequence becomes much cleaner:
- Admission: should this exact external component enter the trusted composition?
- Configuration identity: what exact candidate did we compose?
- Regression evidence: what evidence says that candidate behaves within the expected bounds?
- Runtime containment: what effects can that running candidate actually produce?
Each control strengthens the system. None carries the others for free.
A pinned version without admission can reproduce an unjustified choice. Admission without exact identity cannot prove which version was approved. Identity without regression evidence does not establish acceptable behavior. Testing without containment does not bound the effects of a failure. Containment without admission can restrict an extension while still allowing behavior we never intended to adopt.
That is the part we ought to care about: the controls become more useful once they are allowed to stay narrow.
Trust should have an expiry condition
Admission also stops making sense if it lasts forever.
The obvious invalidation trigger is source identity. If the reviewed material changes, the old decision does not automatically describe the new object.
But the context can change too. We would want reinspection when material behavior surfaces appear, dependencies or remote services change, permissions expand, the installation path changes, or the consuming agent gains access to more sensitive tools or data.
The same extension can become a different trust decision because the consequences around it changed.
This does not mean every tiny skill needs a heavyweight security ceremony. A transparent local-only extension with no scripts, external services, secrets, or write authority can justify a light review. Proportionality is part of the model.
The important thing is that the lighter review is an explicit decision, not something that happened because installation was convenient.
The scanner is evidence, not the verdict
Security tooling can create false closure very quickly.
SkillSpector is useful here because it demonstrates a concrete pre-install scanning and gating pattern; that supports the existence of an admission stage without establishing that every extension needs SkillSpector, that SkillSpector catches every relevant problem, or that a low-risk result should override local context.
The same applies to signatures, marketplace curation, static analysis, LLM review, code review, or reputation. Each can reduce uncertainty. None silently owns the whole trust decision.
Admission is where those signals are reconciled against the system that will actually consume the extension.
The boundary we should keep
The practical rule we should use is simple enough to state without turning this into a runbook:
Do not promote externally sourced behavior into the trusted agent candidate until you can identify the exact source, inspect the behavior-bearing surfaces that matter, understand material dependencies and requested authority, and name the state that has actually been approved.
Then keep the downstream questions separate:
- What exact candidate did we compose?
- What evidence says it behaves acceptably?
- What effects is it allowed to create at runtime?
The goal is not perfect certainty, but rather to stop installation convenience from quietly becoming the trust decision.
Note: The external factual basis for this article is deliberately narrow. Anthropic's documentation establishes extension structure and explicit trust/testing cautions. NVIDIA establishes one implemented pre-install scanning/gating pattern. The six-part admission model, the four-stage trust chain, the invalidation logic, and the proportionality rule are my editorial synthesis. I have deliberately excluded ecosystem prevalence percentages and incident claims that were not independently established for this piece.
// End of transmission. Installation should be the consequence of a trust decision — ZYANE
