← LOGS

The Host Surface Is Part of the Agent Contract

Embedding an agent in Slack, mobile, or another work surface changes more than the interface: it changes context, identity, triggers, visibility, authority, and the record left behind.

Close-up of a thin molded sheet conforming tightly over a raised underlying form, with pronounced ridges, depressions, and pixelated scanline texture.
The surface takes its shape from what lies beneath.

IN BRIEF

Embedding an agent in Slack, mobile, a repository client, or another work surface changes more than where it appears. The host can supply context, identity, trigger conditions, visible state, interruption points, and expectations about what happens next. Those properties become part of the agent's execution contract. Product teams should therefore specify what context is ambient, who the principal is, what starts or stops work, what state people can inspect, where consequential authority is enforced, and what durable record remains. The key boundary is that proximity is not permission: a convenient button, mention, or repository context can express intent without owning unrestricted credentials or final authority.

Putting an agent inside Slack or a mobile app can look like a distribution decision. Same model, same tools, closer interface.

I think that framing is too small.

The host already has its own rules before the agent arrives. Slack knows who is speaking and which thread they are in. A coding client can know which repository was selected. Mobile has a notification model. A pull request has a review state. Once the agent moves into one of those surfaces, those existing semantics start shaping what it can reasonably know, what starts work, what a person can see, and where permission or approval belongs.

So the useful question is not only where can the agent appear? It is: what does the host surface silently contribute to the execution contract?

I am using "host-surface contract" as an analytical frame here, not as vendor terminology or an established standard. The examples below are bounded implementations, not evidence that the industry has converged on one architecture.


Context arrives before the prompt

A standalone chat makes its conversational boundary obvious. An embedded agent often begins with context the host already considers ordinary.

In Cursor's iOS client, the user chooses a repository before launching a cloud agent. The invocation is already attached to a codebase. The same mobile surface can pass visual context into the workflow and keep the control experience coherent while work moves between local and cloud execution.

Moon Bot, a Slack-native coding-agent implementation described by Hugging Face contributors, makes the same issue even clearer. Each Slack thread maps to a persistent agent session. Thread history and tool activity can be restored when the conversation resumes. The thread is not just where responses are displayed; it helps define the continuation boundary of the work.

That creates a distinction I think product teams need to make explicitly:

Ambient host context is not the same thing as portable agent context.

Agent Skills is a useful example of the second category. A skill packages specialized knowledge and workflows into a version-controlled folder, with instructions and resources loaded progressively when needed. That context exists because somebody deliberately packaged it for reuse.

A Slack thread, selected repository, ticket, document, image, or workspace state is different. It exists because of where the agent was invoked.

Both can be useful, but they should not be collapsed into one vague bucket called "context." If a system cannot explain what came from the host, what came from a portable skill, and what the user explicitly supplied, reconstructing why the agent acted becomes much harder than it needs to be.


Identity helps, but it does not grant authority

Embedded surfaces often know more about the person invoking an agent than a generic chat box would.

Moon Bot maps Slack users through Okta groups to access tiers. The interesting part is that this identity is not treated as decorative metadata. It affects which execution environment and credentials are available. Lower access tiers do not simply receive a prompt telling the model to behave; the implementation constrains what credentials exist in the runner.

That separation matters.

The host can help answer who is asking?, but it should not automatically answer what may this action cause?

A user being able to mention an agent in a channel does not imply that every connected tool should be available to that user. A repository being visible does not imply that an agent should be able to merge into its default branch. A mobile button being convenient does not settle whether the principal behind that button has the authority to perform the action.

Microsoft's Agent Governance Toolkit makes this distinction explicit in its own architecture. It separates policy enforcement, agent identity, and audit records from prompt-level instructions, and actions can be allowed, denied, or routed through approval. The repository labels the toolkit Public Preview, so this is not a claim about a finished universal pattern. It is still a useful illustration of the boundary: interface identity and execution authority are related inputs, not synonyms.

For an embedded agent, the host may identify the human, team, channel, repository, document, or task. The execution layer still has to decide what that identity is allowed to cause.


Trigger conditions become product semantics

In a standalone chat, the default trigger is simple: the user sends a message.

Embedded surfaces immediately complicate that.

Moon Bot responds to mentions in Slack channels, while direct messages do not require a mention. Small detail, large consequence. A channel message, a direct message, a mention, a repository event, a button, a schedule, and a background condition are not interchangeable ways of saying "the user asked." They carry different expectations about attention and consent.

I am deliberately not turning this into the separate question of when agents should initiate work proactively. The narrower point is enough: trigger conditions are part of the host-surface contract.

A product specification for an embedded agent should be able to answer some boring questions very precisely:

  • What exact event starts work?
  • Does that event identify a human principal?
  • Can work continue after the initiating event is gone?
  • What pauses or cancels execution?
  • When is a new explicit instruction required?

If those answers are left implicit, "embedded" can quietly become "always listening" or "implicitly authorized." Neither follows from the interface by itself.

Every host arrives with its own idea of what counts as being asked.


Visible state is part of control

Longer-running agents create an attention problem as much as a compute problem.

Cursor's mobile client exposes state outside the main conversation through Live Activities and push notifications, including finished, needs-input, and ready-for-review states. Its agent workflow can also surface demos, screenshots, logs, and diffs for validation. A user does not have to watch every intermediate step to know when the work has crossed a meaningful boundary.

Moon Bot takes a different route: responses can link back to persisted session artifacts and traces. The Slack message is therefore not the complete state; it is an entry point into a more durable record of what the agent did.

Those examples separate two jobs that are easy to blur together:

  1. Attention routing: tell the human when something changed and when a decision is required.
  2. Execution legibility: retain enough evidence that the human can inspect what happened.

More visibility is not automatically better. Exposing every intermediate event can make a system observable and still make the interface unbearable. Showing only done can make the interface calm and still leave nobody able to inspect consequential work.

The product decision is which state is visible by default, which evidence is available on demand, and what survives after the interaction ends.


Proximity is not permission

This is the part I would be most nervous about getting wrong in a product specification.

When an action appears next to the work, the interface makes it feel local. A Slack control beside a thread. A merge action beside a diff. A mobile button beside an agent result. That proximity is useful, but it can also disguise where authority actually lives.

Moon Bot can create pull requests without handing broad GitHub write credentials to the sandboxed agent. Its write path uses dedicated tooling and narrowly scoped tokens outside the sandbox. Cursor can present review and merge actions in the mobile experience while cloud agents run in isolated development environments. Microsoft's governance tooling puts policy decisions around tool execution rather than assuming the model or front end owns the final decision.

These implementations are not the same architecture, and I would not generalize them into one. The transferable point is smaller:

Surface proximity should not be mistaken for authority ownership.

The host can be where a person expresses intent without being where unrestricted credentials live. It can make an action easy to request while the execution layer keeps that action narrow, auditable, revocable, or approval-gated.

That split — intent in one place, enforcement in another — is the pattern When the Product Boundary Moves Into the Control Layer tests directly: a surrounding control belongs to the product when removing it would change what the user can rely on.

That distinction matters especially in surfaces where consequential actions look socially ordinary. Mentioning somebody in Slack is cheap. Asking an agent in the same thread to query sensitive data or open a pull request can look equally cheap even when the permission model absolutely should not be.


Interruption needs a home too

"Start" is only half the control design for asynchronous work.

A person also needs to know where to redirect, pause, veto, or withhold acceptance. Cursor exposes follow-up instructions and a review step before merge. Moon Bot separates read access from dedicated write tools and falls back to a basic access tier when identity configuration cannot be resolved. Microsoft's policy model can require approval for particular actions.

Different mechanisms, same question:

What is the last human-controllable boundary before the action becomes consequential?

Sometimes the gate belongs before a tool runs. Sometimes investigation can proceed freely and the gate belongs before a write. A coding agent might be allowed to produce a draft pull request while merge stays human-controlled. A support or analytics agent might be allowed to read while sending, deleting, or changing data requires another approval.

The exact answer will vary, but what should not vary is whether the answer exists.

Otherwise a user can see the agent and see the action without knowing whether they are looking at a proposal, an in-progress operation, or an effect that has already been authorized.


The record after the interaction is part of the product

It is easy to judge an embedded agent by the immediate exchange: useful Slack reply, completed mobile task, acceptable pull request.

But consequential work has a second audience: the person who needs to understand what happened later.

Moon Bot persists session histories and traces. Cursor can expose logs, screenshots, demos, diffs, and pull requests. Microsoft's toolkit models decision and audit records around actions. Again, these are different systems. What they have in common is an acknowledgement that the visible response is not enough once an agent can operate across tools and time.

A useful test is simple: after the agent finishes, can another person reconstruct the important parts without relying on the agent's own summary?

That does not require a chain-of-thought transcript or maximal telemetry. A bounded record of inputs, actions, tool results, diffs, approvals, and outcomes can be enough. The goal is legibility of consequential work.

A good answer ends the exchange. A good record survives it.


The contract does not travel

Everything above assumes one host. Work does not always stay in one.

A task can start as a mention in a channel, continue in a mobile client, and end at a pull request. Each surface answers the questions above differently. The thread that set the continuation boundary is not holding the review state. The identity a workspace directory resolved is not the identity the repository sees. The notification that said ready for review is not the record somebody reads later.

I do not have a general answer, and none of the implementations above claims one. A handoff between surfaces is a change of contract, not only a change of screen.


Six questions I would put in the spec

The frame reduces to six questions. This is my synthesis from the implementations above, not a standard proposed by any of them.

  1. Context — what does the host make ambient? Which thread, repository, document, ticket, image, prior interaction, or workspace state comes along because of where the agent was invoked? What additional context is loaded explicitly?

  2. Identity — who is the principal? Which human, team, service, or role does the host identify, and how is that identity translated before the agent reaches tools or data?

  3. Trigger — what starts and continues work? Is it a message, mention, button, repository event, schedule, or background condition? What stops the work or requires a new instruction?

  4. Visible state — what can the human see without chasing the agent? Which transitions produce notifications? What evidence is available for inspection? What remains hidden unless requested?

  5. Authority — where are consequential actions allowed, denied, approved, or vetoed? Does the agent propose, draft, execute, or commit the action? Where are permissions actually enforced? Where can a human intervene?

  6. Record — what remains afterwards? Are there durable traces, diffs, decision records, artifacts, or links that let somebody inspect the action later?

A product can answer each question differently. The failure is answering one of them accidentally because the host already happened to have an event hook, identity field, or button.


The interface now participates in execution

I do not think the interesting shift is that agents can appear in more places. It is that the surrounding product starts participating in the conditions under which the agent operates.

In an embedded system, context may arrive before the prompt. Identity may come from an enterprise directory. Triggers may come from host events. State may be split across notifications, traces, and artifacts. Authority may be expressed in the interface and enforced somewhere else.

The model and runtime can remain technically similar while the actual execution contract changes around them.

So "agent + integration" is not enough of a specification. At least for this class of product, the host needs to be specified as part of the system: what it contributes, what it only displays, what it may trigger, what it can authorize, and what evidence it leaves behind.

The interface is not just where the agent speaks anymore; it is one of the places where the terms of action are defined.

Nobody specified the thread, the mention, or the notification. They shaped the contract anyway.


// End of transmission. Specify the surface contract. — ZYANE