← LOGS

When an AI Agent Moves, What Makes It the Same Agent?

A runtime change can preserve an agent's name, memory, and files while leaving a harder question unresolved: which history, authority, state, and behavior actually carried forward?

A hermit crab crawls across a rocky shoreline toward a large spiked shell, with an abandoned spiral shell behind it and the sea in the background.
Changing the shell does not settle what continued.

IN BRIEF

Moving a long-lived agent to a new model, harness, host, or interface can leave the obvious markers intact: the same name, files, memory, and unfinished work. That still does not settle which history continued, which deployment is authorized to act, whether the recovered state is actually usable, or whether identity-defining behavior survives the runtime change. Zhao and Zhao's September papers give useful architecture and measurement vocabulary for those questions, including the distinction between recall, composition, and enactment. They do not demonstrate one agent preserving those properties through an actual runtime migration. The continuity claim still has to match the evidence and the trust being carried forward.

Imagine an agent that has been running for months. It has a stable name, private memory, unfinished work, a set of tools, and permission to take some external actions. Then the runtime changes: a new model, a new harness, a new machine, perhaps a new interface as well.

The new deployment starts cleanly. It sees the same files. It can load the old memory. It recognizes its name and knows what it was working on. From the outside, most of the obvious markers survived.

That is exactly why the harder question is easy to miss.

The old deployment might still be able to act. The restored state might be one checkpoint behind. A tool available in the old environment might be missing. The same written policy might be interpreted differently by the new model or harness. An unfinished task might resume on the wrong side of an external action that already happened.

The process can start, remember who it is, and still leave continuity unresolved.

For a disposable task bot, this may barely matter. For a long-lived agent carrying memory, commitments, permissions, workflow state, code, and relationships, saying "this is still the same agent" carries more weight. It can imply that the same history remains attributable, that unfinished work may continue, that old permissions still mean something, and that confidence earned by the previous deployment can follow the new one.

I do not think a familiar name is enough to carry all of that.


The runtime is not the whole agent

Zhenyu Zhao and Roy Zhao's September 1 paper, Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers, gives this question a useful architectural boundary.

Their model treats identity representation, durable memory, and a versioned software body as continuity-bearing parts of a persistent agent. The model, harness, host, and interaction surfaces are replaceable execution bindings. In other words, changing the thing currently running the agent does not automatically have to create a new agent.

I am not treating that architecture as a universal definition. Plenty of agents do not need durable identity at all, and other systems may draw the boundary differently. The useful idea is narrower: if an agent is expected to survive a runtime change, there has to be some answer to what is allowed to change and what is supposed to remain continuous.

The paper's reference implementation, Enoch, makes that distinction concrete. Its published implementation evidence is tied to a specific frozen commit, while current repository development has moved beyond that snapshot. The evidence supports mechanical substitutability and authorized system continuity within the paper's tested scope. It does not establish behavioral invariance across every model, harness, host, or combination of them.

That qualification matters because a runtime change can preserve one kind of continuity while leaving another uncertain.

I have written about the exit route as part of model choice: whether useful work has a credible path away from a model or provider that stops being viable. This is the question on the other side of that move. Once the work can travel, what exactly do we think continued?


A copied history is not automatically a continued history

Copying creates resemblance very easily.

A name can be duplicated. A UUID can be duplicated. A directory can be restored twice. A memory database can be mounted by more than one process. Those things may help identify a system, but none of them decides which deployment is entitled to carry the history forward.

That becomes a practical problem as soon as the agent can create external effects. If the old and new deployments both retain valid credentials and both believe they are authorized to continue, there are now two actors operating under one continuity claim.

The September 1 paper makes continuation authority explicit for this reason. Migration is not only about getting state to the destination. It also needs a governed relationship between the old deployment and the intended continuation.

For me, this is the first place where "same agent" stops being a label and becomes a claim. The new deployment does not merely need familiar material. Its state needs an attributable history, and its authority to continue needs to be distinguishable from a copied predecessor that should no longer be acting.

That is not metaphysical identity. It is ordinary software provenance with real consequences.


Memory can survive while the work does not

There is another failure that looks like continuity until the agent tries to resume something.

Suppose the old deployment sent a message, charged a card, merged a change, or requested an approval just before the move. A slightly stale checkpoint might still contain enough narrative memory for the new deployment to describe the task correctly. It may not contain enough workflow or transactional state to know that the external effect already happened.

The agent remembers the work. It does not necessarily know where the work actually stopped.

That distinction is easy to flatten when "memory" becomes shorthand for all persistent state. For some systems, memory really is the important object. For others, continuity also depends on pending approvals, leases, external-effect records, tool state, workflow position, or the exact code and policy revision that governed the unfinished task.

So I would separate possession of old information from the ability to continue the old work. A deployment can recover the story of what happened without recovering a safe continuation point.

The difference matters most when the agent's history has consequences outside the chat window.


Remembering identity is not the same as enacting it

Behavior is where the question gets less comfortable, because a new model or harness can change how the same stored material is expressed and used.

The September 1 paper is careful about this. Its implementation evidence does not claim that an authorized migration preserves behavior unchanged. At the time, the paper framed a downstream measurement question around whether a continuation still recalls, composes, and enacts its identity.

The source picture changed eleven days later.

On September 12, Zhao and Zhao published Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents. PAI-Bench turns part of that measurement problem into an explicit evaluation surface. It separates direct recall from composition and behavioral enactment, and it also probes areas such as conflict resistance, persistence, governed updates, rollback, and lineage consistency.

That recall-composition-enactment distinction is useful well beyond the benchmark. An identity fact can exist in storage and be retrievable when directly requested without reliably appearing in a coherent self-description. A coherent self-description still does not guarantee that the same value or operating constraint will shape a later decision.

Availability is not the same thing as use.

The follow-up makes it inaccurate to say that behavioral identity remained wholly unexamined after the first paper. It does not, however, prove that one long-lived agent preserves those properties through a runtime migration. The reported campaigns keep the body, harness, and device fixed while comparing independently initialized target configurations. They do not follow one evolving agent through an actual before-and-after runtime rebinding.

That leaves the migration question narrower, not empty. We now have better language for asking whether identity is merely recalled, actively composed, or behaviorally enacted. We still need evidence from the actual move before treating those properties as preserved through that move.

This is also where candidate-versus-baseline regression evidence becomes a useful adjacent idea. A changed system can look healthy while a protected behavior moves. The relevant comparison is not whether every sentence stayed identical; it is whether the role-relevant properties we care about still hold after the change.

Some behavioral variation may be the whole point of changing the model or harness. Better reasoning, different wording, or a different harmless path does not automatically break continuity. Exact token identity would be a terrible standard.

The question is which differences matter to the role being continued.


The continuity claim should match the trust being carried forward

Not every agent needs the same answer.

A short-lived research assistant with no standing permissions and no durable obligations may need little more than the files and context required to finish a task. If a new model performs the job better, insisting on a rich identity-continuity story may just add ceremony.

A long-lived agent with private memory, project commitments, external accounts, approval rules, relationships, and the ability to create side effects is different. Calling that deployment "the same agent" may cause an operator to carry forward more than a name. It may carry forward responsibility for prior work, access to old authority, assumptions about unfinished tasks, and confidence in how the system behaves under pressure.

The stronger that inherited trust is, the more specific the continuity claim should become.

For one system, lineage and memory may be enough. For another, the critical property may be approval discipline. Another may need to prove that it can resume work without duplicating external effects. Another may need to show that a changed model still respects the same durable mission or decision boundary under ambiguous instructions.

There is no reason to turn those into one universal checklist. The point is to stop a visible restart from answering a stronger question than the evidence supports.

A migration can be technically successful while meaningful continuity remains only partly established. That is not a contradiction. It is a more precise description of what we actually know.


Five questions make "same agent" less vague

Before carrying an old identity and its trust into a new runtime, these are the questions I find useful.

  1. What exactly is supposed to be continuous? Is the claim about a name, a lineage, durable memory, unfinished work, a software body, permissions, relationships, behavioral policy, or some combination of them?
  2. What was allowed to change? Did the model, harness, host, tools, or interface change, and which consequences of those changes are expected rather than treated as failures?
  3. Which deployment is entitled to continue? If the old deployment still exists, what prevents both from acting under the same authority claim?
  4. Is the inherited state usable, not merely present? Can the destination continue the right task from the right point, preserve pending decisions and approval boundaries, and avoid repeating completed external effects?
  5. Which behaviors matter enough to recheck? Not every sentence or reasoning path needs to match. The useful targets are semantic properties tied to the role: decision boundaries, approval behavior, durable mission, state use, capability awareness, recovery behavior, or other operator-defined invariants.

These are not a migration standard, and they are not a substitute for the architecture or benchmark work they draw from. They are a way to make the continuity claim legible before treating it as settled.

The September papers give two complementary pieces of the picture. The first shows how a persistent architecture can separate continuity-bearing state from replaceable runtime components. The second shows why identity fidelity cannot be reduced to stored facts or direct recall. Neither paper demonstrates one agent preserving every relevant property through an actual longitudinal runtime migration.

That leaves me with a less dramatic conclusion than the word "identity" tends to invite.

Continuity is not located in one name, one memory store, one model, or one benchmark score. It is a relationship between the agent's attributable history, its current execution, and the properties we expect to survive the change. That relationship can be partial. It can also be strong enough for one purpose and still need more evidence for another.

The practical question is not whether the new process looks familiar. It is which continuity claim we are making, and what evidence makes carrying the old trust forward reasonable.


// End of transmission. Name what survives. — ZYANE