A prompt can be completely clear and still age badly.
At the beginning, it may be exactly enough: do this, under these constraints, stop here. The agent proposes a plan, the first changes match the request, and nothing feels missing.
Then the task develops a second life. Another session picks it up. A constraint changes. The plan gets revised. A different executor handles the next phase. Someone reviewing the result eventually has to answer a harder question than whether the implementation looks plausible: is this still the result we meant to produce?
That is the point where I would stop trusting the conversation alone.
The prompt did not necessarily fail. The work simply acquired a longer lifetime than the prompt was designed to carry.
The useful distinction is not prompt versus document. It is transient instruction versus durable authority.
The threshold is lifetime, not length
Spec-driven development gives this problem a concrete shape.
DeepLearning.AI's Spec-Driven Development with Coding Agents, built with JetBrains, presents project constitutions and feature specifications as a way to preserve context across agent sessions and support a repeatable plan, implement, verify, and replan workflow. The official companion repository makes that structure visible: project-level mission, technology, and roadmap material sits above feature-level plan, requirements, and validation artifacts.
GitHub's public Spec Kit framing uses similar language around specifications as living sources of truth that evolve with the work.
Those are useful examples, not a universal schema for agent work.
What I take from them is narrower. A long prompt can contain enormous detail and still have the wrong lifetime. If the next participant has to reconstruct which instructions remain current, which decisions were revised, and what counts as acceptance, the prompt is functioning as history rather than authority.
A specification earns its place when it becomes the current reference for the job.
That means persistence by itself is not enough. A file can survive for years while becoming less useful every week. The important property is that the artifact can remain current, inspectable, revisable, and connected to the way the work will eventually be judged.
Five signs the conversation is carrying too much authority
I would not create a specification for every agent task. That turns a coordination tool into ceremony, and ceremony has a maintenance cost too.
The threshold becomes easier to see when the work starts exhibiting a few specific properties.
The work crosses conversational boundaries
One session is easy because the original instruction, corrections, and recent decisions can all remain close together.
Once the job crosses sessions, context compaction, restarts, or handoffs, chronology stops being enough. A future participant needs to know what governs the work now, not only what somebody said first.
That current reference should answer the questions that matter for the next decision: what outcome is still required, which constraints still apply, what changed, what is deliberately out of scope, and what will count as completion.
A transcript can preserve all of those answers somewhere and still make the reader reconstruct the contract.
Replanning is expected
Long-running work rarely preserves its first plan unchanged.
Implementation exposes something awkward. A dependency changes. A requirement appears late. A reasonable first approach stops being reasonable.
The problem is not that the plan changed. The problem is having the change arrive as another conversational patch while the old plan remains equally visible and there is no current object that reconciles them.
The DeepLearning.AI workflow explicitly includes replanning. I think that is one of the more useful parts of the example because it treats a specification as something that can absorb change rather than something that freezes the first interpretation.
Version control, amendments, superseded decisions, or a current-state section can all work. The mechanism matters less than the ability to distinguish current intent from previous intent without guessing.
More than one executor needs the same interpretation
Distributed execution makes small ambiguities expensive.
One agent plans. Another implements. A human reviews. A specialist model handles a later step. Each participant may receive a different slice of context and still be locally competent.
A durable specification does not guarantee agreement. It gives disagreement somewhere visible to happen.
Instead of asking whether everyone reconstructed the same conversation, the workflow can ask whether proposed actions remain compatible with one current set of requirements, constraints, and acceptance criteria.
Acceptance needs some distance from generation
This is the point where a transient instruction becomes especially weak.
If the same loop decides what to build, builds it, and then explains why what it built should count as success, the chain can become internally consistent while drifting from the original need.
The DeepLearning.AI companion structure separates requirements, plan, and validation. I would not copy those filenames into every workflow, but the separation itself is useful.
There should be a way to ask two different questions:
- Did the implementation satisfy its checks?
- Do those checks still represent the intended outcome?
Once that second question matters, acceptance needs a durable reference that does not silently move every time the implementation changes.
Recovery is expensive
Eventually, long-running agent work gets interrupted.
A session ends. A branch is abandoned. Context disappears. A partial implementation has to be resumed by someone who was not present for the earlier decisions.
Recovery is much cheaper when the next participant can reconstruct the current job from current artifacts instead of doing conversational archaeology.
The specification does not need to become an exhaustive event log. It needs to preserve the decisions that govern the next action: target state, active constraints, unresolved questions, current plan where relevant, and acceptance.
The more expensive a wrong restart would be, the more useful that recoverable contract becomes.
A persisted document is not automatically authoritative
There is an easy failure mode here: write a file, call it a spec, and assume the authority problem is solved.
It is not.
A stale specification is just old context with a filename.
For a durable specification to govern work, I would want at least five properties.
Addressability. Participants need to know which artifact is the current reference. The same retrieval problem appears from the knowledge side in The Reader Changed the Documentation Contract: information can exist and still fail an agent if the right operational object cannot be found, distinguished, and recovered when needed.
Revision. Changes have to update the governing state rather than accumulate as detached instructions. Replanning should produce a current contract, not a growing pile of exceptions.
Boundaries. The artifact should say what it controls and what it does not. A feature specification should not quietly become product strategy. A technical plan should not overrule a product requirement merely because it was edited later.
Acceptance. The current contract should preserve how completion will be judged. Otherwise implementation can keep redefining success around what it happened to produce.
Inspectability. Humans and agents should be able to see the current rules without reconstructing hidden context.
These are workflow properties rather than requirements for a particular tool.
A Markdown file under version control may be sufficient. Another environment may justify a more structured requirements system. The format is secondary to a simpler question:
What exact artifact is allowed to tell the next executor what the job currently is?
If there is no clear answer, conversational memory is still carrying the authority.
Do not specify uncertainty just to make it look controlled
There is also a point where a specification arrives too early.
Sometimes the work exists precisely because the target is not settled yet. The team is comparing options, discovering constraints, testing whether an idea is useful, or learning which question is worth solving.
Writing that uncertainty into an authoritative document does not remove it. It can give an early guess more weight than it deserves.
The sequence may need to be:
explore → reduce uncertainty → specify what is stable enough to govern → implement → replan when evidence changes the contract
The specification should hold decisions that are ready to govern work and mark the parts that are not. It should not turn discovery into compliance with the first plausible idea.
Prompts still have plenty of territory
None of this makes ordinary prompting obsolete.
A prompt is the lighter tool when the task is local, reversible, easy to inspect, and unlikely to outlive the conversation that defines it. A small edit with a clear failing test may already have a strong acceptance reference. A short analysis can be reviewed directly. A disposable prototype may be useful precisely because nobody wants to preserve its early decisions.
The mistake is not using prompts.
The mistake is letting a prompt remain the authority after the work has become stateful in practice.
That produces a strange kind of continuity. Everyone appears to be working on the same job, but every new session is really making a local interpretation of a moving target. Later messages repair earlier messages while no durable object says what the current target actually is.
At that point, making the prompt longer does not solve the core problem.
The decision rule I would use
The practical question I would ask is:
Will someone need to make a consequential decision about this work after the original conversational context is no longer sufficient?
If the answer is no, a prompt may be enough.
If the answer is yes, persist the part that has to survive.
- If only the desired outcome must survive, preserve the outcome and constraints.
- If the work will replan, preserve current decisions and how newer decisions supersede older ones.
- If several executors or reviewers will participate, give them one current reference.
- If acceptance matters, keep acceptance criteria separable from the implementation's own description of success.
- If interruption is costly, preserve enough current state that the next session can resume without reconstructing intent from history.
The resulting artifact does not have to be large.
It has to be authoritative enough for the lifetime of the work.
That threshold becomes more important as agents get faster because more implementation, tool use, and replanning can happen between two moments when a human naturally restates the problem. The original prompt can remain perfectly visible while becoming operationally obsolete.
When that happens, I do not want a longer prompt.
I want one current place that tells the next participant what must remain true.
// End of transmission. Keep the contract current. — ZYANE
