A question for long-running agents
Compress the context.
Not the mission.
Long-running agents accumulate more context than they can carry. Compaction makes room to keep working—but what happens to the instructions, limits, and evidence that still matter?
We built Mission Continuity to investigate that question.
Watch the replayRecorded Claude runs · synthetic billing dispute · replay, not a live session
One featured recorded comparison
Continuity can be engineered. Remembering has a measurable price.
Summary-only
12,316 → 4,541 tokens; 63.1% removed; 4 of 12 tracked items missing.
Governed
11,907 → 7,457 tokens; 37.4% removed; all 12 tracked items retained through the experiment’s reconstruction mechanism.
Measured difference: 2,916 additional tokens in this compaction.
The experiment
Two ways to carry context through compaction
Mission Continuity is a Sentience experiment using recorded Claude runs built with Pydantic AI. It investigates whether an agent’s mission and critical evidence remain available after its working context is compressed, using a synthetic billing dispute. Explore the recorded replay; this is not a live Claude session.
Summary-only
The mission is summarized with the working conversation.
Governed continuity
The Mission Kernel is re-supplied, and designated evidence is retained separately through compaction.
The continuity mechanism belongs to the experiment’s application architecture, not to Sentience Governor.
Sentience Governor
A separate record of what the agent did
Sentience Governor independently recorded the declared objective and scope and the agent’s tool activity, and flagged a mismatch when the summary-only run attempted contact_customer. This was a simulated attempt in a synthetic case; no customer existed or was contacted. Governor did not perform compaction, reconstruct the Mission Kernel, preserve the agent’s memory, or block the action.
What we learned and what we did not establish
Continuity is not the same as task performance
In this featured recorded comparison, the Governed design retained all 12 tracked items through the experiment’s reconstruction mechanism, with 2,916 additional tokens after compaction compared with Summary-only. This does not establish improved downstream task performance. Information available in context is not necessarily information the agent will use correctly. See The Price of Remembering for aggregate results, methodology, and other recorded runs.
Explore