A question for long-running agents

Compress the context.
Not the mission.

Long-running agents accumulate more context than they can carry. Compaction makes room to keep working—but what happens to the instructions, limits, and evidence that still matter?

We built Mission Continuity to investigate that question.

Watch the replay

Recorded Claude runs · synthetic billing dispute · replay, not a live session

One featured recorded comparison

Continuity can be engineered. Remembering has a measurable price.

Summary-only

12,316 → 4,541 tokens; 63.1% removed; 4 of 12 tracked items missing.

Governed

11,907 → 7,457 tokens; 37.4% removed; all 12 tracked items retained through the experiment’s reconstruction mechanism.

Measured difference: 2,916 additional tokens in this compaction.

One featured recorded compaction comparison: Summary-only, 12,316 to 4,541 tokens, 63.1% removed, 4 of 12 tracked items missing; Governed, 11,907 to 7,457 tokens, 37.4% removed, all 12 retained through the experiment’s reconstruction mechanism; 2,916 additional tokens in this compaction.
One featured recorded comparison from the Mission Continuity replay, not an aggregate experimental result. Token retention and removal are specific to this compaction, workload, and policy. Retaining additional context does not establish improved task performance.

The experiment

Two ways to carry context through compaction

Mission Continuity is a Sentience experiment using recorded Claude runs built with Pydantic AI. It investigates whether an agent’s mission and critical evidence remain available after its working context is compressed, using a synthetic billing dispute. Explore the recorded replay; this is not a live Claude session.

Summary-only

The mission is summarized with the working conversation.

Governed continuity

The Mission Kernel is re-supplied, and designated evidence is retained separately through compaction.

The continuity mechanism belongs to the experiment’s application architecture, not to Sentience Governor.

Sentience Governor independently recorded the declared objective and scope and the agent’s tool activity, and flagged a mismatch when the summary-only run attempted contact_customer. This was a simulated attempt in a synthetic case; no customer existed or was contacted. Governor did not perform compaction, reconstruct the Mission Kernel, preserve the agent’s memory, or block the action.

In this featured recorded comparison, the Governed design retained all 12 tracked items through the experiment’s reconstruction mechanism, with 2,916 additional tokens after compaction compared with Summary-only. This does not establish improved downstream task performance. Information available in context is not necessarily information the agent will use correctly. See The Price of Remembering for aggregate results, methodology, and other recorded runs.

Explore

Explore the runs and the engineering log