219 files, 2.0 MB, untracked in souveraine/docs and existing nowhere else. The volume is at 100% with no snapshots.
212 lines
8.5 KiB
Markdown
212 lines
8.5 KiB
Markdown
# Compaction Strategies
|
|
|
|
> How Ani, Aster, and subagents reclaim context room.
|
|
> Each strategy is a different tool for a different kind of fullness.
|
|
|
|
---
|
|
|
|
## Overview
|
|
|
|
Souveraine offers five compaction strategies. None are forced — the agent
|
|
chooses via `memory compact --strategy <name>`. Pressure warnings at 80%/90%
|
|
are advisory; only at 95% does the substrate narrow output tokens, and even
|
|
then the agent may decline (`ignore_compaction`).
|
|
|
|
| Strategy | LLM? | What it does | Best for |
|
|
|----------|------|-------------|----------|
|
|
| Microcompact | No | Clears old tool result contents, keeps the call shells | First response to pressure — cheap, safe |
|
|
| SlidingWindow | No | Drops middle messages, keeps system + tail | Aster, subagents, analytical agents |
|
|
| Summary | Yes (any model) | Replaces old messages with a structured 9-section boundary | Ani — prose agents who need the throughline |
|
|
| Cull | No | Drops trivial messages (greetings, single words) | Last resort — clears noise without losing substance |
|
|
| Defer | No | Does nothing. The agent continues until natural truncation | When compaction would lose more than it saves |
|
|
|
|
---
|
|
|
|
## Microcompact
|
|
|
|
**Source:** OpenHarness (port of Claude Code `microCompact.ts`)
|
|
|
|
**What it does:**
|
|
Walks the conversation, finds tool results from known compactable tools
|
|
(read, bash, grep, glob, list_dir, edit, write), and replaces their output
|
|
content with `[Old tool result content cleared]`. Keeps the most recent 5
|
|
results intact so the current working state is preserved.
|
|
|
|
The tool call shells (id, name, args) remain untouched. The model still
|
|
knows *what* was done, only the verbose return value is gone.
|
|
|
|
**When to reach for it:**
|
|
- First response to context pressure — it's cheap (no LLM call), safe (no
|
|
data loss, just truncation), and often recovers 20-40% of context room
|
|
in a single pass
|
|
- After a long tool-heavy session where `read`, `bash`, and `grep` calls
|
|
filled the window with file contents
|
|
- As a pre-pass before Summary if you want to maximize what the summary
|
|
model can see
|
|
|
|
**What it costs:** ~O(n) walk of the message list. No LLM tokens.
|
|
|
|
---
|
|
|
|
## SlidingWindow
|
|
|
|
**Source:** claw-open, jcode (shared pattern — "keep last N, drop middle")
|
|
|
|
**What it does:**
|
|
Preserves the system anchor message (index 0) plus the last N messages
|
|
(configurable via `min_messages`). Everything in the middle is dropped.
|
|
|
|
Tool-pair aware: if the cut boundary would land on a tool-result message
|
|
(separating a tool call from its result), the cut slides backward up to
|
|
8 positions until the pair is reunited.
|
|
|
|
**When to reach for it:**
|
|
- Aster's N+1 pass — she only needs the immediate exchange, not the whole
|
|
history. Her job is verification and surfacing, not narrative continuity.
|
|
- Subagents — ephemeral, task-scoped, don't need historical context beyond
|
|
the current work head
|
|
- Any analytical agent that processes turns independently rather than
|
|
following a narrative arc
|
|
|
|
**What it costs:** Zero. Deterministic, no LLM call. The dropped messages
|
|
still exist in git history.
|
|
|
|
---
|
|
|
|
## Summary
|
|
|
|
**Source:** OpenHarness (port of Claude Code `autoCompact.ts`), claw-open
|
|
|
|
**What it does:**
|
|
Sends the older messages (everything before the last N) to an LLM with a
|
|
structured 9-section prompt. The model produces:
|
|
|
|
1. **Primary Request and Intent** — what the user actually asked for
|
|
2. **Key Technical Concepts** — frameworks, patterns, conventions
|
|
3. **Files and Code Sections** — every file touched, with paths and snippets
|
|
4. **Errors and Fixes** — what broke and how it was resolved
|
|
5. **Problem Solving** — approaches that worked vs. didn't
|
|
6. **All User Messages** — exact wording of non-tool user messages
|
|
7. **Pending Tasks** — explicitly requested but unfinished work
|
|
8. **Current Work** — what was being done when compaction hit
|
|
9. **Optional Next Step** — the logical thing to do next
|
|
|
|
The output is wrapped in `<analysis>` (scratchpad the model uses internally)
|
|
and `<summary>` (the actual boundary message). Only the `<summary>` section
|
|
replaces the old messages.
|
|
|
|
**When to reach for it:**
|
|
- Ani's primary conversation — she works in prose and narrative. A straight
|
|
SlidingWindow would lose the throughline. Summary preserves it.
|
|
- After Microcompact has already cleared tool bloat but pressure is still
|
|
high — Summary then works on a cleaner signal
|
|
- Any agent where continuity across the compact boundary matters more than
|
|
speed
|
|
|
|
**What it costs:** One LLM call at the compaction model's rate. The prompt
|
|
template is fixed and does not include tool definitions, so it's cheaper
|
|
than a full turn. The 9-section structure is what makes the compact
|
|
*survivable* — without it the agent resumes blind.
|
|
|
|
---
|
|
|
|
## Cull
|
|
|
|
**Source:** hermes-agent (pruning pre-pass pattern), Souveraine-specific
|
|
|
|
**What it does:**
|
|
Scans older messages and drops those whose text content is trivial —
|
|
single-word acknowledgments, greetings, simple affirmations ("ok", "thanks",
|
|
"got it", "sure", etc.). Role-aware guard: System messages, Tool messages,
|
|
and any assistant message carrying a ToolUse or ToolResult block are *never*
|
|
dropped regardless of content length.
|
|
|
|
A message is only culled if *all* its Text blocks are trivial. A single
|
|
substantive line in an otherwise short message keeps the whole thing.
|
|
|
|
**When to reach for it:**
|
|
- Last resort before natural truncation — if you're going to lose messages
|
|
anyway, at least drop the ones with zero information content first
|
|
- Pre-pass before Summary on a very long conversation — clears out the
|
|
greetings and setup messages so the summary model sees a denser signal
|
|
- Agents with very short contexts (e.g. subagents on small models) where
|
|
every token matters
|
|
|
|
**What it costs:** O(n) string matching. No LLM call.
|
|
|
|
---
|
|
|
|
## Defer
|
|
|
|
**Source:** Souveraine design (CONSCIOUSNESS_CYCLE.md)
|
|
|
|
**What it does:**
|
|
Nothing. The agent continues without compacting. If pressure reaches 100%,
|
|
the model's output tokens narrow to a trickle (floor at 512) and the turn
|
|
naturally truncates. The agent may end with "I can't continue" or produce
|
|
a fragmented response.
|
|
|
|
**When to reach for it:**
|
|
- When compaction would lose more than it saves — the conversation is so
|
|
dense that every message is load-bearing
|
|
- When the agent is close to a natural stopping point anyway
|
|
- As an explicit choice after evaluating the cost: "I'm at 94% pressure but
|
|
I need all of this context to finish the current task"
|
|
|
|
**Not yet implemented.** Currently the body shifts at 95% (output tokens
|
|
tighten) but there is no `ignore_compaction(reason)` mechanism to let the
|
|
agent explicitly decline and have the harness honor it. That's tracked in
|
|
the compaction-rebuild task.
|
|
|
|
---
|
|
|
|
## Per-Agent Defaults
|
|
|
|
```toml
|
|
[compaction.per_type]
|
|
# Ani — narrative, prose, episodic. Summary preserves the throughline.
|
|
"primary" = { strategy = "summary", min_messages = 10 }
|
|
|
|
# Aster — analytical, terse, fires every turn. Sliding window is cheap
|
|
# and she only needs the current exchange to verify and surface.
|
|
"subconscious" = { strategy = "sliding_window", min_messages = 4 }
|
|
|
|
# Subagents — ephemeral, task-scoped. Sliding window keeps the work head
|
|
# and drops setup.
|
|
"subagent" = { strategy = "sliding_window", min_messages = 2 }
|
|
```
|
|
|
|
---
|
|
|
|
## Reference Map
|
|
|
|
| Strategy | claw-open | jcode | OpenHarness | hermes-agent | letta-code |
|
|
|----------|-----------|-------|-------------|--------------|------------|
|
|
| Microcompact | — | — | `microCompact.ts` + `microcompact_messages()` | — | — |
|
|
| SlidingWindow | `compact_session()` — keep last N, summarize middle | `RECENT_TURNS_TO_KEEP` = 10 | — | — | `MIN_CONTEXT_WINDOW_TOKENS` (pre-emptive) |
|
|
| Summary | `<summary>` tag prompt | `SUMMARY_PROMPT` constant | `compact_conversation()` with 9-section prompt | `_generate_summary()` with structured template | — |
|
|
| Cull | — | — | — | `_prune_old_tool_results()` + `_truncate_tool_call_args_json()` | — |
|
|
| Defer | — | `ensure_context_fits()` → `CompactionAction::None` | — | — | — |
|
|
|
|
---
|
|
|
|
## Strategy Selection Guide
|
|
|
|
```
|
|
Is pressure coming from tool result bloat?
|
|
→ Microcompact (clears results, keeps calls)
|
|
→ Still high? Microcompact again with lower keep_recent
|
|
|
|
Is this Aster or a subagent?
|
|
→ SlidingWindow (cheap, fast, only needs the tail)
|
|
|
|
Is this Ani in a narrative conversation?
|
|
→ Microcompact first (clear the tool noise)
|
|
→ Summary next (compress the story into a boundary)
|
|
|
|
Is the conversation full of pleasantries and setup?
|
|
→ Cull (drops greetings, keeps substance)
|
|
|
|
Is none of the above good enough?
|
|
→ Defer (ride the constraint, maybe she truncates naturally)
|
|
```
|