Every method for cutting token consumption in LLM and agentic pipelines — prompt caching, semantic caching, context compaction, prompt compression, retrieval instead of stuffing, and more — the arithmetic that tells you which one to reach for first, the ways each of them fails, and what changes when the pipeline becomes a fleet of agents.

Read it

Open the monograph →

The monograph lives at its own URL in the warm-paper layout, with figures and worked arithmetic.


← Back to Autonomy