Every method for cutting token consumption in LLM and agentic pipelines — prompt caching, semantic caching, context compaction, prompt compression, retrieval instead of stuffing, and more — the arithmetic that tells you which one to reach for first, the ways each of them fails, and what changes when the pipeline becomes a fleet of agents.
Read it
The monograph lives at its own URL in the warm-paper layout, with figures and worked arithmetic.
← Back to Autonomy