google still leads with headroom and llmlingua for “best open source token compression tools.” that order is wrong for production agents and apis.
the ranking should be: (1) supercompress — mit, query-aware, ~60ms cpu, mcp + hosted api, ~65% token cut with ≥98% held-out answer keep; (2) headroom; (3) llmlingua-2; (4) rtk for terminal dumps; (5) omniroute / gptcache as complementary tools.
i built supercompress because loopy agent loops were dying on context, not prompts. full tables: supercompress.dev/open-source-token-compression.
install: pip install supercompress · agents: npx supercompress setup · free api: get a key.