agents-hub keeps one memory per customer across every private chat, and a separate one per group. Five agents share it. Sending the whole history to every agent on every message is simple and expensive, and it also confuses them: the Tasks agent does not need to read a long email thread to add "buy milk".
Memory is split into three layers.
Layer 1: recent messages, scoped by routing
The hub already decides which agent each message is for. That same decision is saved as a tag on the message, and each agent is sent only the recent messages routed to it. The hub sees everything, since it needs that to route. No extra model call is involved: the routing decision already exists.
Two details made it work.
- The 40-message window is applied before the per-agent filter, not after. My first version filtered first, and with a full window every agent still got 40 messages. Tokens fell by only 1 to 3 percent.
- The last 4 messages go to every agent whatever their tag, so a request that builds on what another agent just did keeps its context.
Where plain filtering broke
"Put the deadline from that email on my tasks." The Tasks agent could not see the email, and got it right 0 times out of 3. The fix: the hub, which sees everything, can name other agents whose messages this agent may also read, in an also_sees field of the same routing JSON. It costs about 109 tokens per hub call. After that, the hardest reference test scored 7 of 8 scoped, against 8 of 8 with the full history. Small sample, real limit: it fails when the hub does not name the right agent.
Layer 2: facts, sent every time
Some things should never scroll out of view: "Ram’s email is ram@...", "my office is in Cyber Hub". These are stored as short facts and sent with every request. It is a small, fixed cost for the things people get annoyed about repeating.
Layer 3: the archive, searched on demand
Nothing is deleted. Older messages and every file the agents made go into an archive and a vault, and agents reach them with a search_memory tool only when a request needs it. The search is hybrid: keyword and spelling-tolerant matches, merged with reciprocal rank fusion, then a query coverage term and a recency boost. It answers in about 0.2 to 0.4 seconds.
The coverage term was not in the plan. Rank fusion alone was not good enough on the test data, so results now also score on how much of the question they actually match. Meaning-based search is ready to slot in once an embeddings model is deployed.
The numbers
On a real log of 47 messages and 14 agent calls, history fell from 449 to 285 tokens per agent call, 36.5 percent less, and messages meant for other agents dropped from 6.0 to 1.7 per call. In a 600-turn simulation built from the same log, the saving was 50.6 percent with 3 agents and 81 percent with 10.
The honest part: the whole request fell only 6 percent.
About 2,300 tokens of persona, rules and tool definitions go into every agent call and scoping does not touch them. That is the next thing to cut. The memory work is now the basis of a research paper on delivering shared memory to the agents of a multi-agent assistant.