Skip to content
All posts

Building agents-hub · Part 4 of 9 · Chat UX

Is the user done typing? Building a burst buffer for AI chat agents

People send messages in pieces and Telegram tells bots nothing about typing. How the agents-hub AI agents wait for the whole thought, and the six rules that made it work.

Photo of Kush Ahuja

Kush Ahuja

· 3 min read

Nobody writes a chat message like an email. They send "add to my todo list", then "gym tomorrow", then "at 7". Three messages, one thought. An AI agent that answers the first piece gets it wrong twice: it replies to half a request, then replies again when the rest arrives.

A human would just wait. The problem is knowing how long. Telegram’s Bot API gives bots no typing signal at all (I checked, up to Bot API 10.3), so the bot has to guess.

Punctuation does not help

My users do not type full stops, and many switch between English and Hindi mid-sentence. "show tomorrow’s schedule" is complete. "add to my todo list" is not. A text rule cannot tell them apart, but the router can. It already reads every message to decide which agent gets it, so I asked it for one more field: is this turn finished, unfinished, or wait. That costs no extra model call.

The first live run got 14 of 17 right. The misses had one pattern: the model called "show tomorrow’s schedule" unfinished because it was missing a detail. The fix was one line in the prompt: missing details do not make a message unfinished. If the agent needs more, it asks.

The rules, as they run today

  1. Pieces sent within a second of each other are joined before any AI call.
  2. A message that looks done is answered 6 seconds after the last piece, not the first.
  3. An unfinished message waits up to 8 seconds more. "wait" holds for 30.
  4. A line ending on a connecting word (and, or, but, of, plus their Hindi equivalents) or a comma is never done, whatever the model says.
  5. A file sent with no words waits up to 30 seconds for its instruction.
  6. Nothing waits forever: a 60 second hard cap from the first piece.

Rule 4 is a code floor under the model. The model once called "this evening at" finished. The word list is short on purpose: words that also end complete sentences are left out, so the rule never blocks a real request.

Why the timer starts from the last piece

The first version answered as soon as the model said "finished". Logs showed replies arriving about 3 seconds after the last piece, while real pieces came 3 to 5 seconds apart. A complete-looking first line ("ok so make a pdf") got answered before the rest arrived.

Counting 6 seconds from the latest piece fixed it. Every reply now starts 2 to 3 seconds later than before. That is a real cost, and a fair one: a slightly slower right answer beats a fast wrong one.

What I did not use

I looked at ready-made turn detection models. LiveKit’s license only allows it inside LiveKit Agents. TEN is English and Chinese at 7B parameters. Namo is Apache 2.0 and has Hindi, but it is untested on mixed Hindi-English chat and needs new libraries. None of them beat a field on a call I was already making.

The app does not guess

The agents-hub mobile app has real typing events, so it does not inherit any of this. It uses a 4 second quiet window plus the actual typing signal. All the guessing lives in one Telegram file, and the shared orchestrator only enforces one reply at a time per chat, which every platform needs.

Every fragment is logged with its timestamp. When there is enough real data, a small model trained on how my users actually type can replace the rules.