Most agent replies take a few seconds. The Docs agent is different. It writes Python in a sandbox, renders a PDF or a Word file, checks it and fixes it. A real job takes one to four minutes.
On 28 September I asked it to edit a report. The edit took four minutes. I sent "where is the report?" and it waited behind the job with nothing on screen but "typing...". I could not tell if anything was happening. If I could not, a customer never would.
Let go of the chat after 15 seconds
Rounds used to run strictly one at a time per chat. That rule is still right for quick work: a Gmail draft, then "send it", must happen in order. So it stays for anything short. A round that is still working after 15 seconds stops holding the chat and carries on by itself.
One status message, edited
The job posts one message and edits it as it goes: the steps so far and the time elapsed, then a tick, an undo arrow or a warning at the end. I considered a new message per step. For a four-minute job that is four or five pings in a row, which is worse than silence.
The hub knows what is running
Running jobs are listed in one place, and the router sees that list when it plans the next message. That gives three behaviours:
- "Where is it?" or "hurry up": the hub answers from the job’s steps, without disturbing it.
- A change to the same work ("make the title red too"): the job is stopped and started again with the original request plus the change, on the same files.
- Anything else: runs alongside, in parallel.
Restart versus queue was a real trade-off. Waiting for the job to finish and then applying the change is cheaper and wastes no sandbox time, but the right file arrives later. I chose restart, because the customer is waiting for the file, not for my compute bill. Every agent request carries an abort signal, so stopping is clean.
The limits
Running jobs live in the server process, so a restart kills them. The request itself is saved when the round starts, not when it ends, so "try again" always works. A /reset while a long job runs cannot stop that job from posting its reply afterwards. Both go away when jobs move to a real queue, which is on the list for scaling past one server.