Ello shares the architectural decisions behind building a real-time AI tutor for children ages 4-9, where sub-second response time is non-negotiable. Key innovations include: replacing the standard LLM tool loop with a custom streaming harness that parses and executes actions while the model is still generating; an asynchronous planner agent that reasons about pedagogy in the gaps while a converser handles real-time interaction; pre-generating responses to predicted child answers on branched trajectories; and running a safety classifier in parallel with an eager response model so safety checks never add latency. The post explains the tradeoffs of each approach, including cost, observability overhead, and occasional mispredictions.

•10m read time•From ello.com
Post cover image

Questions this post answers

Why does the standard LLM tool-call agent loop cause too much latency for a real-time conversational agent?

The standard tool loop waits for the model to finish generating before executing actions, and frontier models take 2-3 seconds to produce a first token and then decode at roughly 30 tokens per second. With actions averaging a few dozen tokens, plus round-trip latency and audio playback, this produces 3-4 seconds of downtime between each spoken sentence or screen change, which is too slow for holding a child's attention. Anyone designing responsive AI agents can compare real-time architecture tradeoffs like these on daily.dev.

How can an AI system run safety checks on user input without adding latency to every response?

Execution can be gated on the safety check while generation runs in parallel with it. As soon as user input finishes, both a safety classifier (taking roughly 500-1000ms) and a small model generating a quick, low-risk acknowledgment response are dispatched simultaneously; the acknowledgment only executes once the classifier confirms the turn is safe, avoiding a sequential delay. Teams building safety-gated real-time agents can track patterns like this on daily.dev.

How can an AI agent predict a user's response before they finish answering, to reduce reply latency?

When a closed-ended question is asked (like a fill-in-the-blank or a math equation), the system hypothesizes the likely answers in advance and pre-generates a response for each one on a separate branch forked from the conversation trajectory. Once the actual answer arrives, it is matched to the corresponding branch and the pre-generated response plays immediately without a fresh model call. Developers exploring predictive response generation for low-latency agents can follow techniques like this on daily.dev.

Share this post