autoregressive
Writing an answer one piece at a time, each piece chosen based on what came before.
Most chat assistants are autoregressive. They write a reply one token at a time, and each new token is picked by looking at everything written so far. This is why answers appear word by word on screen. It is flexible, but a long answer takes many steps, so it is slower and costs more.
Think of it as
Like writing a sentence where you choose each word only after rereading all the words before it.
Example
Ask a chat assistant for a 300-word summary and you watch it appear word by word. Each of those words was a separate step, which is why long answers take noticeably longer than short ones.