Same next-token loop, richer input and output. Images and audio get encoded into tokens the transformer already consumes — so the new skill isn't the model, it's knowing when a multimodal model beats a classic OCR / ASR pipeline. That call is all this page installs.
Tap a stage. The whole trick is that pixels and sound become tokens too.
One architecture, richer input — vision and voice are not new models, just new token streams into the loop you already know.
Tap a card to flip it. Two honest columns.
The classic line: reach for the model when the input is messy or needs reasoning, the pipeline when it's clean and high-volume.
Tap a stage. Three boxes — but the words were never the hard part.
The latency budget is everything: every hop adds delay, and a beat too slow feels broken. The hard part is turn-taking and interruptions — knowing when to stop and listen — not transcribing the words. Newer speech-to-speech models collapse the three boxes into one, cutting the lag.
Pick the situation you actually have. Get the boring, correct answer.