ByteDance's Seed team released SeedRealtime on Wednesday, a model that natively fuses audio, video and text in one end-to-end system for full-duplex interaction: it keeps watching and listening even while it speaks, instead of taking turns. The company says it reads sound, visuals and timing together to work out who is addressing it and what they want, and it is already live in the Doubao assistant app. It extends April's voice-only Seeduplex model to the visual world.