Models 🇺🇸 10.08.2026 09:01

ByteDance Seed Unveils SeedRealtime: A Full-Duplex Audio-Visual LLM That Watches, Listens, and Speaks in One Model

ByteDanceByteDance
ByteDance's Seed team has introduced SeedRealtime, a native audio-visual full-duplex LLM that fuses audio, video, and text in a unified architecture for real-time, continuous multimodal interaction. It aims to replace cascaded ASR, VLM, and TTS pipelines with a single end-to-end model, enabling proactive speech and natural turn-taking. The model is currently deployed in ByteDance's Doubao app, but no technical report, parameters, weights, or API have been released.
ByteDance's Seed team has unveiled SeedRealtime, a native audio-visual full-duplex large language model that integrates audio, video, and text in a single end-to-end architecture. Unlike traditional cascaded systems that chain automatic speech recognition (ASR), vision-language models (VLM), and text-to-speech (TTS) modules, SeedRealtime performs perception, understanding, decision-making, and expression in parallel, and also handles turn-taking internally, eliminating the need for an external voice-activity detector. The company claims three key breakthroughs: joint audio-visual understanding, proactive interaction, and natural conversational timing. In demonstrations, the model successfully matched names to faces in noisy environments, proactively reminded users when a specific object appeared in camera view, corrected espresso-making steps based on visual state, and suppressed unrelated chatter while recalling off-screen information. SeedRealtime is currently deployed in ByteDance's Doubao app, but the company has not published a technical report, parameter count, open weights, or an API, making it unavailable for third-party integration. ByteDance's internal evaluations report that pacing issues are halved compared to cascaded stacks, but no benchmark or latency figures were provided.
Abbreviations
ASR = Automatic Speech Recognition — Автоматическое распознавание речи
VLM = Vision-Language Model — Модель «зрение-язык»
TTS = Text-to-Speech — Синтез речи из текста
LLM = Large Language Model — Большая языковая модель
API = Application Programming Interface — Программный интерфейс приложения
Source: MarkTechPost — original
Our earlier posts on this topic ↓
Fresh news