arxiv:2609.21465
Published on Sep 18
· Submitted by
Haolin HE on Sep 21
· Qwen
Upvote
31
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, a good reply often needs to account for the user's surroundings, facial expressions, and nearby objects, and such responses can be expressed in many different ways, making keyword matching unreliable for evaluating reply quality. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models' basic dialogue abilities across five ability categories. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.
View arXiv page View PDF GitHub 33 Add to collection
Community
Paper submitter 1 day ago
•
edited 1 day ago
Huggingface Repo: https://hf-mirror.com/datasets/Harland/OmniVChat Github Repo: https://github.com/HarlandZZC/OmniVChat
Paper submitter 1 day ago
Native audio-visual dialogue without ASR/intermediate text has long suffered from two big headaches: data scarcity and open-ended evaluation.
This paper tackles both nicely:
OmniVChat-Studio & Bench: Uses multi-agent synthesis to create rich single/multi-turn Audio-visual dialogues across 5 capability categories.
OmniVChat-RL: Balances reply accuracy, efficiency, and natural dialogue style.
Solid Sim-to-Real Transfer: Tested on Qwen3-Omni-Instruct, showing clear improvements on both synthetic and human-recorded benchmarks.
about 12 hours ago
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
-
Training-Free Speech-Centric Omni Understanding with Frozen VLMs (2026)
-
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence (2026)
-
Realtime-Venus: A full-duplex interaction system with asynchronous delegation (2026)
-
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos (2026)
-
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming (2026)
-
Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on HF Mirror checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
about 7 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from HF Mirror
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
· Sign up or log in to comment
Upvote
31
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.21465 in a model README.md to link it from this page.
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.21465 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.