Alibaba Cloud has launched Qwen2.5-Omni-7B, a unified end-to-end multimodal model in the Qwen series. Uniquely designed for comprehensive multimodal perception, it can process diverse inputs, including text, images, audio, and videos, while generating real-time text and natural speech responses. This sets a new standard for optimal deployable multimodal AI for edge devices like mobile phones and laptops.
Qwen2.5-Omni-7B delivers remarkable performance across all modalities, rivaling specialized single-modality models of comparable size. Notably, it sets a new benchmark in real-time voice interaction, natural and robust speech generation, and end-to-end speech instruction following.
Qwen2.5-Omni-7B was pre-trained on a vast, diverse dataset, including image-text, video-text, video-audio, audio-text, and text data, ensuring robust performance across tasks.
Read More News on Business News Malaysia
Read More News on Business News Malaysia
The opening ceremony witnessed the presence of senior representatives from diplomatic missions, government agencies, trade…
Malaysian businesses prepare for US-Iran tensions by strengthening risk strategies to ensure stability amid potential…
The US will impose 10% to 12.5% tariffs on 60 trading partners over forced labor…
KIPREIT posts record FY26, strong outlook with Setapak Central acquisition, AEIs, and rental reversions supporting…
Analysts expect cautious trading in KLCI today due to Wall Street declines, new tariffs, and…
Rystad warns oil prices hinge on resilience of flows, with escalation risks tightening markets and…
This website uses cookies.