Anyone building retrieval systems lives with the same awkwardness: one model for text, another for images, a third for audio. Cross-modal retrieval means either stitching together multiple embedding vectors or giving up entirely. On September 21, Alibaba's ATH-MaaS team published a technical report on arXiv (2609.25165) proposing a way out of this fragmentation: Ovis-Omni-Embedding-3B, a unified embedding model covering text, images, video, audio, visual documents, and interleaved multimodal inputs.
No Tower Assembly, Just an Omni Backbone
Most multimodal embedding models follow an "assembly" approach: a text tower plus a vision tower, each encoding separately before alignment. The Ovis team went the opposite way. They took a pretrained Qwen2.5-Omni-3B omni-modal model as the foundation, removed the speech-generation Talker module and the language-modeling head, and kept the native text tokenizer, vision encoder, audio encoder, and the shared Thinker backbone. The final-layer hidden state at the last non-padding token is used directly as the retrieval embedding.
The training recipe has three planks, all centered on "unified": contrastive training with low-rank initialization for adaptation; a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data, with homogeneous-source sampling ensuring informative in-batch negatives; and focal loss to emphasize hard examples plus similarity-based Embedding Distillation transferring fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition allows flexible dimensionality — 2048 dimensions can compress to 1024, 512, 256, or even 128 with limited performance loss.
The Numbers: Audio Leads the Pack
The self-reported scorecard (GitHub README): on MMEB-v3, which spans 190 datasets, Ovis-Omni-Embedding-3B scores 58.46 overall, 5.19 points above the strongest compared baseline's 53.27, ranking first on every modality group. By group, audio shows the largest margin (50.08 vs 43.17, +6.91), agent retrieval comes second (45.52 vs 39.42, +6.10), while the image group leads by only 3.72. Across the 31 aggregate and sub-task entries in the full comparison, it ranks first on 22 and second on 8. MultiConIR is the only entry where it falls outside the top two.
Three supplementary benchmarks are worth noting: 57.29 on MAEB (vs LCO-Embedding-Omni-7B's 53.54) for audio embedding, and 61.77 on MVEB (vs 57.58) for video. But on RTEB, a text-only retrieval benchmark, its edge over Qwen3-Embedding-4B is just 67.35 vs 67.27 — a 0.08-point gap that is essentially noise.
Cold Water and Silver Linings
Two buckets of cold water first. All scores are self-reported by the team; evaluations were run locally and inserted into the corresponding leaderboard snapshots, with no third-party replication yet. Second, and more critical: the README explicitly states model weights are "not open-sourced yet" and will be released in the near future. A model claiming to unify omni-modal retrieval loses much of its industry impact if the weights arrive late or stingy — a 3B-parameter, 2048-dimension embedding only matters if others can actually deploy it.
There are two silver linings. First, 3B parameters is lightweight for the embedding track, yet it swallows six modality groups at once; if weights land as promised, it is a ready-made foundation for RAG, multimodal knowledge bases, and agent memory retrieval. Second, the elastic dimensionality design (2048 down to 128) directly parallels Matryoshka-style approaches, offering an official escape hatch for storage-sensitive deployments.
So what? Embedding models are moving from "one expert per modality" toward "one backbone for everything." Ovis shows the omni-backbone-to-retrieval path works, but the 0.08-point text squeaker and the unreleased weights both remind us: unification is the trend, the moat is not built yet. The day the weights actually drop, this report becomes worth a second read.