LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Quick Answer
LookME introduces a novel framework for enhancing multimodal embeddings in Vision-Language Models (VLMs) through a hierarchical two-level lookup method.
Quick Take
This approach improves efficiency and performance by prioritizing critical embeddings and enabling on-demand loading, outperforming traditional text-only PLE methods in various visual benchmarks.
Key Points
- LookME enables efficient lookup of multimodal embeddings from large-scale embedding tables.
- The framework uses a coarse-to-fine strategy for scene-level to primitive-level lookups.
- It integrates a sparse injection strategy to prioritize critical embeddings in layers.
- Experiments show LookME outperforms text-only PLE methods on visual benchmarks.
- The approach supports partitioned storage, enhancing deployment in resource-constrained environments.
DeepSignal Analysis
What happened
LookME is a new framework designed to enhance multimodal embeddings in Vision-Language Models (VLMs) using a hierarchical two-level lookup method. This method aims to improve efficiency and performance by prioritizing critical embeddings and allowing for on-demand loading, which is particularly beneficial in resource-constrained environments.
Key evidence
- LookME introduces a hierarchical two-level lookup method that performs lookups from scene-level to intra-scene primitive-level, enhancing the retrieval of multimodal embeddings.
- The framework integrates a sparse injection strategy that prioritizes critical embeddings over larger sets, improving efficiency and performance in VLMs.
- Experiments conducted on multiple visual benchmarks demonstrate that LookME outperforms traditional text-only Per-Layer Embedding methods.
Why it matters
The development of LookME addresses significant challenges in deploying Vision-Language Models in environments with limited resources. By improving the efficiency of embedding retrieval, it allows for better performance without the high memory costs associated with full model loading. This advancement could lead to broader applications of VLMs in real-world scenarios where computational resources are constrained.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance limits deployment in resource-constrained environments due to the trade-off between high memory usage from full loading and increased latency from on-demand loading. Recently, the Per-Layer Embedding (PLE) architecture addresses this by scaling models with large external embedding tables stored in ROM and performing lightweight lookup to retrieve relevant embeddings to enhance token representations. Nevertheless, existing PLE-style methods are primarily designed for text embeddings due to the convenience of ID-based retrieval, limiting their effectiveness in VLMs where multimodal embeddings contain richer information for visual tasks. In this paper, we propose LookME, the first framework that enables lookup-based enhancement for multimodal embeddings in VLMs while supporting partitioned storage and on-demand loading. To efficiently lookup arbitrary continuous multimodal embeddings from large-scale embedding tables, we propose a hierarchical two-level lookup method employing a coarse-to-fine strategy that performs lookups from the scene-level to the intra-scene primitive-level. Furthermore, we integrate the lookup method with a sparse injection strategy, which adaptively prioritizes critical embeddings over voluminous multimodal embeddings within layers, and facilitates embedding table reuse across neighboring layers, improving the trade-off among efficiency, model size, and performance. Experiments on multiple visual benchmarks show that LookME outperforms text-only PLE-style methods, validating the effectiveness of lookup-based multimodal embedding enhancement.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.16305 [cs.CV] |
| (or arXiv:2607.16305v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16305 arXiv-issued DOI via DataCite |
Submission history
From: Zeyu Xu [view email]
[v1]
Tue, 14 Jul 2026 08:03:00 UTC (87 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.