search
multimodal AI
Trends
- 1Google releases EmbeddingGemma 2, an open lightweight multimodal embedding modelโEmbeddingGemma 2: An open, lightweight multimodal embedding model
Google has released EmbeddingGemma 2, an open, lightweight multimodal embedding model announced on the company's developer blog. The model is designed to convert text and other inputs into embeddings for search and retrieval tasks while remaining small enough to run on modest hardware. Developers are discussing the release, with attention on its open availability and what a small multimodal embedding model means for building search, RAG, and classification applications without heavy compute.
- 2
German AI company Aleph Alpha has released a technical report on Kolibri, its multimodal foundation model. The PDF, published on the company's website, is being discussed by technology readers, with many weighing how the German challenger's approach and capabilities compare to larger US-based AI labs.
- 3
Google has released EmbeddingGemma 2, a new embedding model the company says can run directly on devices and handle multimodal inputs. The launch is being discussed among developers and AI watchers as part of the push to bring capable AI models to phones and other consumer hardware without relying on cloud servers. Reaction online has focused on what the model means for privacy, latency and on-device applications.
- 4UniEvo-VL Uses Self-Distillation for Multimodal Self-ImprovementโUniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement
A new paper introduces UniEvo-VL, a multimodal AI model trained through self-distillation, allowing it to improve its own performance without relying on large amounts of externally labeled data. The approach is being discussed among researchers as an example of growing interest in self-improving model training methods, and the paper is available on arXiv.
- 5TwelveLabs launches Pegasus 1.6 video model for physical AIโผPegasus 1.6 brings video understanding to physical AI, says TwelveLabs
TwelveLabs has released Pegasus 1.6, a video understanding model aimed at physical AI applications such as robotics. The company says the model can analyze video input to help machines and robots interpret real-world visual environments, extending multimodal AI beyond screen-based tasks into embodied systems operating in physical spaces.
- 6Mistral AI launches trillion-parameter Large 4 modelโผMistral AI has launched its Large 4 multimodal model, nicknamed le Chonk, with one trillion parameters (49 billion activ
French AI company Mistral has released Large 4, a multimodal model with one trillion parameters, of which 49 billion are active at inference, jokingly nicknamed le Chonk. The company says it trained the model on 3,800 NVIDIA GPUs in its European data centres and reports an 82% score on a vulnerability reproduction-and-patching test.
- 7Run LLMs Locally adds EmbeddingGemma2 multimodal embeddingsโNew update: Run LLMs Locally Added EmbeddingGemma2 using llama.cpp. It includes multi modality, allowing to generate emb
A developer released an update to Run LLMs Locally adding EmbeddingGemma2 support via llama.cpp. The new version handles multimodal inputs, generating embedding vectors from text, images, video, and audio, which makes it possible to build local search engines over documents and photos. The project is available on GitHub, and it is drawing attention among people interested in running AI tools offline on their own hardware.
- 8Reka AI unveils Rho-1 omni-model spanning text to robot controlโผReka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model
AI startup Reka AI has introduced Rho-1, an omni-model designed to process text, images, and video while also controlling robots, all within a single model. The announcement is drawing attention for combining multimodal understanding with real-world robotic control, a capability usually handled by separate systems. Observers are watching closely to see how Rho-1 performs against established multimodal AI offerings.
- 9Google DeepMind releases open EmbeddingGemma 2 embedding modelโGoogle DeepMind has released EmbeddingGemma 2, an open embedding model that maps text, code, images,... # ai # automatio
Google DeepMind has released EmbeddingGemma 2, an open embedding model that maps text, code, images and other inputs into shared representations, allowing search and retrieval across different data types. The model is designed to run on-device rather than in the cloud, making it free and practical for local applications. Developers and tech commentators are highlighting its multimodal capabilities and its usefulness for search, coding and automation tools.
- 10Mistral AI launches Mistral Large 4 public previewโMistral AI officially launched the Mistral Large 4 public preview. Discover how this trillion-parameter multimodal model
Mistral AI has officially launched the public preview of Mistral Large 4, a trillion-parameter multimodal model. The French AI company says the new system is built to rival top closed-source models from competitors, marking a significant step in Europe's push into frontier-scale artificial intelligence. Early reactions in tech communities focus on its scale and open-ecosystem implications.
- 11AI forces a rethink of what it means to understand imagesโWhat does it mean to understand an image in the age of AI? Visual epistemology shifts attention from simply recognizing
Researchers and commentators are debating how image understanding should be defined now that multimodal AI systems can interpret pictures. The discussion draws on visual epistemology, shifting the focus from what images contain to how seeing, knowledge, belief and interpretation interact to produce meaning, and questioning whether machines genuinely understand images or merely simulate recognition.
- 12Google launches EmbeddingGemma 2 for on-device multimodal searchโBring multimodal semantic search to the edge with EmbeddingGemma 2
Google has announced EmbeddingGemma 2, a model designed to bring multimodal semantic search to edge devices. The release lets developers run embedding-based search across text, images and other media locally, without sending data to the cloud. It signals Google's continued push to move AI capabilities onto phones and other resource-constrained hardware.
- 13New VA-Bench Shows AI Models Struggle With Robot TasksโDalian University of Technology's VA-Bench: Top Multimodal Models Finish Only Half of Robot Tasks
Dalian University of Technology has introduced VA-Bench, a benchmark for evaluating multimodal AI models on robotics tasks. Results show that even the top-performing models complete only about half of the tested tasks, highlighting a significant gap between current multimodal capabilities and practical robot control. The benchmark is drawing attention as a measure of how far AI still is from reliable real-world robotics.
- 14BostonGene to Present AI Approach in Cancer Drug DevelopmentโผBostonGene Chief Medical Officer to Discuss Biologically Grounded AI and Multimodal Data in Oncology Drug Development at the 2nd Annual Global Cancer Research & Innovation Symposium
BostonGene's Chief Medical Officer is scheduled to speak at the 2nd Annual Global Cancer Research & Innovation Symposium, where the focus will be on biologically grounded artificial intelligence and the use of multimodal data in oncology drug development. The company argues that combining AI with biological context can improve how cancer therapies are discovered and developed.