MikeTrendsTrends right now

search

multimodal AI

Trends

  1. 1
    Google releases EmbeddingGemma 2, an open lightweight multimodal embedding modelโ—EmbeddingGemma 2: An open, lightweight multimodal embedding modelYhnTechnology41337 min ago

    Google has released EmbeddingGemma 2, an open, lightweight multimodal embedding model announced on the company's developer blog. The model is designed to convert text and other inputs into embeddings for search and retrieval tasks while remaining small enough to run on modest hardware. Developers are discussing the release, with attention on its open availability and what a small multimodal embedding model means for building search, RAG, and classification applications without heavy compute.

  2. 2
    Aleph Alpha Publishes Tech Report for Kolibri Modelโ–ผKolibri โ€“ Tech Report [pdf]YhnTechnology1099 h ago

    German AI company Aleph Alpha has released a technical report on Kolibri, its multimodal foundation model. The PDF, published on the company's website, is being discussed by technology readers, with many weighing how the German challenger's approach and capabilities compare to larger US-based AI labs.

  3. 3

    Google has released EmbeddingGemma 2, a new embedding model the company says can run directly on devices and handle multimodal inputs. The launch is being discussed among developers and AI watchers as part of the push to bring capable AI models to phones and other consumer hardware without relying on cloud servers. Reaction online has focused on what the model means for privacy, latency and on-device applications.

  4. 4
    UniEvo-VL Uses Self-Distillation for Multimodal Self-Improvementโ—UniEvo-VL: Self-Distillation Training for Multimodal Model Self-ImprovementYhn1319 h ago

    A new paper introduces UniEvo-VL, a multimodal AI model trained through self-distillation, allowing it to improve its own performance without relying on large amounts of externally labeled data. The approach is being discussed among researchers as an example of growing interest in self-improving model training methods, and the paper is available on arXiv.

  5. 5
    TwelveLabs launches Pegasus 1.6 video model for physical AIโ–ผPegasus 1.6 brings video understanding to physical AI, says TwelveLabsโœ‰newsTechnologyRobotics10 h ago

    TwelveLabs has released Pegasus 1.6, a video understanding model aimed at physical AI applications such as robotics. The company says the model can analyze video input to help machines and robots interpret real-world visual environments, extending multimodal AI beyond screen-based tasks into embodied systems operating in physical spaces.

  6. 6
    Mistral AI launches trillion-parameter Large 4 modelโ–ผMistral AI has launched its Large 4 multimodal model, nicknamed le Chonk, with one trillion parameters (49 billion activMmastodonTechnology11 d ago

    French AI company Mistral has released Large 4, a multimodal model with one trillion parameters, of which 49 billion are active at inference, jokingly nicknamed le Chonk. The company says it trained the model on 3,800 NVIDIA GPUs in its European data centres and reports an 82% score on a vulnerability reproduction-and-patching test.

  7. 7
    Run LLMs Locally adds EmbeddingGemma2 multimodal embeddingsโ—New update: Run LLMs Locally Added EmbeddingGemma2 using llama.cpp. It includes multi modality, allowing to generate embMmastodonTechnologyAI41 d ago

    A developer released an update to Run LLMs Locally adding EmbeddingGemma2 support via llama.cpp. The new version handles multimodal inputs, generating embedding vectors from text, images, video, and audio, which makes it possible to build local search engines over documents and photos. The project is available on GitHub, and it is drawing attention among people interested in running AI tools offline on their own hardware.

  8. 8
    Reka AI unveils Rho-1 omni-model spanning text to robot controlโ–ผReka AI's omni-model Rho-1 handles text, images, video, and robot control in a single modelโœ‰newsTechnologyRobotics1 d ago

    AI startup Reka AI has introduced Rho-1, an omni-model designed to process text, images, and video while also controlling robots, all within a single model. The announcement is drawing attention for combining multimodal understanding with real-world robotic control, a capability usually handled by separate systems. Observers are watching closely to see how Rho-1 performs against established multimodal AI offerings.

  9. 9
    Google DeepMind releases open EmbeddingGemma 2 embedding modelโ—Google DeepMind has released EmbeddingGemma 2, an open embedding model that maps text, code, images,... # ai # automatioMmastodonBusiness317 h ago

    Google DeepMind has released EmbeddingGemma 2, an open embedding model that maps text, code, images and other inputs into shared representations, allowing search and retrieval across different data types. The model is designed to run on-device rather than in the cloud, making it free and practical for local applications. Developers and tech commentators are highlighting its multimodal capabilities and its usefulness for search, coding and automation tools.

  10. 10
    Mistral AI launches Mistral Large 4 public previewโ—Mistral AI officially launched the Mistral Large 4 public preview. Discover how this trillion-parameter multimodal modelMmastodonTechnologyAI316 h ago

    Mistral AI has officially launched the public preview of Mistral Large 4, a trillion-parameter multimodal model. The French AI company says the new system is built to rival top closed-source models from competitors, marking a significant step in Europe's push into frontier-scale artificial intelligence. Early reactions in tech communities focus on its scale and open-ecosystem implications.

  11. 11
    AI forces a rethink of what it means to understand imagesโ—What does it mean to understand an image in the age of AI? Visual epistemology shifts attention from simply recognizingMmastodonTechnologyAI01 d ago

    Researchers and commentators are debating how image understanding should be defined now that multimodal AI systems can interpret pictures. The discussion draws on visual epistemology, shifting the focus from what images contain to how seeing, knowledge, belief and interpretation interact to produce meaning, and questioning whether machines genuinely understand images or merely simulate recognition.

  12. 12
    Google launches EmbeddingGemma 2 for on-device multimodal searchโ—Bring multimodal semantic search to the edge with EmbeddingGemma 2โœ‰newsTechnologyAI23 h ago

    Google has announced EmbeddingGemma 2, a model designed to bring multimodal semantic search to edge devices. The release lets developers run embedding-based search across text, images and other media locally, without sending data to the cloud. It signals Google's continued push to move AI capabilities onto phones and other resource-constrained hardware.

  13. 13
    New VA-Bench Shows AI Models Struggle With Robot Tasksโ—Dalian University of Technology's VA-Bench: Top Multimodal Models Finish Only Half of Robot Tasksโœ‰newsTechnologyRobotics2 d ago

    Dalian University of Technology has introduced VA-Bench, a benchmark for evaluating multimodal AI models on robotics tasks. Results show that even the top-performing models complete only about half of the tested tasks, highlighting a significant gap between current multimodal capabilities and practical robot control. The benchmark is drawing attention as a measure of how far AI still is from reliable real-world robotics.

  14. 14
    BostonGene to Present AI Approach in Cancer Drug Developmentโ–ผBostonGene Chief Medical Officer to Discuss Biologically Grounded AI and Multimodal Data in Oncology Drug Development at the 2nd Annual Global Cancer Research & Innovation Symposiumโœ‰newsScienceBiology5 d ago

    BostonGene's Chief Medical Officer is scheduled to speak at the 2nd Annual Global Cancer Research & Innovation Symposium, where the focus will be on biologically grounded artificial intelligence and the use of multimodal data in oncology drug development. The company argues that combining AI with biological context can improve how cancer therapies are discovered and developed.