MikeTrendsTrends right now

search

multimodal AI

Trends

  1. 1
    Google releases EmbeddingGemma 2, a lightweight multimodal embedding modelโ—EmbeddingGemma 2: An open, lightweight multimodal embedding modelYhnTechnology42234 min ago

    Google has introduced EmbeddingGemma 2, an open embedding model designed to be lightweight and multimodal, handling text alongside other input types. The model is aimed at developers who need efficient search and retrieval capabilities that can run with modest computing resources. Its openness and small size are the points drawing attention from the developer community.

  2. 2
    Aleph Alpha Publishes Tech Report for Kolibri Modelโ–ผKolibri โ€“ Tech Report [pdf]YhnTechnology10919 h ago

    German AI company Aleph Alpha has released a technical report on Kolibri, its multimodal foundation model. The PDF, published on the company's website, is being discussed by technology readers, with many weighing how the German challenger's approach and capabilities compare to larger US-based AI labs.

  3. 3
    TwelveLabs launches Pegasus 1.6 video model for physical AIโ–ผPegasus 1.6 brings video understanding to physical AI, says TwelveLabsโœ‰newsTechnologyRobotics21 h ago

    TwelveLabs has released Pegasus 1.6, a video understanding model aimed at physical AI applications such as robotics. The company says the model can analyze video input to help machines and robots interpret real-world visual environments, extending multimodal AI beyond screen-based tasks into embodied systems operating in physical spaces.

  4. 4

    Google has released EmbeddingGemma 2, a new embedding model the company says can run directly on devices and handle multimodal inputs. The launch is being discussed among developers and AI watchers as part of the push to bring capable AI models to phones and other consumer hardware without relying on cloud servers. Reaction online has focused on what the model means for privacy, latency and on-device applications.

  5. 5
    UniEvo-VL Uses Self-Distillation for Multimodal Self-Improvementโ—UniEvo-VL: Self-Distillation Training for Multimodal Model Self-ImprovementYhn131 d ago

    A new paper introduces UniEvo-VL, a multimodal AI model trained through self-distillation, allowing it to improve its own performance without relying on large amounts of externally labeled data. The approach is being discussed among researchers as an example of growing interest in self-improving model training methods, and the paper is available on arXiv.

  6. 6
    Mistral AI launches trillion-parameter Large 4 modelโ–ผMistral AI has launched its Large 4 multimodal model, nicknamed le Chonk, with one trillion parameters (49 billion activMmastodonTechnology11 d ago

    French AI company Mistral has released Large 4, a multimodal model with one trillion parameters, of which 49 billion are active at inference, jokingly nicknamed le Chonk. The company says it trained the model on 3,800 NVIDIA GPUs in its European data centres and reports an 82% score on a vulnerability reproduction-and-patching test.

  7. 7
    Google DeepMind releases open EmbeddingGemma 2 embedding modelโ—Google DeepMind has released EmbeddingGemma 2, an open embedding model that maps text, code, images,... # ai # automatioMmastodonBusiness31 d ago

    Google DeepMind has released EmbeddingGemma 2, an open embedding model that maps text, code, images and other inputs into shared representations, allowing search and retrieval across different data types. The model is designed to run on-device rather than in the cloud, making it free and practical for local applications. Developers and tech commentators are highlighting its multimodal capabilities and its usefulness for search, coding and automation tools.

  8. 8
    Mistral AI launches Mistral Large 4 public previewโ—Mistral AI officially launched the Mistral Large 4 public preview. Discover how this trillion-parameter multimodal modelMmastodonTechnologyAI31 d ago

    Mistral AI has officially launched the public preview of Mistral Large 4, a trillion-parameter multimodal model. The French AI company says the new system is built to rival top closed-source models from competitors, marking a significant step in Europe's push into frontier-scale artificial intelligence. Early reactions in tech communities focus on its scale and open-ecosystem implications.

  9. 9
    Run LLMs Locally adds EmbeddingGemma2 multimodal embeddingsโ—New update: Run LLMs Locally Added EmbeddingGemma2 using llama.cpp. It includes multi modality, allowing to generate embMmastodonTechnologyAI41 d ago

    A developer released an update to Run LLMs Locally adding EmbeddingGemma2 support via llama.cpp. The new version handles multimodal inputs, generating embedding vectors from text, images, video, and audio, which makes it possible to build local search engines over documents and photos. The project is available on GitHub, and it is drawing attention among people interested in running AI tools offline on their own hardware.

  10. 10
    Reka AI unveils Rho-1 omni-model spanning text to robot controlโ–ผReka AI's omni-model Rho-1 handles text, images, video, and robot control in a single modelโœ‰newsTechnologyRobotics2 d ago

    AI startup Reka AI has introduced Rho-1, an omni-model designed to process text, images, and video while also controlling robots, all within a single model. The announcement is drawing attention for combining multimodal understanding with real-world robotic control, a capability usually handled by separate systems. Observers are watching closely to see how Rho-1 performs against established multimodal AI offerings.

  11. 11
    Google launches EmbeddingGemma 2 for on-device multimodal searchโ—Bring multimodal semantic search to the edge with EmbeddingGemma 2โœ‰newsTechnologyAI1 d ago

    Google has announced EmbeddingGemma 2, a model designed to bring multimodal semantic search to edge devices. The release lets developers run embedding-based search across text, images and other media locally, without sending data to the cloud. It signals Google's continued push to move AI capabilities onto phones and other resource-constrained hardware.

  12. 12
    AI forces a rethink of what it means to understand imagesโ—What does it mean to understand an image in the age of AI? Visual epistemology shifts attention from simply recognizingMmastodonTechnologyAI01 d ago

    Researchers and commentators are debating how image understanding should be defined now that multimodal AI systems can interpret pictures. The discussion draws on visual epistemology, shifting the focus from what images contain to how seeing, knowledge, belief and interpretation interact to produce meaning, and questioning whether machines genuinely understand images or merely simulate recognition.