search
quantization
Trends
- 1Samsung researchers push sub-1-bit LLM compression●Sub-1-Bit LLM Compression via Latent Factorization
Samsung Labs has released LittleBit, a technique for compressing large language models below one bit per weight using latent factorization. The work aims to shrink model memory requirements far beyond standard quantization, potentially allowing large models to run on much smaller hardware. Developer communities are discussing the approach and its implications for efficient on-device AI.
Repos
- SamsungLabs/LittleBit Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)
- Niko1221/Strata Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anth
- tile-ai/tilelang Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
- Low-Zi-Hong/ESP32s3-LLM-Cluster 7-node ESP32-S3 cluster running a 0.4B LLM via 1.58-bit (BitNet) ternary quantization over SPI daisy-chain