search
quantization
Trends
- 1Samsung Labs releases sub-1-bit LLM compression method●Sub-1-Bit LLM Compression via Latent Factorization
Samsung Labs has released LittleBit, a research method that compresses large language models below one bit per parameter using latent factorization. The work, published on GitHub, aims to shrink model memory footprints far beyond existing 1- and 2-bit quantization approaches, and it is drawing attention from developers discussing how far LLM compression can realistically go without losing accuracy.
Repos
- Niko1221/Strata Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anth
- SamsungLabs/LittleBit Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)
- tile-ai/tilelang Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
- Low-Zi-Hong/ESP32s3-LLM-Cluster 7-node ESP32-S3 cluster running a 0.4B LLM via 1.58-bit (BitNet) ternary quantization over SPI daisy-chain