MikeTrendsTrends right now

search

quantization

Trends

  1. 1
    180B-Parameter AI Model Claims to Run on 8GB Laptop GPUโ—VIDRAFT์˜ POCKET-Darwin-180B-GGUF๋กœ 180B ํŒŒ๋ผ๋ฏธํ„ฐ MoE ๋ชจ๋ธ์„ 8GB VRAM ๋…ธํŠธ๋ถ์—์„œ ๊ตฌ๋™. 4๋น„ํŠธ ์–‘์žํ™”์™€ llama.cpp๋กœ 250๋ฐฐ ํ•˜๋“œ์›จ์–ด ์š”๊ตฌ์‚ฌํ•ญ ์ ˆ๊ฐ, ๋กœ์ปฌ AI ๋ฐฐํฌ ์™„MmastodonTechnologySoftware321 h ago

    VIDRAFT has released POCKET-Darwin-180B-GGUF, a quantized build that reportedly runs a 180-billion-parameter mixture-of-experts model on a laptop with just 8GB of VRAM. Using 4-bit quantization and llama.cpp, the project claims hardware requirements are cut by roughly 250 times, making local deployment of very large AI models feasible on consumer machines.

  2. 2
    Laser Gives Scientists Precise Nanoscale Control of Phononsโ–ผLaser Provides Precise Control over Phonons at Nanoscale Levelโœ‰newsSciencePhysics1 d ago

    Researchers report using laser light to precisely control phonons โ€” quantized vibrations in materials โ€” at the nanoscale. The advance could enable faster, smaller devices by manipulating heat and sound-like vibrations directly within chips and other nanostructures. It marks a step forward in phononics, a field seeking to use vibrations for information processing and thermal management alongside conventional electronics.

  3. 3
    The eternal question: will this AI model fit on my GPU?โ—Have you ever found yourself stuck in this question: "will this model fit on my GPU?". A lot of us have. The honest answMmastodonTechnologyAI21 d ago

    A discussion is circulating among AI hobbyists and developers about whether a given machine-learning model will fit on a specific graphics card. The oft-repeated answer โ€” that it depends on quantization and context length โ€” is accurate but unhelpful for people who just want a quick yes or no before downloading a model. The exchange reflects a common frustration in the local AI community.

  4. 4
    180B-parameter LLM runs locally on a laptop without a GPUโ—GPU ์—†์ด ์†Œ๋น„์ž์šฉ ๋…ธํŠธ๋ถ์—์„œ 180์–ต ํŒŒ๋ผ๋ฏธํ„ฐ LLM์„ ๊ตฌ๋™ํ•˜๋Š” POCKET-Darwin-180B. 4๋น„ํŠธ GGUF ์–‘์žํ™”๋กœ 360GBโ†’111GB ์••์ถ•, ์•ฝ $1,400 ํ•˜๋“œ์›จ์–ด๋กœ ๋กœ์ปฌ ์ถ”๋ก  ๊ฐ€๋Šฅ. # ai #MmastodonTechnologyAI34 d ago

    A project called POCKET-Darwin-180B is drawing attention for running a 180-billion-parameter language model on consumer hardware with no discrete GPU. Using 4-bit GGUF quantization, the model is compressed from roughly 360GB down to 111GB, enabling local inference on hardware costing about $1,400. Commenters in AI and open-source circles are highlighting it as a sign that frontier-scale models may soon run off the cloud.

  5. 5

    Hobbyists and independent developers are sharing methods to make AI models run significantly faster on consumer-grade computers, without specialized data-center equipment. The discussion centers on optimization tricks such as quantization, caching and smarter memory use that let large language models run on ordinary laptops and desktops. Commenters are trading benchmarks and configuration tips, with many arguing that capable local AI no longer needs expensive hardware.

  6. 6
    Quantized 27B Model Claimed to Match Frontier AI on Coding Taskโ—A 27B Quantized LLM Is Said To Match Frontier AI Models In Just One Task From A Coding Benchmark, Making It A More Believable Claimโœ‰newsTechnologyAI5 d ago

    A quantized 27-billion-parameter language model is reported to match frontier AI models on a single task from a coding benchmark. The narrow, specific nature of the claim makes it more believable than sweeping benchmark-superiority claims, but it also means the result says little about overall performance. Readers are debating how much weight such partial benchmark results deserve in judging open and smaller models.

  7. 7
    AllenAI releases open-source AstaBrief 8B for automated science reportsโ—ํ•ต์‹ฌ ๋‚ด์šฉ AllenAI๊ฐ€ ๊ณต๊ฐœํ•œ AstaBrief 8B ๋Š” ๊ณผํ•™ ๋…ผ๋ฌธ์„ ์ธ์šฉ ๊ธฐ๋ฐ˜ ๋ณด๊ณ ์„œ๋กœ ์ž๋™ ์ƒ์„ฑํ•ด ์ฃผ๋Š” ๋ชจ๋ธ์ด๋‹ค. ๊ธฐ์กด Claude ๊ธฐ๋ฐ˜ โ€˜Thinking modeโ€™๋ณด๋‹ค 3.5๋ฐฐ ๋น ๋ฅธ 51์ดˆ ์•ˆ์— ์ „์ฒด ๋ณด๊ณ ์„œMmastodonTechnologySoftware34 d ago

    AllenAI has released AstaBrief 8B, an open model that automatically generates citation-based reports from scientific papers. It produces a full report in about 51 seconds, roughly 3.5 times faster than Claude's Thinking mode. Both model weights and training data are open, allowing local execution; developers note an 8B quantized version can run on a low-budget VPS with 4-6 GB of memory.

  8. 8
    Tether pushes 13-billion parameter BitNet b1.58 model to the edgeโ—Tether is pushing the 13-billion parameter BitNet b1.58 LLM to the edge.โœ‰newsTechnologyAI6 d ago

    Tether, the company behind the USDT stablecoin, is developing BitNet b1.58, a 13-billion parameter large language model built on 1.58-bit quantization designed to run efficiently on edge devices with limited hardware. The move signals Tether's expansion beyond crypto into artificial intelligence, drawing attention for its unconventional low-precision approach to AI inference.

Repos