Mmastodon TechnologyTechnology first seen 14 h ago, last 14 h ago, peak #1
Project claims 125B model runs fast on RTX 4090
Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s Article URL: https:// github.com/Niko1221/Strat
A GitHub project called Strata claims it can run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. The claim is drawing attention from developers on Hacker News, where it has collected 15 points and a handful of comments, with discussion focused on whether such speed and memory efficiency on consumer hardware is realistic.
Why now: Running a 125B model at 100 tokens per second on a single consumer GPU would be a notable efficiency breakthrough, prompting skepticism and interest.
Qwen 3.8 Flash NextStrataNiko1221Nvidia RTX 4090Hacker News
Rank over time, top of the chart is #1. 5 snapshots from 14 h ago to 14 h ago.
Evidence
API: https://socialmediatrends-api.osmike.com/v1/trends/1022056