search
RTX 4090
Trends
- 1Running Qwen 3.8 Flash Next 125B fast on an RTX 4090●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A newly shared open-source project, Strata, claims to run Alibaba's Qwen 3.8 Flash Next, a 125-billion-parameter model, on a single consumer RTX 4090 graphics card at speeds around 100 tokens per second. If the results hold up, it would make large-model inference far cheaper for hobbyists and small teams, and developers on Hacker News are actively weighing in on the performance claims.
- 2Project claims 125B model runs fast on RTX 4090●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s Article URL: https:// github.com/Niko1221/Strat
A GitHub project called Strata claims it can run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. The claim is drawing attention from developers on Hacker News, where it has collected 15 points and a handful of comments, with discussion focused on whether such speed and memory efficiency on consumer hardware is realistic.
- 3Qwen 3.8 Flash Next Runs at 100 T/s on One RTX 4090●Qwen 3.8 Flash Next on a Single RTX 4090: How Consumer‑Grade GPUs Reach 100 T/s By Senior Editor – October 2026 “A singl
Reports circulating in tech circles claim that Alibaba's Qwen 3.8 Flash Next, a 125-billion-parameter model, can run at roughly 100 trillion tokens per second on a single consumer RTX 4090 GPU — a throughput previously associated with multi-node H100 clusters. Enthusiasts are discussing what this means for local AI inference and the collapsing cost barrier between consumer and data-center hardware.