MikeTrendsTrends right now

Yhn SportFootball first seen 8 h ago, last 2 min ago, peak #2

Running a 125B Qwen model fast on an RTX 4090

Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

A new open-source project, Strata, claims to run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 graphics card at roughly 100 tokens per second. The claim is drawing attention among AI enthusiasts because a model of that size is normally considered far too large for consumer hardware, suggesting new efficiency techniques for local inference.

Why now: Local LLM enthusiasts are excited by the prospect of running a very large model cheaply on a single consumer GPU at high speed.

QwenNvidia RTX 4090StrataGitHub

Open on hn →

Rank over time, top of the chart is #1. 12 snapshots from 1 h ago to 2 min ago.

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/988765