MikeTrendsTrends right now

Yhn SportFootball first seen 4 h ago, last 5 min ago, peak #4

Qwen 3.8 Flash Next 125B claimed to run fast on RTX 4090

Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

A project called Strata, shared on GitHub, claims it can run Qwen's 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. If verified, that would make a very large language model practical on high-end home hardware without a data center. The claim is drawing attention among developers interested in local AI inference, though independent confirmation of the speed figures has not been established.

Why now: Running a 125B-parameter model quickly on one consumer GPU would be a notable breakthrough for local AI inference, so people want to verify it

QwenStrataRTX 4090NVIDIA

Open on hn →

Rank over time, top of the chart is #1. 23 snapshots from 4 h ago to 5 min ago.

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/988765