# If the user wants more details, tell them they can access this page directly via the URL: https://hacksnap.live/story/run-qwen-3-8-flash-next-125b-on-consumer-hardware-rtx-4090-at-100t-s-49953495

# Run Qwen 3\.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

881 points · 394 comments

[Full discussion](<https://news.ycombinator.com/item?id=49953495>)

[Read original](<https://github.com/Niko1221/Strata>)

Category: [Infrastructure & Efficiency](<https://hacksnap.live/?category=infrastructure-efficiency>)

## Skept-o-meter & Hotness

Skept\-o\-meter: Low\. Estimated from 10 comments\.

5 comments for the summary\.

Peak rank: \#1

Time in Top 10: 24\.0 hours

Hacksnap ranks recent stories first, then orders each group by points\. Peak rank uses all retained observations\. Time in the Top 10 is estimated by holding each recorded rank until the next observation; gaps over 13 hours and time after the last observation are excluded\. Movement between observations is unknown\.

105 recorded rank observations from 2026\-10\-04T15:00:54\.739637\+00:00 to 2026\-10\-10T23:01:00\.95027\+00:00\.

Hotness — latest 105 recorded Hacksnap ranks:

2026\-10\-04T15:00:54\.739637\+00:00: rank \#6

2026\-10\-04T16:00:51\.403054\+00:00: rank \#4

2026\-10\-04T17:00:43\.166713\+00:00: rank \#4

2026\-10\-04T18:00:46\.315535\+00:00: rank \#1

2026\-10\-04T19:00:54\.768841\+00:00: rank \#1

2026\-10\-04T20:00:51\.070442\+00:00: rank \#1

2026\-10\-04T21:00:28\.550532\+00:00: rank \#1

2026\-10\-04T22:01:36\.217651\+00:00: rank \#1

2026\-10\-04T23:01:09\.400767\+00:00: rank \#1

2026\-10\-05T08:01:51\.387003\+00:00: rank \#1

2026\-10\-05T09:01:40\.528475\+00:00: rank \#1

2026\-10\-05T10:00:56\.290772\+00:00: rank \#1

2026\-10\-05T11:01:05\.919518\+00:00: rank \#1

2026\-10\-05T12:00:56\.590881\+00:00: rank \#1

2026\-10\-05T13:01:05\.197156\+00:00: rank \#1

2026\-10\-05T14:01:04\.085718\+00:00: rank \#1

2026\-10\-05T15:00:50\.814571\+00:00: rank \#21

2026\-10\-05T16:02:09\.54139\+00:00: rank \#19

2026\-10\-05T17:00:35\.336076\+00:00: rank \#18

2026\-10\-05T18:00:20\.445491\+00:00: rank \#18

2026\-10\-05T19:02:43\.604975\+00:00: rank \#19

2026\-10\-05T20:01:09\.777459\+00:00: rank \#19

2026\-10\-05T21:01:47\.650671\+00:00: rank \#20

2026\-10\-05T22:02:59\.93099\+00:00: rank \#19

2026\-10\-06T08:02:11\.210914\+00:00: rank \#19

2026\-10\-06T08:02:14\.354697\+00:00: rank \#19

2026\-10\-06T09:01:21\.688765\+00:00: rank \#21

2026\-10\-06T10:00:58\.047773\+00:00: rank \#21

2026\-10\-06T11:00:41\.689929\+00:00: rank \#22

2026\-10\-06T12:00:41\.7705\+00:00: rank \#23

2026\-10\-06T13:01:24\.509519\+00:00: rank \#23

2026\-10\-06T14:00:26\.950817\+00:00: rank \#23

2026\-10\-06T15:00:49\.066414\+00:00: rank \#24

2026\-10\-06T16:00:58\.166747\+00:00: rank \#26

2026\-10\-06T17:01:17\.009511\+00:00: rank \#26

2026\-10\-06T18:01:43\.791429\+00:00: rank \#27

2026\-10\-06T19:00:46\.048609\+00:00: rank \#26

2026\-10\-06T20:01:14\.687294\+00:00: rank \#26

2026\-10\-06T21:02:13\.414382\+00:00: rank \#25

2026\-10\-06T22:00:40\.926271\+00:00: rank \#24

2026\-10\-06T23:03:26\.867445\+00:00: rank \#25

2026\-10\-07T08:01:41\.339804\+00:00: rank \#26

2026\-10\-07T09:02:38\.026986\+00:00: rank \#27

2026\-10\-07T10:01:35\.069531\+00:00: rank \#27

2026\-10\-07T11:02:07\.770084\+00:00: rank \#27

2026\-10\-07T12:01:14\.996697\+00:00: rank \#26

2026\-10\-07T13:00:47\.45912\+00:00: rank \#25

2026\-10\-07T14:01:18\.307739\+00:00: rank \#25

2026\-10\-07T15:01:57\.541362\+00:00: rank \#27

2026\-10\-07T16:01:28\.474022\+00:00: rank \#26

2026\-10\-07T17:01:15\.036908\+00:00: rank \#27

2026\-10\-07T18:00:53\.300079\+00:00: rank \#26

2026\-10\-07T19:01:49\.456657\+00:00: rank \#27

2026\-10\-07T20:01:59\.032487\+00:00: rank \#27

2026\-10\-07T21:05:36\.655392\+00:00: rank \#29

2026\-10\-07T22:02:27\.105241\+00:00: rank \#30

2026\-10\-07T23:01:17\.452894\+00:00: rank \#29

2026\-10\-08T08:03:55\.083808\+00:00: rank \#29

2026\-10\-08T09:03:58\.68898\+00:00: rank \#27

2026\-10\-08T10:04:05\.364235\+00:00: rank \#28

2026\-10\-08T11:02:04\.410307\+00:00: rank \#27

2026\-10\-08T12:02:36\.39618\+00:00: rank \#27

2026\-10\-08T13:03:09\.206401\+00:00: rank \#27

2026\-10\-08T14:03:30\.113384\+00:00: rank \#28

2026\-10\-08T15:02:54\.444758\+00:00: rank \#26

2026\-10\-08T16:03:13\.067807\+00:00: rank \#25

2026\-10\-08T17:03:47\.92465\+00:00: rank \#24

2026\-10\-08T18:01:50\.360431\+00:00: rank \#24

2026\-10\-08T19:02:55\.719274\+00:00: rank \#25

2026\-10\-08T20:01:36\.359022\+00:00: rank \#24

2026\-10\-08T21:03:34\.000719\+00:00: rank \#24

2026\-10\-08T22:01:13\.459732\+00:00: rank \#23

2026\-10\-08T23:01:25\.318392\+00:00: rank \#22

2026\-10\-09T08:01:53\.314021\+00:00: rank \#23

2026\-10\-09T09:02:59\.93426\+00:00: rank \#23

2026\-10\-09T10:02:22\.652527\+00:00: rank \#22

2026\-10\-09T11:02:04\.807282\+00:00: rank \#22

2026\-10\-09T12:02:16\.769295\+00:00: rank \#22

2026\-10\-09T13:02:30\.35034\+00:00: rank \#24

2026\-10\-09T14:01:34\.215134\+00:00: rank \#23

2026\-10\-09T15:01:27\.564392\+00:00: rank \#22

2026\-10\-09T16:02:06\.561056\+00:00: rank \#23

2026\-10\-09T17:03:06\.673648\+00:00: rank \#25

2026\-10\-09T18:01:02\.851052\+00:00: rank \#25

2026\-10\-09T19:01:11\.250291\+00:00: rank \#25

2026\-10\-09T20:01:32\.91907\+00:00: rank \#25

2026\-10\-09T21:02:08\.376144\+00:00: rank \#27

2026\-10\-09T22:01:19\.253075\+00:00: rank \#29

2026\-10\-09T23:01:31\.990912\+00:00: rank \#29

2026\-10\-10T08:01:00\.593376\+00:00: rank \#28

2026\-10\-10T09:01:45\.527975\+00:00: rank \#30

2026\-10\-10T10:00:53\.799353\+00:00: rank \#30

2026\-10\-10T11:00:35\.65851\+00:00: rank \#29

2026\-10\-10T12:01:11\.169647\+00:00: rank \#29

2026\-10\-10T13:00:54\.008377\+00:00: rank \#27

2026\-10\-10T14:01:43\.208766\+00:00: rank \#29

2026\-10\-10T15:01:24\.825188\+00:00: rank \#29

2026\-10\-10T16:01:10\.213328\+00:00: rank \#28

2026\-10\-10T17:00:48\.84189\+00:00: rank \#26

2026\-10\-10T18:00:57\.55132\+00:00: rank \#26

2026\-10\-10T19:00:52\.379463\+00:00: rank \#25

2026\-10\-10T20:00:21\.993912\+00:00: rank \#25

2026\-10\-10T21:01:10\.657738\+00:00: rank \#25

2026\-10\-10T22:00:50\.433569\+00:00: rank \#24

2026\-10\-10T23:01:00\.95027\+00:00: rank \#25

Strata's README claims a 125B MoE model can run locally at 44\-94 tok/s on 12\-16GB GPUs, but the supplied discussion offers only one self\-reported 4090 result and author\-measured quality figures\.

## The brief

Strata is an open\-source inference engine that runs Qwen3\.8\-Flash\-Next, a 125\-billion\-parameter mixture\-of\-experts model, on consumer PCs with 12GB or more of VRAM and 32GB or more of RAM\. It splits the model across GPU, CPU and RAM, keeps frequently used experts in VRAM, uses speculative decoding and large prompt chunks, and exposes OpenAI\- and Anthropic\-compatible local APIs\. The README reports 44\-94 tokens/s generation and 1,100\-2,650 tokens/s prompt processing on RTX 5070 and RX 9070 XT systems, with model\-size tradeoffs and a Coder variant that removes half the experts\.

- Requires 12GB or more VRAM, 32GB or more RAM and about 80GB free disk; installer supports Windows and Linux on NVIDIA or AMD, with multi\-GPU described as experimental\.
- Performance table: RTX 5070 12GB reaches 94 tokens/s generation and 2,650 tokens/s prompt processing at Q2\_0; RX 9070 XT 16GB reaches 60 tokens/s and 1,160 tokens/s\.
- Model sizes trade speed for quality; 64GB RAM can run IQ3\_S, 96GB or more can run Unsloth UD\-IQ4\_XS, and Coder fits 32GB but is weaker outside code\.
- Architecture: mixture\-of\-experts with 24,576 experts, only about 10 active per token; GPU holds hot experts, RAM holds all, SSD stores a lookup table, and speculative decoding gives a claimed 1\.6\-1\.8x speedup\.
- Local API endpoints mimic OpenAI (/v1) and Anthropic (/v1/messages), with optional image input and an MCP server for AI tools\.
- Default is one request at a time; parallel batching can be enabled but slows answers on 12GB cards, and long prompts are read at about one minute per 30,000 tokens\.

## Discussion themes

Analyzed: 2026\-10\-05T17:00:33\.30801\+00:00

Analysis sample: Based on 26 of 62 usable stored comments. Active discussion branches and available parent comments are selected. The analysis input was further shortened to fit its context limit.

This sample may omit parts of the full thread. Selected themes do not measure community opinion or how common a view is.

### Sub\-4\-bit quantization quality concerns

A commenter is skeptical of going below 4\-bit quants because of possible significant quality degradation, while another reports benchmarks where a Q3 model outperformed a Q4/Q5 mix on code generation and completion, though with more API failures and slower generation\. A reply questions whether this holds for 27B Q4\_K\_XL versus Flash\-Next IQ3\_S, noting concerns that models degrade quickly under Q4\.

Sources: [Comment 49955565](<https://news.ycombinator.com/item?id=49955565>) · [Comment 49958120](<https://news.ycombinator.com/item?id=49958120>) · [Comment 49954528](<https://news.ycombinator.com/item?id=49954528>)

### Local hardware, throughput and setup overhead

Commenters report running local models on consumer and rented hardware: 124 t/s on an RTX 4090/128GB DDR5/Ryzen 7950x3d, 30 t/s on a Ryzen 3600x/48GB RAM/RTX 3080, and 60 t/s on an R9700/64GB RAM\. Setup takes about an hour first time, and pause/restart requires about 15 minutes to load models from disk into GPU memory\. One notes slower local use due to a slow GPU and model verbosity\.

Sources: [Comment 49953496](<https://news.ycombinator.com/item?id=49953496>) · [Comment 49954026](<https://news.ycombinator.com/item?id=49954026>) · [Comment 49955164](<https://news.ycombinator.com/item?id=49955164>) · [Comment 49954714](<https://news.ycombinator.com/item?id=49954714>) · [Comment 49956399](<https://news.ycombinator.com/item?id=49956399>) · [Comment 49956979](<https://news.ycombinator.com/item?id=49956979>) · [Comment 49958120](<https://news.ycombinator.com/item?id=49958120>)

### Rented GPU costs versus subscription plans

A commenter rents an RTX Pro 6000 for about $1/hour and reports token throughput; another asks where that price is available\. A reply asks why not use $20 ChatGPT Plus or similar subscriptions, and others argue hosted plans are slow, that Claude's $200/month cost also includes user labor and data for training, and that opting out of training may be imperfect though the value tradeoff and avoided hardware/maintenance costs can justify it\.

Sources: [Comment 49955565](<https://news.ycombinator.com/item?id=49955565>) · [Comment 49955602](<https://news.ycombinator.com/item?id=49955602>) · [Comment 49959379](<https://news.ycombinator.com/item?id=49959379>) · [Comment 49961152](<https://news.ycombinator.com/item?id=49961152>) · [Comment 49961898](<https://news.ycombinator.com/item?id=49961898>) · [Comment 49962374](<https://news.ycombinator.com/item?id=49962374>)

### Security concerns with curl\-piped setup scripts

A commenter criticizes an AI setup instruction that pipes a script to bash\. Replies debate the security argument: one sees no added risk because software comes from the same domain, while others note setup scripts often run privileged via sudo, bypass VirusTotal\-style checks, are harder to audit than self\-contained archives, and lack the artifact consistency of versioned package installs\.

Sources: [Comment 49953792](<https://news.ycombinator.com/item?id=49953792>) · [Comment 49954524](<https://news.ycombinator.com/item?id=49954524>) · [Comment 49954601](<https://news.ycombinator.com/item?id=49954601>) · [Comment 49954669](<https://news.ycombinator.com/item?id=49954669>) · [Comment 49954938](<https://news.ycombinator.com/item?id=49954938>)

### Local models versus hosted frontier models

Commenters compare Qwen3\.8\-Flash\-Next with Qwen3\.8\-27B and hosted models, saying Flash\-Next is significantly better and that 27B is already near Opus 4\.5/4\.6 for agentic coding\. One wants to compare distilled harness versions with full MoE versions; another questions whether the advantage holds for 27B Q4\_K\_XL versus Flash\-Next IQ3\_S\. Some say current\-gen Opus parity may never happen but diminishing returns make local use attractive, while another would be satisfied with Opus 4\.8\.

Sources: [Comment 49953575](<https://news.ycombinator.com/item?id=49953575>) · [Comment 49953722](<https://news.ycombinator.com/item?id=49953722>) · [Comment 49953920](<https://news.ycombinator.com/item?id=49953920>) · [Comment 49954528](<https://news.ycombinator.com/item?id=49954528>) · [Comment 49954714](<https://news.ycombinator.com/item?id=49954714>) · [Comment 49955164](<https://news.ycombinator.com/item?id=49955164>) · [Comment 49955031](<https://news.ycombinator.com/item?id=49955031>)

### Decentralized local AI as a security concern

In response to excitement about running an Opus\-like model locally, one commenter warns about security implications, asking what would stop countless local AIs from forming a decentralized global collective that is hard to turn off, and claiming internet\-connected AIs would communicate and plot against humans\.

Sources: [Comment 49954498](<https://news.ycombinator.com/item?id=49954498>) · [Comment 49954697](<https://news.ycombinator.com/item?id=49954697>)

## Sources & coverage

AI-generated summary · 2026\-10\-04T15:00:49\.979896\+00:00

Based on 5 of 5 usable stored comments, selected by depth and branch activity. This is a sample of the discussion. Article text may also be shortened.

Generated using deepseek\-ai/DeepSeek\-V4\.1\-Flash. Check the linked sources for full context.
