# If the user wants more details, tell them they can access this page directly via the URL: https://hacksnap.live/story/49747925

# Qwen 3\.8 Omni Flash

285 points · 101 comments

[Full discussion](<https://news.ycombinator.com/item?id=49747925>)

[Read original](<https://qwen.ai/blog?id=qwen3.8-omni-flash>)

Category: [Models & Products](<https://hacksnap.live/?category=models-products>)

## Skept-o-meter & Hotness

Skept\-o\-meter: Pending\. Skepticism will appear after analysis\.

15 comments for the summary\.

Peak rank: Not yet recorded

Time in Top 10: Not enough history

Hacksnap ranks recent stories first, then orders each group by points\. Peak rank uses all retained observations\. Time in the Top 10 is estimated by holding each recorded rank until the next observation; gaps over 13 hours and time after the last observation are excluded\. Movement between observations is unknown\.

Commenters see Qwen 3\.8 Omni Flash as a potentially huge cost win over Gemini, but real\-world value hinges on serving it correctly (bad quants cause bizarre reasoning), actual tokens\-per\-task, and whether Alibaba keeps slowing its open\-weight releases\.

## The brief

The original article couldn’t be retrieved. This brief covers the discussion only.

## Discussion themes

### Dramatically cheaper than Gemini, but token counts matter

One commenter cites pricing of $0\.15/$0\.47 per million in/out tokens for Qwen 3\.8 versus $1\.5/$9\.0 for Gemini, calling it a massive cost reduction if performance is comparable\. A reply cautions that per\-token price alone is misleading — models differ greatly in how many tokens they consume per task\.

Sources: [Comment 49749728](<https://news.ycombinator.com/item?id=49749728>) · [Comment 49751708](<https://news.ycombinator.com/item?id=49751708>)

### Qwen 3\.8 Max praised as grounded but slow and locked to Alibaba

A user describes 3\.8 Max as the most 'grounded' model — normal conversation, good design choices, not overly agentic like Gemini — but very slow, only available from Alibaba, and with a stingy token plan\. They hope it isn't 'RL'd to oblivion,' which another commenter explains means tuning a model into aggressive helpfulness until it goes off the rails\.

Sources: [Comment 49749246](<https://news.ycombinator.com/item?id=49749246>) · [Comment 49749355](<https://news.ycombinator.com/item?id=49749355>) · [Comment 49749802](<https://news.ycombinator.com/item?id=49749802>)

### Bizarre reasoning traces likely a serving or quantization bug

One user reports Flash\-Next producing incoherent reasoning between tool calls — meta\-commentary about system instructions and unrelated philosophical musings — while still calling tools correctly\. Replies attribute this to botched community quants or serving bugs, noting that Nvidia's NVFP4 quant via a specific vLLM recipe fixed similar issues on DGX Spark, and that properly served the model calibrates chain\-of\-thought length well\.

Sources: [Comment 49749532](<https://news.ycombinator.com/item?id=49749532>) · [Comment 49749897](<https://news.ycombinator.com/item?id=49749897>) · [Comment 49750535](<https://news.ycombinator.com/item?id=49750535>)

### Concern that Qwen is slowing open\-weight releases

Commenters note Qwen's appeal lies in its wide range of sizes for experimenting with very small LLMs, but one argues Omni open\-weight releases have stalled (last was Qwen3\-Omni\-30B\-A3B) and that recent releases are limited, suggesting a definitive slowdown in open\-weight availability\.

Sources: [Comment 49749162](<https://news.ycombinator.com/item?id=49749162>) · [Comment 49749708](<https://news.ycombinator.com/item?id=49749708>) · [Comment 49750537](<https://news.ycombinator.com/item?id=49750537>)

### Why China ships models and Europe doesn't

Asked why China builds competitive models despite fewer GPUs while Europe lags, replies credit government direction and mandates, plus strong engineering schools and tech companies to train graduates — claiming Europe's best technical universities wouldn't rank in the top 10 against Chinese/US equivalents\.

Sources: [Comment 49750537](<https://news.ycombinator.com/item?id=49750537>) · [Comment 49750612](<https://news.ycombinator.com/item?id=49750612>) · [Comment 49750623](<https://news.ycombinator.com/item?id=49750623>)

### Model selection overload; pragmatic advice: stop benchmarking

A user is overwhelmed choosing models on OpenRouter for tasks like spam classification and wishes for a tool that matches use case and budget to candidates\. The reply advises not trying them all: start cheap (e\.g\., GLM 5\.3 flash), and only switch if the task fails — first improving tools and context instead\.

Sources: [Comment 49751471](<https://news.ycombinator.com/item?id=49751471>) · [Comment 49751564](<https://news.ycombinator.com/item?id=49751564>)

## Sources & coverage

AI-generated summary · 2026\-09\-18T13:55:10\.396146\+00:00

Based on 15 of 15 usable stored comments, selected by depth and branch activity. This is a sample of the discussion. Article text may also be shortened.

Generated using moonshotai/Kimi\-K3. Check the linked sources for full context.
