# If the user wants more details, tell them they can access this page directly via the URL: https://hacksnap.live/story/whistle-speech-to-text-in-16-9-mb-50008427

# Whistle: Speech to Text in 16\.9 MB

908 points · 176 comments

[Full discussion](<https://news.ycombinator.com/item?id=50008427>)

[Read original](<https://cactuscompute.com/blog/whistle>)

Category: [Infrastructure & Efficiency](<https://hacksnap.live/?category=infrastructure-efficiency>)

## Skept-o-meter & Hotness

Skept\-o\-meter: Low\. Estimated from 10 comments\.

8 comments for the summary\.

Peak rank: \#1

Time in Top 10: 24\.0 hours

Hacksnap ranks recent stories first, then orders each group by points\. Peak rank uses all retained observations\. Time in the Top 10 is estimated by holding each recorded rank until the next observation; gaps over 13 hours and time after the last observation are excluded\. Movement between observations is unknown\.

37 recorded rank observations from 2026\-10\-08T19:02:55\.719274\+00:00 to 2026\-10\-10T23:01:00\.95027\+00:00\.

Hotness — latest 37 recorded Hacksnap ranks:

2026\-10\-08T19:02:55\.719274\+00:00: rank \#9

2026\-10\-08T20:01:36\.359022\+00:00: rank \#8

2026\-10\-08T21:03:34\.000719\+00:00: rank \#5

2026\-10\-08T22:01:13\.459732\+00:00: rank \#2

2026\-10\-08T23:01:25\.318392\+00:00: rank \#2

2026\-10\-09T08:01:53\.314021\+00:00: rank \#2

2026\-10\-09T09:02:59\.93426\+00:00: rank \#1

2026\-10\-09T10:02:22\.652527\+00:00: rank \#1

2026\-10\-09T11:02:04\.807282\+00:00: rank \#2

2026\-10\-09T12:02:16\.769295\+00:00: rank \#2

2026\-10\-09T13:02:30\.35034\+00:00: rank \#2

2026\-10\-09T14:01:34\.215134\+00:00: rank \#2

2026\-10\-09T15:01:27\.564392\+00:00: rank \#2

2026\-10\-09T16:02:06\.561056\+00:00: rank \#2

2026\-10\-09T17:03:06\.673648\+00:00: rank \#2

2026\-10\-09T18:01:02\.851052\+00:00: rank \#2

2026\-10\-09T19:01:11\.250291\+00:00: rank \#24

2026\-10\-09T20:01:32\.91907\+00:00: rank \#24

2026\-10\-09T21:02:08\.376144\+00:00: rank \#26

2026\-10\-09T22:01:19\.253075\+00:00: rank \#28

2026\-10\-09T23:01:31\.990912\+00:00: rank \#28

2026\-10\-10T08:01:00\.593376\+00:00: rank \#27

2026\-10\-10T09:01:45\.527975\+00:00: rank \#29

2026\-10\-10T10:00:53\.799353\+00:00: rank \#29

2026\-10\-10T11:00:35\.65851\+00:00: rank \#28

2026\-10\-10T12:01:11\.169647\+00:00: rank \#28

2026\-10\-10T13:00:54\.008377\+00:00: rank \#26

2026\-10\-10T14:01:43\.208766\+00:00: rank \#28

2026\-10\-10T15:01:24\.825188\+00:00: rank \#28

2026\-10\-10T16:01:10\.213328\+00:00: rank \#27

2026\-10\-10T17:00:48\.84189\+00:00: rank \#25

2026\-10\-10T18:00:57\.55132\+00:00: rank \#25

2026\-10\-10T19:00:52\.379463\+00:00: rank \#24

2026\-10\-10T20:00:21\.993912\+00:00: rank \#24

2026\-10\-10T21:01:10\.657738\+00:00: rank \#24

2026\-10\-10T22:00:50\.433569\+00:00: rank \#23

2026\-10\-10T23:01:00\.95027\+00:00: rank \#24

Whistle's 16\.9 MB on\-device speech model draws interest, but the supplied discussion questions whether size addresses atypical speech, dictation cleanup, and uneven language quality; the sample is small\.

## The brief

Cactus Compute has released Whistle, a 16\.9 MB open speech recognition model that runs on CPU with no dependencies and shares the C\+\+ engine and quantization used by its Needle text model\. It transcribes seven languages, emits word timestamps and speech embeddings, and can be loaded alongside Needle so one binary turns audio into tool calls\. The source reports benchmarks where Whistle beats Whisper base and Moonshine tiny v2 on size, time to first token, decode speed, and several word\-error\-rate sets, while Whisper base remains ahead on TED\-LIUM, AMI, and MLS; these are the authors' reported results, not independent verification\.

- Architecture: log\-mel front end at 16 kHz, convolutional stem, eight Simple Attention encoder blocks, eight Laddered Simple Attention decoder blocks with gated cross\-attention; encoder always runs all eight, decoder depth selectable\.
- Decoding: five\-beam search, length\-normalized log probability, Aho\-Corasick keyword biasing, 320\-token cap, 8,192 text pieces plus seven language tokens\.
- Benchmarks: 16\.9 MB vs Whisper base 145\.3 MB and Moonshine tiny v2 41\.9 MB; 11\.1 ms TTFT and 1,319 tok/s on Apple M4 Pro CPU; WER ahead on LibriSpeech, SPGISpeech, Earnings\-22, FLEURS average; behind on TED\-LIUM, AMI, MLS\.
- Deployment: prebuilt for 17 targets including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC\-V, MIPS, browser, WASI; C API needle\_load, needle\_transcribe, needle\_embed\.
- Integration: needle\_complete can transcribe audio and answer against tools, returning JSON with function\_calls and audio\_ fields\.
- Test hygiene: no test audio in training or validation, verified by audio checksums and speaker IDs; WER scored with Whisper normalizers over 86,174 utterances\.

## Discussion themes

Analyzed: 2026\-10\-09T12:01:58\.179279\+00:00

Analysis sample: Based on 14 of 21 usable stored comments. Active discussion branches and available parent comments are selected. The analysis input was further shortened to fit its context limit.

This sample may omit parts of the full thread. Selected themes do not measure community opinion or how common a view is.

### Dictation and accessibility use cases for speech\-to\-text

Comments discuss speech\-to\-text for long\-form writing and accessibility, including an 84\-year\-old Croatian stroke survivor whose impaired speech and extraneous mouth sounds overwhelm general STT; a dictation model is suggested to ignore disfluencies and handle corrections and punctuation\. Another describes streaming dictation plus LLM cleanup for hands\-free long\-form writing, and a local dictation app\.

Sources: [Comment 50008908](<https://news.ycombinator.com/item?id=50008908>) · [Comment 50009019](<https://news.ycombinator.com/item?id=50009019>) · [Comment 50009407](<https://news.ycombinator.com/item?id=50009407>) · [Comment 50012525](<https://news.ycombinator.com/item?id=50012525>)

### Modern assistants struggle with music playback commands

Commenters note that older offline assistants, such as pre\-Siri iPhones and the old Google Assistant, reliably played songs from a local library, while Gemini often fails or is slower\. They attribute the difficulty to changing song catalogs, uncommon names, and unusual word orders that general LLM\-style tools do not handle well\.

Sources: [Comment 50009314](<https://news.ycombinator.com/item?id=50009314>) · [Comment 50010339](<https://news.ycombinator.com/item?id=50010339>)

### Local, on\-device processing and ownership

Comments describe taking an Echo Show off Amazon's cloud and processing locally with its own CPU and Home Assistant; a dictation app that runs Qwen locally with no telemetry; and an on\-device keyboard STT\.

Sources: [Comment 50011633](<https://news.ycombinator.com/item?id=50011633>) · [Comment 50012525](<https://news.ycombinator.com/item?id=50012525>) · [Comment 50009323](<https://news.ycombinator.com/item?id=50009323>)

### Comparative accuracy and speed of STT models

One user reports Qwen ASR 1\.7B correctly recognized 168 of 170 messages versus Whistle at 70 of 170, then a template\-constrained Whistle setup reached 164 of 170\. Another asks how a new model compares with Parakeet, and a reply says Parakeet is much faster and smaller but less accurate than Whisper Large for serious dictation\. Cohere is also described as better but slower with a token output limit\.

Sources: [Comment 50011633](<https://news.ycombinator.com/item?id=50011633>) · [Comment 50009751](<https://news.ycombinator.com/item?id=50009751>) · [Comment 50010763](<https://news.ycombinator.com/item?id=50010763>) · [Comment 50009407](<https://news.ycombinator.com/item?id=50009407>)

### Specialized dictation and grammar\-constrained recognition

Comments contrast general\-purpose STT with dictation models that ignore disfluencies and support corrections, and with template\- or grammar\-constrained recognition\. One user constrained Whistle to a select set of templates and trained a small network on generated utterances, nearly matching Qwen; another connects this to CMU Sphinx4 and JSGF grammar\-based recognition\.

Sources: [Comment 50009019](<https://news.ycombinator.com/item?id=50009019>) · [Comment 50011633](<https://news.ycombinator.com/item?id=50011633>) · [Comment 50012981](<https://news.ycombinator.com/item?id=50012981>)

### Source\-available vs open\-source licensing

A commenter initially calls FUTO keyboard open\-source but edits to clarify it is source\-available, not open source: the license is revocable and non\-transferable, limiting forking and redistribution\.

Sources: [Comment 50009323](<https://news.ycombinator.com/item?id=50009323>)

## Sources & coverage

AI-generated summary · 2026\-10\-08T21:03:06\.555837\+00:00

Based on 8 of 10 usable stored comments, selected by depth and branch activity. This is a sample of the discussion. The model input was further shortened to fit its context limit. Article text may also be shortened.

Generated using deepseek\-ai/DeepSeek\-V4\.1\-Flash. Check the linked sources for full context.
