Whistle: Speech To Text In 16.9 MB
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB speech recognition model designed to run locally on a CPU without external dependencies. The company reports support for seven languages and strong results on several speech benchmarks, but its performance figures come from company testing and differ across datasets.

Cactus Compute released Whistle, a speech recognition model packaged in a 16.9 MB file and designed to transcribe audio locally on a CPU. The company says the model handles seven languages, produces its first token in 11.1 milliseconds in a 10-second audio test on an Apple M4 Pro, and can run in the same C++ engine as its Needle model.

Whistle accepts 16 kHz mono audio clips of up to 30 seconds in English, German, French, Spanish, Italian, Dutch and Polish. Cactus says it detects the language automatically unless the user specifies one. Along with transcripts, the model can return word-level timestamps and probabilities, or produce speech embeddings without decoding words.

The company says audio stays on the device in its browser demonstration, where the first use downloads the model. It also describes Whistle as running without dependencies and as using the same container, quantization and CPU engine as Needle. These are product and implementation details supplied by Cactus; the release material does not provide an independent audit of the browser demonstration or privacy behavior.

In Cactus’s comparison, Whistle recorded lower word error rates than the compared systems on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base performed better on TED-LIUM, AMI and the MLS average. The report notes that benchmark coverage differs: Moonshine tiny v2 is English-only, and some Whisper results were not published or used a different AMI subset.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact speech-to-text model that runs on-device using the company’s C++ engine.

Local Transcription on Smaller Devices

A model that fits in 16.9 MB could make speech recognition easier to deploy on phones, wearables, robots, smart-home equipment and other devices with limited resources. Local processing may also be useful where connectivity is unreliable or where developers want audio handled on the device, though the release does not quantify power use, memory requirements or privacy protections beyond describing local processing.

The reported speed results speak to responsiveness, not every part of the user experience. Cactus measured 11.1 ms to first token for 10 seconds of audio on an Apple M4 Pro, with 5.9 ms for five seconds and 36.3 ms for 30 seconds. The company says its decode rate was 1,319 tokens per second in the same comparison. Those results may not carry over to other processors, devices or application workloads.

For developers already using Needle, shared engine components could reduce the work of combining speech recognition with on-device language-model features. Cactus says one binary can process a clip into tool calls, but the release does not describe a complete application or provide an independent demonstration of that workflow.

Amazon

on-device speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Fits Cactus’s Models

Whistle is presented as a speech model for mobile and embedded uses, rather than a general-purpose cloud transcription service. Its encoder converts audio into frames, and its decoder generates text; the model can also return frame-based embeddings. According to Cactus, several attention blocks are shared with Needle’s implementation, while speech-specific gated cross-attention connects the audio representation to the decoder.

The release says Whistle uses five-beam decoding, supports keyword biasing and limits transcripts to 320 text tokens. The company also offers adjustable decoder depth: models trained at different decoder depths can be selected at load time, while all eight encoder blocks continue to run. This gives developers a way to trade some decoding capacity for a smaller or faster configuration, although the report does not quantify the trade-offs at each setting.

Cactus compared Whistle with Whisper base and Moonshine tiny v2 using each system’s official runtime and default settings, according to its report. The timing test used 10 seconds of audio on an Apple M4 Pro. Cactus says Whisper pads every input to 30 seconds, while Whistle’s time to first token varies with clip length. Timing figures therefore reflect both model behavior and runtime choices, not a controlled comparison on every possible device.

““Speech recognition model” for “mobiles, wearables, robots, smart home, automotive and microcontrollers.””

— Cactus Compute, in its October 2, 2026 release

Performance Beyond Company Tests

The published material is a company report; it does not identify independent testing or peer-reviewed evaluation. Cactus provides selected benchmark results and device-specific timing measurements, but the material does not give enough detail to establish how performance changes across different CPUs, microphones, accents, noisy environments or longer recordings.

Several practical details remain unreported, including the model’s memory and battery use, licensing terms, exact hardware requirements and availability outside the described release and browser demo. The report also does not establish whether the seven-language results are comparable across datasets or real-world conditions. Claims about local processing describe the product’s design, but readers do not receive a third-party verification of data handling.

Deployment and Independent Testing

Cactus’s release points to use through its C++ engine and browser demonstration, but it does not set out a future release timetable or identify a forthcoming independent evaluation. The next useful evidence would include reproducible tests on a range of target devices, with hardware, runtime settings and dataset details reported alongside accuracy, speed and resource use.

Developers considering Whistle can compare its results with their own audio and hardware, paying attention to the company’s stated limits: clips up to 30 seconds, seven supported languages and configurable decoder depth. Whether the small file and reported latency translate into reliable performance in specific products remains to be established.

Key Questions

What is Whistle?

Whistle is an on-device speech recognition model released by Cactus Compute. It converts audio into text and can also return word timestamps, probabilities and speech embeddings.

How large is the model, and which languages does it support?

Cactus says the model file is 16.9 MB. It supports English, German, French, Spanish, Italian, Dutch and Polish, with automatic language detection unless a language is specified.

Does Whistle send audio to a server?

Cactus says audio in its browser demonstration stays on the device after the model is downloaded. The release does not provide an independent privacy audit or establish the data handling of every possible deployment.

How fast is Whistle?

In Cactus’s test on an Apple M4 Pro CPU, Whistle reached its first token in 11.1 milliseconds for a 10-second clip. The company reported 5.9 milliseconds for five seconds and 36.3 milliseconds for 30 seconds; results on other hardware may differ.

How did Whistle compare with Whisper and Moonshine?

Cactus reported lower word error rates for Whistle on several listed benchmarks, including LibriSpeech, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base scored better on TED-LIUM, AMI and the MLS average. The report warns that benchmark availability and, for AMI, dataset subsets differ.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

URXR One Kickstarter Secures $1.3M To Revolutionize Ultralight Tethered VR Headsets

URXR One’s Kickstarter campaign raises $1.3 million to develop a lightweight tethered VR headset with passthrough technology, aiming to transform VR experiences.

Valve Announces Steam Frame Price, Release Date, And Accessories – Pre-orders Now Open

Valve announces Steam Frame VR headset pricing, release schedule, accessories, and pre-order details, with no immediate options for pre-order.

Need Assistance With Virtual Reality? Here’s What You Should Do

Learn how to resolve common VR support problems with practical steps and expert advice, based on recent community reports.

Experience The Future Of VR: Discovery Rogue Planet Launches Today

Discovery Rogue Planet, an upcoming virtual reality experience, launches today, offering players a new exploration adventure into a mysterious alien world.