Run Cactus Whistle Locally on Windows 11, Mac and Kali Linux: Free Offline Speech-to-Text Install Guide (16.9 MB, Beats Whisper Base)
Cactus Whistle is a free, open speech-to-text model that fits in a single 16.9 MB file and installs locally on Windows 11, Mac or Kali Linux in one command, running on an ordinary processor with no graphics card and no internet. It was released on October 2, 2026 under Apache 2.0 and is trending on Hugging Face this week. The surprise is the comparison its makers invite: Whistle is 8.6 times smaller than OpenAI's Whisper base, the model most "free transcription" tools quietly run, reaches its first word 6.6 times faster, and on the two standard LibriSpeech tests it makes fewer errors: 4.31% of words wrong on clean audiobook speech, 10.49% on the harder set. It is not better everywhere, and this page shows exactly where Whisper still wins. The catches are real too: Whistle hears seven languages, English, German, French, Spanish, Italian, Dutch and Polish, it listens to 30 seconds at a time, and it is a model with a Python package, not an app with a record button.
Jake runs a phone repair shop and spends an hour a day driving between suppliers. That hour is when he remembers things, so his phone is full of voice notes: "order three iPhone 17 screens," "call Mrs. Patel about the Pixel," "check why the Galaxy came back." He never listens to them again, because listening takes as long as recording. His friend Ethan, a developer, had one condition before helping: nothing recorded in Jake's car would be uploaded anywhere. This page is what they built. It covers what Whistle is in plain English, the honest benchmarks against Whisper and Moonshine, the free dictation your computer already has, how to install Whistle on Windows, Mac and Kali Linux, how to transcribe recordings longer than 30 seconds, how Jake's voice notes became a daily to-do list, how Whistle turns a spoken command into an action with its sibling model Needle, the mistakes that make it look bad, and what the cloud alternatives cost.
New to running AI on your own machine? Our plain-English series on running AI locally explains the words, models, parameters and quantization, in about ten minutes. Whistle is one of the gentlest first models there is, so you can also just follow the steps below.
What Cactus Whistle is, in plain English
Speech-to-text, also called transcription or voice-to-text, means turning recorded speech into written words. For years the free way to do it on your own computer was OpenAI's Whisper, released in 2022, which is excellent and also heavy: its smallest useful version is a 145 MB download that wants a decent processor, and its best versions want a graphics card with 10 GB of memory.
Whistle comes from Cactus Compute, a small company that builds AI for phones, watches, cars and the chips inside smart-home devices. Its whole philosophy is "make it tiny and make it run anywhere," and Whistle is that philosophy applied to speech. The entire model is one 16.9 MB file, stored at 2 to 4 bits per number, and it runs on the same small C++ engine as Cactus's other model, Needle, with no dependencies and no graphics card.
Whistle does three jobs, all on your device:
- Transcription: 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish. It detects the language unless you name it. Silence returns an empty transcript rather than an invented sentence, which matters more than it sounds; many speech models hallucinate words into quiet audio.
- Word timestamps: every word with its start time, end time and probability, taken from the model's own attention, so an app can highlight words as they are spoken, jump to a word, or cut a recording on a word.
- Speech embedding: the raw encoder output, one row per 80 milliseconds, for matching and searching audio without ever writing a transcript.
One feature deserves its own paragraph: keyword biasing. You hand Whistle a list of the names, places and product words your users actually say, and it favors them while it decodes. Every speech model mangles proper nouns; "Mrs. Patel" becomes "Mrs. Battle," and "Pixel 10" becomes "pixel ten" or worse. Giving Whistle Jake's supplier and product names was the single change that made his transcripts usable.
The facts:
- Size: one 16.9 MB file,
whistle.cact. - Input: 16 kHz mono, up to 30 seconds per pass.
- Languages: English, German, French, Spanish, Italian, Dutch, Polish.
- Design: a log-mel front end and a convolutional stem feed an audio encoder; a Needle-shaped decoder reads it through gated cross-attention at every layer. The decoder is "laddered," so every depth from 2 layers up is a working model, chosen at load time with
--audio-depth. Shallower is faster and slightly less accurate. - Speed: 11.1 ms to the first word and 1,319 tokens per second on an Apple M4 Pro processor, on a 10-second clip. The first-word delay grows with clip length: 5.9 ms at 5 seconds, 11.1 ms at 10, 36.3 ms at 30.
- License: Apache 2.0, free for personal and commercial use.
- Runs on: Windows (x64 and ARM), Mac, Linux (x86-64, ARM64, ARMv7, RISC-V, even the MIPS chips in cameras and routers), Android, iOS, Apple Watch and TV, and in the browser through WebAssembly.
What Whistle is not
- Not an app. There is no window with a record button. You use it from Python, from a command line, or inside something you build. The
needle whistle playgroundcommand does give you a quick microphone demo. - Not for every language. Seven European languages. For Hindi, Tamil, Chinese, Japanese, Arabic and the other 90-odd languages Whisper handles, Whisper or a cloud service is still the answer.
- Not long-form by itself. Thirty seconds per pass. Longer recordings are split into pieces first, which this page shows how to do.
- Not a meeting transcriber. It does not tell speakers apart, and on the AMI meeting benchmark it gets about one word in four wrong, as its own chart shows.
Whistle vs Whisper vs Moonshine: the honest numbers
Cactus published word error rates, the percentage of words a model gets wrong, on nine standard test sets, measured over 86,174 recordings and scored with Whisper's own text normalizers. The Whisper and Moonshine figures are the ones those models' authors published, not Cactus's own runs of them. Lower is better.
| Test set (what it is) | Whistle word error rate | Who is ahead |
|---|---|---|
| LibriSpeech test-clean (clear audiobook reading) | 4.31% | Whistle, ahead of Whisper base |
| LibriSpeech test-other (harder, accented reading) | 10.49% | Whistle |
| SPGISpeech (earnings calls, clean) | 7.65% | Whistle, ahead of Moonshine |
| Earnings-22 (earnings calls, varied accents) | 19.01% | Whistle, ahead of Moonshine |
| AMI (meeting room recordings) | 26.07% | Whisper base |
| AMI cleaned | 22.87% | Only Whistle reported |
| TED-LIUM (TED talks) | 7.61% | Whisper base |
| FLEURS (read speech, average of the seven languages) | 21.4% | Whistle |
| MLS (audiobooks, six languages, no English) | 24.9% | Whisper base |
Cactus's own summary is fair: Whistle is ahead on LibriSpeech clean and other, on SPGISpeech, on Earnings-22 and on the FLEURS average; Whisper base is ahead on TED-LIUM, on AMI and on the MLS average. Moonshine is English-only, so it has no bars on the multilingual sets, and Whisper never published SPGISpeech, Earnings-22 or AMI-cleaned numbers. Cactus also checked that none of the test audio appears in Whistle's training data, by comparing audio checksums and speaker IDs, which is more than most model cards bother to say.
Read the pattern, not the wins. Whistle is strongest on clear speech from one person, which is exactly what a voice note, a dictated letter or a customer call is. It is weakest on meetings and on non-English audiobooks, where a 22 to 26% error rate means a transcript you can search but would not publish. And remember that Whisper base is Whisper's second-smallest model; Whisper small, medium and large are more accurate than base and much larger than Whistle. Nobody is claiming a 16.9 MB file beats a 1.5 GB one.
Size and speed
| Whistle | Whisper base | Moonshine tiny v2 | |
|---|---|---|---|
| Size on disk | 16.9 MB | 145.3 MB | 41.9 MB |
| Time to first word, 10 s clip | 11.1 ms | 73.2 ms | 22.8 ms |
| Decode speed, tokens per second | 1,319 | 266 | 262 |
| Precision | 2 to 4 bit | fp32 on CPU | int8 |
All three were measured on an Apple M4 Pro, each on its official runtime at default settings: Whistle's C++ engine with 5 beams, the openai-whisper package, and moonshine-voice. A 10-second clip decoding at 1,319 tokens per second means the transcript is finished almost before you have lifted your finger off the key. On an older or cheaper processor, expect slower numbers in the same proportions.
Where Whistle sits in the Whisper family
Since "which Whisper is best" is a common question, here are OpenAI's own Whisper sizes, so you can see what Whistle is being compared with:
| Whisper model | Parameters | Memory needed | Relative speed |
|---|---|---|---|
| tiny | 39M | about 1 GB | about 10x |
| base (the one Whistle is compared with) | 74M | about 1 GB | about 7x |
| small | 244M | about 2 GB | about 4x |
| medium | 769M | about 5 GB | about 2x |
| large | 1,550M | about 10 GB | 1x |
Whisper is MIT-licensed and handles about a hundred languages, which is why it remains the default for anything multilingual or long-form. Whistle's pitch is different: the accuracy of Whisper base on clear speech, in a file 8.6 times smaller, fast enough for a watch.
Do you need a model at all? Free speech-to-text you already own
Before installing anything, a quick honesty check, because most "speech to text free" searches are answered by a key you already have:
- Windows 11 voice typing: press Win+H in any text box and talk. It punctuates, it is free, and it works offline for the main languages once the speech pack is downloaded. Our old guide to Windows speech recognition and voices covers the settings behind it.
- Mac: Dictation in System Settings, with a keyboard shortcut you choose.
- Phones: the microphone key on the keyboard, on both Android and iPhone.
- Word and Google Docs: both have a Dictate or Voice typing button.
Those are for dictating: you talk, text appears. They are poor at transcribing: turning a recording you already have into text, in bulk, with timestamps, inside your own program. "Transcribe audio to text" is one of the most searched phrases in this whole subject, and it is the job Whistle is for: a folder of voice notes, customer calls, lecture clips or interviews, processed on your machine, with the names you care about spelled right.
How to install Whistle locally on Windows 11, Mac and Kali Linux
Whistle ships inside the cactus-needle Python package, which also carries the engine. There is nothing else to install: no PyTorch, no CUDA, no graphics driver. The package covers Windows x64 and ARM, Mac on Apple Silicon, and Linux on x86-64 and ARM64.
Step 1: a virtual environment
Make a virtual environment first, so this install cannot disturb anything else, and because Kali and other recent Linux systems refuse system-wide pip installs with "externally-managed-environment." Our guide to that pip error explains why that refusal is a feature.
python3 -m venv ~/whistle
source ~/whistle/bin/activate
pip install cactus-needle
On Windows, the activate line is whistle\Scripts\activate. If you want to record from a microphone or feed in files that are not already 16 kHz WAV, install the extra: pip install "cactus-needle[mic]", with the quotes, because Kali's zsh reads square brackets as a file pattern without them.
Step 2: transcribe a clip
import needle
result = needle.transcribe("clip.wav")
print(result["text"])
# turn off the kitchen lights
The first call downloads the engine and the 16.9 MB model once and caches them in ~/.cache/cactus-needle/v3/. Every call returns the text, the detected language, the milliseconds to the first token and the decoder's tokens per second. Three options cover most needs:
r = needle.transcribe(
"note.wav",
language="en", # skip detection when you know the language
keywords=["Patel", "Pixel 10", "iPhone 17", "Galaxy S26"], # names it must get right
word_timestamps=True, # start, end and probability per word
)
print(r["text"], r["language"], r["ttft_ms"], r["decode_tps"])
for w in r["words"]: # each entry carries the word's start, end and probability
print(w)
needle.transcribe() is the entry point for speech; the model loads once per process and stays loaded, so a loop over many clips does not pay the start-up cost again. needle.Whistle() gives you the model as an object for the speech-embedding feature, embed(audio), or to hold a particular weights file. Two commands are handy on day one. needle whistle playground transcribes from your microphone, press Enter to record and Enter again to stop, and takes --language en, --keywords "..." and --word-timestamps. needle whistle compare clip.wav runs the same clip through Whistle, Whisper tiny, Whisper base and Moonshine tiny v2 and shows each one's timing; it needs pip install "cactus-needle[mic,compare]".
The command-line engine (no Python at all)
For scripts, servers or small boards, the bare engine is one small binary per platform plus the model file:
needle download linux-x86_64 # or macos-arm64, windows-x86_64, linux-arm64 for a Raspberry Pi
needle download whistle
./linux-x86_64/needle --model whistle.cact --audio clip.wav --audio-word-timestamps
On Windows the binary is needle.exe. The bare engine reads no environment variables; every behavior is a compiled default or an explicit flag, so the same file gives the same transcript on every machine, which is a quiet gift when you are debugging.
Fully offline, including air-gapped machines
Inference never touches the network; only the first download does. For a computer with no internet at all, the Python package's own offline notes give the recipe:
- On a connected machine, run
pip download cactus-needle -d wheelsto collect the package, andneedle download whistle --out modelsto fetch the speech weights.needle fetchorneedle download needle3pulls the engine and Needle's weights if you want those too. - Copy the
wheelsandmodelsfolders, and the cache folder~/.cache/cactus-needle/v3/, to the offline machine. - Install there with
pip install --no-index --find-links wheels cactus-needle. - Point the package at the weights with the environment variable
NEEDLE_WHISTLE_WEIGHTS=/path/to/whistle.cact, or passweights=toneedle.Whistle(). - Set
HF_HUB_OFFLINE=1, so a missing file fails at once with a clear message instead of hanging on a download that can never finish.
That is the setup for a clinic laptop, a courtroom, a factory floor or any place where the rule is "nothing leaves the building." The whole kit is a few tens of megabytes.
Transcribe audio to text from any source: phone notes, WhatsApp, Zoom, video
"Transcribe audio to text" covers a dozen file types, and almost none of them arrive as 16 kHz mono WAV. ffmpeg converts all of them with one pattern, -ac 1 -ar 16000, which means one channel at 16,000 samples per second. Install it once: on Windows winget install Gyan.FFmpeg, on Mac brew install ffmpeg, on Kali and Debian sudo apt install ffmpeg.
| Source | Usual file | Convert with |
|---|---|---|
| iPhone or Android voice memo | .m4a | ffmpeg -i note.m4a -ac 1 -ar 16000 note.wav |
| WhatsApp or Telegram voice note | .opus or .ogg | ffmpeg -i note.opus -ac 1 -ar 16000 note.wav |
| Zoom, Teams or Meet recording | .m4a or .mp4 | ffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 meeting.wav (-vn drops the video) |
| A video file or screen recording | .mp4, .mkv, .mov | Same as above, with -vn |
| Windows Sound Recorder | .m4a (or .wav at 44.1 kHz) | ffmpeg -i rec.m4a -ac 1 -ar 16000 rec.wav |
| A dictation machine or old recorder | .mp3, .wma, .dss | Same command; ffmpeg reads nearly everything except some proprietary .dss files |
Remember the meeting caveat before you transcribe a Zoom call: several voices through laptop microphones is Whistle's weakest case. One person talking into a phone is its best.
Transcribe a voice note on Windows 11, step by step
For anyone who has never opened a terminal, here is the whole path from a phone recording to text on a Windows laptop:
- Install Python from python.org, ticking "Add python.exe to PATH" in the installer, and ffmpeg with
winget install Gyan.FFmpegin a terminal. Close and reopen the terminal afterward so it finds both. - Make a folder, for example
C:\Transcribe, and copy the voice note into it. - Open a terminal in that folder: type
cmdin the folder's address bar in File Explorer and press Enter. - Create and activate a virtual environment:
python -m venv venv, thenvenv\Scripts\activate. - Install Whistle:
pip install cactus-needle. - Convert the note:
ffmpeg -i note.m4a -ac 1 -ar 16000 note.wav. If it is longer than 30 seconds, use the splitting command in the next section instead. - Transcribe:
python -c "import needle; print(needle.transcribe('note.wav')['text'])". The first run downloads the 16.9 MB model; later runs are instant.
Mac and Kali users follow the same seven steps with source venv/bin/activate in step 4 and the package manager of their system in step 1. If your microphone itself is the thing not working, our fix for USB audio Code 10 errors on Windows 11 is the page you need before any model.
How to transcribe recordings longer than 30 seconds
Whistle listens to 30 seconds at a time, so a 12-minute voice note has to be cut up first. The simplest reliable tool is ffmpeg, which is free on every platform. This one command converts any recording, including the .m4a files phones produce, to mono 16 kHz WAV and splits it into 28-second pieces:
mkdir -p chunks
ffmpeg -loglevel error -i note.m4a -ac 1 -ar 16000 -f segment -segment_time 28 chunks/part%03d.wav
Then transcribe the pieces in order and join them:
from pathlib import Path
import needle
parts = []
for wav in sorted(Path("chunks").glob("part*.wav")):
r = needle.transcribe(str(wav), language="en", keywords=["Patel", "Pixel 10", "iPhone 17"])
parts.append(r["text"])
print(" ".join(parts))
One honest limitation: a cut every 28 seconds will sometimes land in the middle of a word, and that word may come out wrong or missing. For voice notes and dictation it hardly matters. For an interview you will quote, split on silence instead: ffmpeg's silencedetect filter lists the quiet moments, and cutting at the nearest pause under 30 seconds keeps every word whole. Whistle's word timestamps also tell you when a chunk's last word ends early, which is a sign the cut fell on speech.
A real example: Jake's voice notes become a to-do list
Here is the whole of what Ethan built, in one script that runs on Jake's laptop every evening. It watches the folder where his phone syncs voice notes, converts and chunks anything new, transcribes it with Jake's product and supplier names as keywords, and appends the text to a dated file. Nothing leaves the laptop.
import datetime as dt, subprocess, tempfile
from pathlib import Path
import needle
NOTES = Path.home() / "VoiceNotes" # where the phone syncs .m4a files
DONE = NOTES / "done"; DONE.mkdir(exist_ok=True)
KEYWORDS = ["Patel", "Pixel 10", "iPhone 17", "Galaxy S26", "Mobile Parts Direct", "screen", "battery"]
for m4a in sorted(NOTES.glob("*.m4a")):
with tempfile.TemporaryDirectory() as tmp:
subprocess.run(["ffmpeg", "-loglevel", "error", "-i", str(m4a), "-ac", "1", "-ar", "16000",
"-f", "segment", "-segment_time", "28", f"{tmp}/p%03d.wav"], check=True)
text = " ".join(needle.transcribe(str(w), language="en", keywords=KEYWORDS)["text"]
for w in sorted(Path(tmp).glob("p*.wav")))
day = dt.date.fromtimestamp(m4a.stat().st_mtime)
with open(NOTES / f"{day}.md", "a", encoding="utf-8") as f:
f.write(f"\n- [ ] ({m4a.stem}) {text.strip()}\n")
m4a.rename(DONE / m4a.name)
Each note becomes a checkbox line in that day's Markdown file, so Jake's "order three iPhone 17 screens" is now something he can tick off. The keyword list was the difference between useful and not: before it, "Mrs. Patel" came out three different ways in one week. After it, every mention matched. The car noise still costs a few words on the motorway, which is the AMI lesson in miniature: clear, close speech is where a small model shines.
Once the notes are text, the next step is searching them by meaning. Our guide to EmbeddingGemma 2 shows how to build a private search box over transcripts and the original audio, also offline.
From spoken words to actions: Whistle plus Needle
The reason Cactus built Whistle on the same engine as Needle becomes clear when you load both at once. Needle is Cactus's tiny tool-calling model: give it the functions your app exposes and it picks the right one and fills in the arguments. With Whistle beside it, a clip goes in and a function call comes out, in one pass, with the transcript attached:
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav
{"function_calls": [{"name": "set_lights", "arguments": {"room": "kitchen", "on": false}}],
"confidence": 0.94,
"audio_text": "turn off the kitchen lights",
"audio_language": "en"}
That is a complete offline voice assistant in two files totaling under 50 MB, and it runs on a Raspberry Pi. Needle's full model is 29 MB, and like Whistle it is laddered, so a 4-layer version of about 8 MB works for simple command sets. A request no tool covers returns an empty list rather than a guess, and the confidence number lets your app act, confirm or refuse. For a shop, "book a screen repair for Thursday" becoming a calendar entry without a cloud account is the kind of thing that used to need a product team. If the decisions after the transcript are the hard part, our guide to decision models in Ollama covers the other small models built for exactly that.
What word timestamps and embeddings are for
Two of Whistle's three jobs are easy to overlook, so here is what people build with them.
Word timestamps give every word a start time, an end time and a probability. With those you can make captions for a video, since each caption line is just a run of words and their times; build a "click a word to jump there" player for lectures and interviews; cut the dead air out of a recording automatically, because the gaps between word end-times are the silences; and spot the words the model was unsure of, by their low probability, so a human checks only those. Jake's version was simpler: the words with low probability in his transcripts were almost always supplier names he had not yet added to the keyword list.
Speech embeddings are the encoder's output before any words are decoded: one row of numbers per 80 milliseconds of audio. Two clips of the same phrase produce similar rows, which lets an app match or search audio without ever transcribing it: find every voice note that mentions the same thing, detect a repeated jingle, or check whether a recorded phrase matches a stored one. It is a research-flavored feature, but it comes free with the same 16.9 MB file.
Choosing a decoder depth
Whistle's decoder is laddered: every depth from 2 layers upward was trained as a working model, and the command-line engine picks one at load time with --audio-depth. The encoder always runs all eight of its blocks, so the saving is only in the decoder. Cactus publishes no accuracy figures per depth, so the honest advice is: on a laptop, use the default and never think about it; on a watch or a microcontroller, start shallow, run needle whistle compare on ten of your own clips, and go one step deeper until the errors stop falling.
Why Whistle sounds wrong: the common mistakes
Speech models rarely crash. They return confident text with the wrong words in it, so these are worth reading before you judge the model.
Audio at the wrong sample rate
Whistle wants 16 kHz mono. Phone recordings are usually 44.1 or 48 kHz stereo. The base package accepts only 16 kHz WAV or raw samples; either install the [mic] extra, which handles other rates, or convert with the ffmpeg command above. Feeding 48 kHz audio into a 16 kHz model produces slowed, garbled text, not an error.
Clips over 30 seconds
Anything past 30 seconds is not transcribed. Split first.
An unsupported language
Seven languages only. Hindi, Tamil, Chinese, Japanese, Arabic, Portuguese and the rest come out as nonsense or as the nearest supported language. Use Whisper or a cloud service for those.
Meetings and crowded rooms
On the AMI meeting benchmark Whistle gets 26% of words wrong, and it does not separate speakers. Use it for one clear voice near the microphone; use a larger Whisper or a dedicated meeting tool for the conference room.
Leaving out keywords
Names and product words are where every small speech model fails. Pass the ones that matter in keywords=[...]. Do not pass a dictionary: a short list of the words that must be right works better than a long list of words that might appear.
Running at the shallowest ladder depth
Whistle's decoder runs at any depth from 2 layers up, chosen with --audio-depth. Shallower is faster and less accurate. If a tutorial set a tiny depth for a microcontroller, raise it on a laptop.
"needle: command not found"
The needle command lives inside the virtual environment you installed it in. Activate the environment first, or call it by its full path.
Expecting an app
Whistle is a model and a package. For a window with a record button, use Windows voice typing or a dictation app, or build the small script above once and reuse it.
Judging it on faded, distant audio
A note recorded with the phone in a pocket on a motorway will be rough with any model. Record close and clear, and the small model keeps up with the big ones.
Whistle vs Whisper, whisper.cpp, Moonshine, Windows voice typing and the cloud
| Option | What it is | Choose it when |
|---|---|---|
| Whistle | 16.9 MB open model, 7 languages, CPU only, Apache 2.0 | Short clear clips, keyword accuracy, tiny devices, voice commands with Needle |
| Whisper (OpenAI) | Open models from 39M to 1.55B parameters, about 100 languages, MIT | Other languages, long recordings, best accuracy with a GPU |
| whisper.cpp | Whisper rewritten in C++ for CPUs and phones | You want Whisper's languages without PyTorch |
| Moonshine tiny v2 | 41.9 MB English-only model for devices | English only on a small device; Whistle is ahead on the shared tests |
| Windows voice typing (Win+H) | Built-in dictation | You want to talk into a document right now |
| MAI-Transcribe-2 (Microsoft) | Cloud API, 60 languages, 2.5% error on its streaming test | Best accuracy and many languages, and uploading is acceptable |
| Amazon Transcribe | AWS's managed service, billed per second | You already run on AWS and want speaker labels, redaction or call analytics |
A fair ten-minute bake-off on your own recordings
Benchmarks are other people's audio. Before you commit, test on yours. Cactus ships the tool for it:
- Collect ten clips of under 30 seconds that represent your real use: your voice, your room, your names and product words.
- Write down what was said in each, carefully. This is your answer key.
- Install the comparison extras:
pip install "cactus-needle[mic,compare]". - Run
needle whistle compare clip.wavfor each clip. It prints Whistle, Whisper tiny, Whisper base and Moonshine tiny v2 side by side, with each one's timing. - Count the wrong words per model against your answer key, then run Whistle again with your
keywordslist and count once more. - Decide on the whole picture: accuracy on your audio, speed on your machine, the languages you need, and whether anything may leave the building.
What the cloud costs, for comparison
Whistle costs nothing per hour of audio. For scale, here are Amazon Transcribe's on-demand prices in US East (N. Virginia), from AWS's official price list published September 11, 2026: standard batch transcription is $0.0001 per second, which is $0.36 per hour of audio, and streaming is $0.0001667 per second, $0.60 per hour. Microsoft's MAI-Transcribe-2 is $0.10 an hour for recorded audio and $0.54 for streaming through the end of 2026. Jake's five hours of voice notes a month would cost under $2 on either, so price is not the argument for Whistle. Privacy, no account, no internet and the ability to run on a $50 board are. If you do want a managed service on AWS, our explainer on Hugging Face models on AWS covers hosting open models there as well.
Cactus Whistle: frequently asked questions
What is Cactus Whistle?
Whistle is a free, open speech-to-text model from Cactus Compute, released October 2, 2026 under Apache 2.0. It is a single 16.9 MB file that transcribes English, German, French, Spanish, Italian, Dutch and Polish on any CPU, with word timestamps and keyword biasing.
Is Whistle better than Whisper?
On five of eight shared tests, yes against Whisper base: both LibriSpeech sets, SPGISpeech, Earnings-22 and FLEURS. Whisper base is ahead on TED-LIUM, AMI meetings and MLS. Larger Whisper models are more accurate than base and far larger than Whistle.
How do I install Whistle?
Create a Python virtual environment, run pip install cactus-needle, then call needle.transcribe("clip.wav")["text"]. The first call downloads the engine and the 16.9 MB model. Add the [mic] extra for microphone input and other sample rates.
Does Whistle need a GPU?
No. It runs on the processor alone, at 11.1 ms to the first word and 1,319 tokens per second on an Apple M4 Pro. It also runs on Raspberry Pi boards, phones, watches and in the browser.
What languages does Whistle support?
English, German, French, Spanish, Italian, Dutch and Polish. It detects the language unless you set one. For other languages, use Whisper or a cloud service.
How long a recording can Whistle transcribe?
Up to 30 seconds in one pass. For longer audio, split it into pieces with ffmpeg, transcribe each piece in order and join the text. Splitting on silence keeps words whole.
Can Whistle transcribe audio to text offline?
Yes, completely. After the first download, nothing is sent anywhere. Set HF_HUB_OFFLINE=1 and copy the cache folder for a machine with no internet at all.
What audio format does Whistle need?
16 kHz mono. Convert phone recordings with ffmpeg -i note.m4a -ac 1 -ar 16000 note.wav, or install the [mic] extra, which handles other sample rates.
Does Whistle give word timestamps?
Yes. Pass word_timestamps=True and each word returns with its start time, end time and probability, read from the decoder's own attention.
What is keyword biasing in Whistle?
A list of names, places and product words you pass with keywords=[...]. Whistle favors them while decoding, so proper nouns that small models usually mangle come out right.
Is Whistle free for commercial use?
Yes. It is released under the Apache 2.0 license, with no revenue limit.
How do I run Whistle locally on Windows 11?
Install Python, create a virtual environment with python -m venv venv and venv\Scripts\activate, run pip install cactus-needle, convert your recording to 16 kHz mono WAV with ffmpeg, and call needle.transcribe("note.wav"). The 16.9 MB model downloads once; everything after that is offline.
Can Whistle run on Windows?
Yes. The cactus-needle package supports Windows x64 and ARM, and the standalone engine is needle.exe in the windows-x86_64 or windows-arm64 folder.
Can Whistle run on Kali Linux or a Raspberry Pi?
Yes. On Kali, install cactus-needle inside a virtual environment. For a Raspberry Pi, download the linux-arm64 engine folder and run the needle binary with whistle.cact.
Does Whistle separate speakers?
No. It transcribes one stream of speech without speaker labels, and its meeting-room accuracy is about 74%. For meetings with several speakers, use a larger Whisper or a cloud service with diarization.
What is Needle, and how does it work with Whistle?
Needle is Cactus's 8 to 29 MB tool-calling model. Loaded with Whistle in the same engine, a voice clip goes in and a function call with arguments comes out, plus the transcript, making an offline voice assistant.
How does Whistle compare with Amazon Transcribe?
Whistle is free and offline, seven languages, no speaker labels. Amazon Transcribe is a managed service at $0.36 per hour for batch and $0.60 for streaming in US East, with many languages, speaker labels and redaction.
By the end of the week Jake had stopped dreading his own voice notes. Each evening a Markdown file appeared with the day's reminders as checkboxes, Mrs. Patel's name spelled correctly in every one, and nothing he had said in the car had gone further than the laptop under the counter. "It's smaller than a photo," he said, looking at the 16.9 MB file, as if it might be a trick. It is not a trick. It is what happens when a company builds for watches and cameras first, and laptops turn out to be the easy case.
If you keep one line from this page
Record close and clear, pass the names that matter as keywords, and split anything over 30 seconds.
Seven languages, one clear voice, nothing uploaded. For meetings and other languages, reach for Whisper.
Revision note. Written October 9, 2026, while Whistle sits near the top of Hugging Face's trending list. If your phone is full of voice notes you never play back, tonight is a good night to let a very small model listen to them for you.
