Run Cactus Whistle Locally on Windows 11, Mac and Kali Linux: Free Offline Speech-to-Text Install Guide (16.9 MB, Beats Whisper Base)

Logeshwaran
—
Run Cactus Whistle Locally on Windows 11, Mac and Kali Linux: Free Offline Speech-to-Text Install Guide (16.9 MB, Beats Whisper Base)

Cactus Whistle is a free, open speech-to-text model that fits in a single 16.9 MB file and installs locally on Windows 11, Mac or Kali Linux in one command, running on an ordinary processor with no graphics card and no internet. It was released on October 2, 2026 under Apache 2.0 and is trending on Hugging Face this week. The surprise is the comparison its makers invite: Whistle is 8.6 times smaller than OpenAI's Whisper base, the model most "free transcription" tools quietly run, reaches its first word 6.6 times faster, and on the two standard LibriSpeech tests it makes fewer errors: 4.31% of words wrong on clean audiobook speech, 10.49% on the harder set. It is not better everywhere, and this page shows exactly where Whisper still wins. The catches are real too: Whistle hears seven languages, English, German, French, Spanish, Italian, Dutch and Polish, it listens to 30 seconds at a time, and it is a model with a Python package, not an app with a record button.

Jake runs a phone repair shop and spends an hour a day driving between suppliers. That hour is when he remembers things, so his phone is full of voice notes: "order three iPhone 17 screens," "call Mrs. Patel about the Pixel," "check why the Galaxy came back." He never listens to them again, because listening takes as long as recording. His friend Ethan, a developer, had one condition before helping: nothing recorded in Jake's car would be uploaded anywhere. This page is what they built. It covers what Whistle is in plain English, the honest benchmarks against Whisper and Moonshine, the free dictation your computer already has, how to install Whistle on Windows, Mac and Kali Linux, how to transcribe recordings longer than 30 seconds, how Jake's voice notes became a daily to-do list, how Whistle turns a spoken command into an action with its sibling model Needle, the mistakes that make it look bad, and what the cloud alternatives cost.

⚡ Quick Answer

• What is Cactus Whistle? → A 16.9 MB open speech-to-text model for seven languages that runs on any CPU, with word timestamps and keyword biasing. What it is.

• Is it better than Whisper? → Smaller and faster, ahead on five of eight published tests, behind on three, including meetings. The numbers.

• Install → pip install cactus-needle, then needle.transcribe("clip.wav")["text"]. Steps for Windows, Mac and Kali.

• Recordings over 30 seconds? → split them into chunks first; a six-line script does it. Long recordings.

• Just want to dictate? → Windows 11 has voice typing built in (Win+H). Whistle is for transcribing recordings and building things. Free built-in options.

No GPU, no account, no upload. The model is smaller than a single photo from a modern phone.

New to running AI on your own machine? Our plain-English series on running AI locally explains the words, models, parameters and quantization, in about ten minutes. Whistle is one of the gentlest first models there is, so you can also just follow the steps below.

What Cactus Whistle is, in plain English

Speech-to-text, also called transcription or voice-to-text, means turning recorded speech into written words. For years the free way to do it on your own computer was OpenAI's Whisper, released in 2022, which is excellent and also heavy: its smallest useful version is a 145 MB download that wants a decent processor, and its best versions want a graphics card with 10 GB of memory.

Whistle comes from Cactus Compute, a small company that builds AI for phones, watches, cars and the chips inside smart-home devices. Its whole philosophy is "make it tiny and make it run anywhere," and Whistle is that philosophy applied to speech. The entire model is one 16.9 MB file, stored at 2 to 4 bits per number, and it runs on the same small C++ engine as Cactus's other model, Needle, with no dependencies and no graphics card.

Run Cactus Whistle Locally on Windows 11, Mac and Kali Linux: Free Offline Speech-to-Text Install Guide (16.9 MB, Beats Whisper Base)

Whistle does three jobs, all on your device:

  • Transcription: 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish. It detects the language unless you name it. Silence returns an empty transcript rather than an invented sentence, which matters more than it sounds; many speech models hallucinate words into quiet audio.
  • Word timestamps: every word with its start time, end time and probability, taken from the model's own attention, so an app can highlight words as they are spoken, jump to a word, or cut a recording on a word.
  • Speech embedding: the raw encoder output, one row per 80 milliseconds, for matching and searching audio without ever writing a transcript.

One feature deserves its own paragraph: keyword biasing. You hand Whistle a list of the names, places and product words your users actually say, and it favors them while it decodes. Every speech model mangles proper nouns; "Mrs. Patel" becomes "Mrs. Battle," and "Pixel 10" becomes "pixel ten" or worse. Giving Whistle Jake's supplier and product names was the single change that made his transcripts usable.

The facts:

  • Size: one 16.9 MB file, whistle.cact.
  • Input: 16 kHz mono, up to 30 seconds per pass.
  • Languages: English, German, French, Spanish, Italian, Dutch, Polish.
  • Design: a log-mel front end and a convolutional stem feed an audio encoder; a Needle-shaped decoder reads it through gated cross-attention at every layer. The decoder is "laddered," so every depth from 2 layers up is a working model, chosen at load time with --audio-depth. Shallower is faster and slightly less accurate.
  • Speed: 11.1 ms to the first word and 1,319 tokens per second on an Apple M4 Pro processor, on a 10-second clip. The first-word delay grows with clip length: 5.9 ms at 5 seconds, 11.1 ms at 10, 36.3 ms at 30.
  • License: Apache 2.0, free for personal and commercial use.
  • Runs on: Windows (x64 and ARM), Mac, Linux (x86-64, ARM64, ARMv7, RISC-V, even the MIPS chips in cameras and routers), Android, iOS, Apple Watch and TV, and in the browser through WebAssembly.

What Whistle is not

  • Not an app. There is no window with a record button. You use it from Python, from a command line, or inside something you build. The needle whistle playground command does give you a quick microphone demo.
  • Not for every language. Seven European languages. For Hindi, Tamil, Chinese, Japanese, Arabic and the other 90-odd languages Whisper handles, Whisper or a cloud service is still the answer.
  • Not long-form by itself. Thirty seconds per pass. Longer recordings are split into pieces first, which this page shows how to do.
  • Not a meeting transcriber. It does not tell speakers apart, and on the AMI meeting benchmark it gets about one word in four wrong, as its own chart shows.

Whistle vs Whisper vs Moonshine: the honest numbers

Cactus published word error rates, the percentage of words a model gets wrong, on nine standard test sets, measured over 86,174 recordings and scored with Whisper's own text normalizers. The Whisper and Moonshine figures are the ones those models' authors published, not Cactus's own runs of them. Lower is better.

Test set (what it is)Whistle word error rateWho is ahead
LibriSpeech test-clean (clear audiobook reading)4.31%Whistle, ahead of Whisper base
LibriSpeech test-other (harder, accented reading)10.49%Whistle
SPGISpeech (earnings calls, clean)7.65%Whistle, ahead of Moonshine
Earnings-22 (earnings calls, varied accents)19.01%Whistle, ahead of Moonshine
AMI (meeting room recordings)26.07%Whisper base
AMI cleaned22.87%Only Whistle reported
TED-LIUM (TED talks)7.61%Whisper base
FLEURS (read speech, average of the seven languages)21.4%Whistle
MLS (audiobooks, six languages, no English)24.9%Whisper base

Cactus's own summary is fair: Whistle is ahead on LibriSpeech clean and other, on SPGISpeech, on Earnings-22 and on the FLEURS average; Whisper base is ahead on TED-LIUM, on AMI and on the MLS average. Moonshine is English-only, so it has no bars on the multilingual sets, and Whisper never published SPGISpeech, Earnings-22 or AMI-cleaned numbers. Cactus also checked that none of the test audio appears in Whistle's training data, by comparing audio checksums and speaker IDs, which is more than most model cards bother to say.

Read the pattern, not the wins. Whistle is strongest on clear speech from one person, which is exactly what a voice note, a dictated letter or a customer call is. It is weakest on meetings and on non-English audiobooks, where a 22 to 26% error rate means a transcript you can search but would not publish. And remember that Whisper base is Whisper's second-smallest model; Whisper small, medium and large are more accurate than base and much larger than Whistle. Nobody is claiming a 16.9 MB file beats a 1.5 GB one.

Size and speed

WhistleWhisper baseMoonshine tiny v2
Size on disk16.9 MB145.3 MB41.9 MB
Time to first word, 10 s clip11.1 ms73.2 ms22.8 ms
Decode speed, tokens per second1,319266262
Precision2 to 4 bitfp32 on CPUint8

All three were measured on an Apple M4 Pro, each on its official runtime at default settings: Whistle's C++ engine with 5 beams, the openai-whisper package, and moonshine-voice. A 10-second clip decoding at 1,319 tokens per second means the transcript is finished almost before you have lifted your finger off the key. On an older or cheaper processor, expect slower numbers in the same proportions.

Where Whistle sits in the Whisper family

Since "which Whisper is best" is a common question, here are OpenAI's own Whisper sizes, so you can see what Whistle is being compared with:

Whisper modelParametersMemory neededRelative speed
tiny39Mabout 1 GBabout 10x
base (the one Whistle is compared with)74Mabout 1 GBabout 7x
small244Mabout 2 GBabout 4x
medium769Mabout 5 GBabout 2x
large1,550Mabout 10 GB1x

Whisper is MIT-licensed and handles about a hundred languages, which is why it remains the default for anything multilingual or long-form. Whistle's pitch is different: the accuracy of Whisper base on clear speech, in a file 8.6 times smaller, fast enough for a watch.

Do you need a model at all? Free speech-to-text you already own

Before installing anything, a quick honesty check, because most "speech to text free" searches are answered by a key you already have:

  • Windows 11 voice typing: press Win+H in any text box and talk. It punctuates, it is free, and it works offline for the main languages once the speech pack is downloaded. Our old guide to Windows speech recognition and voices covers the settings behind it.
  • Mac: Dictation in System Settings, with a keyboard shortcut you choose.
  • Phones: the microphone key on the keyboard, on both Android and iPhone.
  • Word and Google Docs: both have a Dictate or Voice typing button.

Those are for dictating: you talk, text appears. They are poor at transcribing: turning a recording you already have into text, in bulk, with timestamps, inside your own program. "Transcribe audio to text" is one of the most searched phrases in this whole subject, and it is the job Whistle is for: a folder of voice notes, customer calls, lecture clips or interviews, processed on your machine, with the names you care about spelled right.

How to install Whistle locally on Windows 11, Mac and Kali Linux

Whistle ships inside the cactus-needle Python package, which also carries the engine. There is nothing else to install: no PyTorch, no CUDA, no graphics driver. The package covers Windows x64 and ARM, Mac on Apple Silicon, and Linux on x86-64 and ARM64.

Step 1: a virtual environment

Make a virtual environment first, so this install cannot disturb anything else, and because Kali and other recent Linux systems refuse system-wide pip installs with "externally-managed-environment." Our guide to that pip error explains why that refusal is a feature.

python3 -m venv ~/whistle
source ~/whistle/bin/activate
pip install cactus-needle

On Windows, the activate line is whistle\Scripts\activate. If you want to record from a microphone or feed in files that are not already 16 kHz WAV, install the extra: pip install "cactus-needle[mic]", with the quotes, because Kali's zsh reads square brackets as a file pattern without them.

Step 2: transcribe a clip

import needle

result = needle.transcribe("clip.wav")
print(result["text"])
# turn off the kitchen lights

The first call downloads the engine and the 16.9 MB model once and caches them in ~/.cache/cactus-needle/v3/. Every call returns the text, the detected language, the milliseconds to the first token and the decoder's tokens per second. Three options cover most needs:

r = needle.transcribe(
    "note.wav",
    language="en",                                  # skip detection when you know the language
    keywords=["Patel", "Pixel 10", "iPhone 17", "Galaxy S26"],   # names it must get right
    word_timestamps=True,                           # start, end and probability per word
)
print(r["text"], r["language"], r["ttft_ms"], r["decode_tps"])
for w in r["words"]:          # each entry carries the word's start, end and probability
    print(w)

needle.transcribe() is the entry point for speech; the model loads once per process and stays loaded, so a loop over many clips does not pay the start-up cost again. needle.Whistle() gives you the model as an object for the speech-embedding feature, embed(audio), or to hold a particular weights file. Two commands are handy on day one. needle whistle playground transcribes from your microphone, press Enter to record and Enter again to stop, and takes --language en, --keywords "..." and --word-timestamps. needle whistle compare clip.wav runs the same clip through Whistle, Whisper tiny, Whisper base and Moonshine tiny v2 and shows each one's timing; it needs pip install "cactus-needle[mic,compare]".

The command-line engine (no Python at all)

For scripts, servers or small boards, the bare engine is one small binary per platform plus the model file:

needle download linux-x86_64      # or macos-arm64, windows-x86_64, linux-arm64 for a Raspberry Pi
needle download whistle
./linux-x86_64/needle --model whistle.cact --audio clip.wav --audio-word-timestamps

On Windows the binary is needle.exe. The bare engine reads no environment variables; every behavior is a compiled default or an explicit flag, so the same file gives the same transcript on every machine, which is a quiet gift when you are debugging.

Fully offline, including air-gapped machines

Inference never touches the network; only the first download does. For a computer with no internet at all, the Python package's own offline notes give the recipe:

  1. On a connected machine, run pip download cactus-needle -d wheels to collect the package, and needle download whistle --out models to fetch the speech weights. needle fetch or needle download needle3 pulls the engine and Needle's weights if you want those too.
  2. Copy the wheels and models folders, and the cache folder ~/.cache/cactus-needle/v3/, to the offline machine.
  3. Install there with pip install --no-index --find-links wheels cactus-needle.
  4. Point the package at the weights with the environment variable NEEDLE_WHISTLE_WEIGHTS=/path/to/whistle.cact, or pass weights= to needle.Whistle().
  5. Set HF_HUB_OFFLINE=1, so a missing file fails at once with a clear message instead of hanging on a download that can never finish.

That is the setup for a clinic laptop, a courtroom, a factory floor or any place where the rule is "nothing leaves the building." The whole kit is a few tens of megabytes.

Transcribe audio to text from any source: phone notes, WhatsApp, Zoom, video

"Transcribe audio to text" covers a dozen file types, and almost none of them arrive as 16 kHz mono WAV. ffmpeg converts all of them with one pattern, -ac 1 -ar 16000, which means one channel at 16,000 samples per second. Install it once: on Windows winget install Gyan.FFmpeg, on Mac brew install ffmpeg, on Kali and Debian sudo apt install ffmpeg.

SourceUsual fileConvert with
iPhone or Android voice memo.m4affmpeg -i note.m4a -ac 1 -ar 16000 note.wav
WhatsApp or Telegram voice note.opus or .oggffmpeg -i note.opus -ac 1 -ar 16000 note.wav
Zoom, Teams or Meet recording.m4a or .mp4ffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 meeting.wav (-vn drops the video)
A video file or screen recording.mp4, .mkv, .movSame as above, with -vn
Windows Sound Recorder.m4a (or .wav at 44.1 kHz)ffmpeg -i rec.m4a -ac 1 -ar 16000 rec.wav
A dictation machine or old recorder.mp3, .wma, .dssSame command; ffmpeg reads nearly everything except some proprietary .dss files

Remember the meeting caveat before you transcribe a Zoom call: several voices through laptop microphones is Whistle's weakest case. One person talking into a phone is its best.

Transcribe a voice note on Windows 11, step by step

For anyone who has never opened a terminal, here is the whole path from a phone recording to text on a Windows laptop:

  1. Install Python from python.org, ticking "Add python.exe to PATH" in the installer, and ffmpeg with winget install Gyan.FFmpeg in a terminal. Close and reopen the terminal afterward so it finds both.
  2. Make a folder, for example C:\Transcribe, and copy the voice note into it.
  3. Open a terminal in that folder: type cmd in the folder's address bar in File Explorer and press Enter.
  4. Create and activate a virtual environment: python -m venv venv, then venv\Scripts\activate.
  5. Install Whistle: pip install cactus-needle.
  6. Convert the note: ffmpeg -i note.m4a -ac 1 -ar 16000 note.wav. If it is longer than 30 seconds, use the splitting command in the next section instead.
  7. Transcribe: python -c "import needle; print(needle.transcribe('note.wav')['text'])". The first run downloads the 16.9 MB model; later runs are instant.

Mac and Kali users follow the same seven steps with source venv/bin/activate in step 4 and the package manager of their system in step 1. If your microphone itself is the thing not working, our fix for USB audio Code 10 errors on Windows 11 is the page you need before any model.

How to transcribe recordings longer than 30 seconds

Whistle listens to 30 seconds at a time, so a 12-minute voice note has to be cut up first. The simplest reliable tool is ffmpeg, which is free on every platform. This one command converts any recording, including the .m4a files phones produce, to mono 16 kHz WAV and splits it into 28-second pieces:

mkdir -p chunks
ffmpeg -loglevel error -i note.m4a -ac 1 -ar 16000 -f segment -segment_time 28 chunks/part%03d.wav

Then transcribe the pieces in order and join them:

from pathlib import Path
import needle

parts = []
for wav in sorted(Path("chunks").glob("part*.wav")):
    r = needle.transcribe(str(wav), language="en", keywords=["Patel", "Pixel 10", "iPhone 17"])
    parts.append(r["text"])
print(" ".join(parts))

One honest limitation: a cut every 28 seconds will sometimes land in the middle of a word, and that word may come out wrong or missing. For voice notes and dictation it hardly matters. For an interview you will quote, split on silence instead: ffmpeg's silencedetect filter lists the quiet moments, and cutting at the nearest pause under 30 seconds keeps every word whole. Whistle's word timestamps also tell you when a chunk's last word ends early, which is a sign the cut fell on speech.

A real example: Jake's voice notes become a to-do list

Here is the whole of what Ethan built, in one script that runs on Jake's laptop every evening. It watches the folder where his phone syncs voice notes, converts and chunks anything new, transcribes it with Jake's product and supplier names as keywords, and appends the text to a dated file. Nothing leaves the laptop.

import datetime as dt, subprocess, tempfile
from pathlib import Path
import needle

NOTES = Path.home() / "VoiceNotes"           # where the phone syncs .m4a files
DONE = NOTES / "done"; DONE.mkdir(exist_ok=True)
KEYWORDS = ["Patel", "Pixel 10", "iPhone 17", "Galaxy S26", "Mobile Parts Direct", "screen", "battery"]

for m4a in sorted(NOTES.glob("*.m4a")):
    with tempfile.TemporaryDirectory() as tmp:
        subprocess.run(["ffmpeg", "-loglevel", "error", "-i", str(m4a), "-ac", "1", "-ar", "16000",
                        "-f", "segment", "-segment_time", "28", f"{tmp}/p%03d.wav"], check=True)
        text = " ".join(needle.transcribe(str(w), language="en", keywords=KEYWORDS)["text"]
                        for w in sorted(Path(tmp).glob("p*.wav")))
    day = dt.date.fromtimestamp(m4a.stat().st_mtime)
    with open(NOTES / f"{day}.md", "a", encoding="utf-8") as f:
        f.write(f"\n- [ ] ({m4a.stem}) {text.strip()}\n")
    m4a.rename(DONE / m4a.name)

Each note becomes a checkbox line in that day's Markdown file, so Jake's "order three iPhone 17 screens" is now something he can tick off. The keyword list was the difference between useful and not: before it, "Mrs. Patel" came out three different ways in one week. After it, every mention matched. The car noise still costs a few words on the motorway, which is the AMI lesson in miniature: clear, close speech is where a small model shines.

Once the notes are text, the next step is searching them by meaning. Our guide to EmbeddingGemma 2 shows how to build a private search box over transcripts and the original audio, also offline.

From spoken words to actions: Whistle plus Needle

The reason Cactus built Whistle on the same engine as Needle becomes clear when you load both at once. Needle is Cactus's tiny tool-calling model: give it the functions your app exposes and it picks the right one and fills in the arguments. With Whistle beside it, a clip goes in and a function call comes out, in one pass, with the transcript attached:

needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav
{"function_calls": [{"name": "set_lights", "arguments": {"room": "kitchen", "on": false}}],
 "confidence": 0.94,
 "audio_text": "turn off the kitchen lights",
 "audio_language": "en"}

That is a complete offline voice assistant in two files totaling under 50 MB, and it runs on a Raspberry Pi. Needle's full model is 29 MB, and like Whistle it is laddered, so a 4-layer version of about 8 MB works for simple command sets. A request no tool covers returns an empty list rather than a guess, and the confidence number lets your app act, confirm or refuse. For a shop, "book a screen repair for Thursday" becoming a calendar entry without a cloud account is the kind of thing that used to need a product team. If the decisions after the transcript are the hard part, our guide to decision models in Ollama covers the other small models built for exactly that.

What word timestamps and embeddings are for

Two of Whistle's three jobs are easy to overlook, so here is what people build with them.

Word timestamps give every word a start time, an end time and a probability. With those you can make captions for a video, since each caption line is just a run of words and their times; build a "click a word to jump there" player for lectures and interviews; cut the dead air out of a recording automatically, because the gaps between word end-times are the silences; and spot the words the model was unsure of, by their low probability, so a human checks only those. Jake's version was simpler: the words with low probability in his transcripts were almost always supplier names he had not yet added to the keyword list.

Speech embeddings are the encoder's output before any words are decoded: one row of numbers per 80 milliseconds of audio. Two clips of the same phrase produce similar rows, which lets an app match or search audio without ever transcribing it: find every voice note that mentions the same thing, detect a repeated jingle, or check whether a recorded phrase matches a stored one. It is a research-flavored feature, but it comes free with the same 16.9 MB file.

Choosing a decoder depth

Whistle's decoder is laddered: every depth from 2 layers upward was trained as a working model, and the command-line engine picks one at load time with --audio-depth. The encoder always runs all eight of its blocks, so the saving is only in the decoder. Cactus publishes no accuracy figures per depth, so the honest advice is: on a laptop, use the default and never think about it; on a watch or a microcontroller, start shallow, run needle whistle compare on ten of your own clips, and go one step deeper until the errors stop falling.

Why Whistle sounds wrong: the common mistakes

Speech models rarely crash. They return confident text with the wrong words in it, so these are worth reading before you judge the model.

Audio at the wrong sample rate

Whistle wants 16 kHz mono. Phone recordings are usually 44.1 or 48 kHz stereo. The base package accepts only 16 kHz WAV or raw samples; either install the [mic] extra, which handles other rates, or convert with the ffmpeg command above. Feeding 48 kHz audio into a 16 kHz model produces slowed, garbled text, not an error.

Clips over 30 seconds

Anything past 30 seconds is not transcribed. Split first.

An unsupported language

Seven languages only. Hindi, Tamil, Chinese, Japanese, Arabic, Portuguese and the rest come out as nonsense or as the nearest supported language. Use Whisper or a cloud service for those.

Meetings and crowded rooms

On the AMI meeting benchmark Whistle gets 26% of words wrong, and it does not separate speakers. Use it for one clear voice near the microphone; use a larger Whisper or a dedicated meeting tool for the conference room.

Leaving out keywords

Names and product words are where every small speech model fails. Pass the ones that matter in keywords=[...]. Do not pass a dictionary: a short list of the words that must be right works better than a long list of words that might appear.

Running at the shallowest ladder depth

Whistle's decoder runs at any depth from 2 layers up, chosen with --audio-depth. Shallower is faster and less accurate. If a tutorial set a tiny depth for a microcontroller, raise it on a laptop.

"needle: command not found"

The needle command lives inside the virtual environment you installed it in. Activate the environment first, or call it by its full path.

Expecting an app

Whistle is a model and a package. For a window with a record button, use Windows voice typing or a dictation app, or build the small script above once and reuse it.

Judging it on faded, distant audio

A note recorded with the phone in a pocket on a motorway will be rough with any model. Record close and clear, and the small model keeps up with the big ones.

Whistle vs Whisper, whisper.cpp, Moonshine, Windows voice typing and the cloud

OptionWhat it isChoose it when
Whistle16.9 MB open model, 7 languages, CPU only, Apache 2.0Short clear clips, keyword accuracy, tiny devices, voice commands with Needle
Whisper (OpenAI)Open models from 39M to 1.55B parameters, about 100 languages, MITOther languages, long recordings, best accuracy with a GPU
whisper.cppWhisper rewritten in C++ for CPUs and phonesYou want Whisper's languages without PyTorch
Moonshine tiny v241.9 MB English-only model for devicesEnglish only on a small device; Whistle is ahead on the shared tests
Windows voice typing (Win+H)Built-in dictationYou want to talk into a document right now
MAI-Transcribe-2 (Microsoft)Cloud API, 60 languages, 2.5% error on its streaming testBest accuracy and many languages, and uploading is acceptable
Amazon TranscribeAWS's managed service, billed per secondYou already run on AWS and want speaker labels, redaction or call analytics

A fair ten-minute bake-off on your own recordings

Benchmarks are other people's audio. Before you commit, test on yours. Cactus ships the tool for it:

  1. Collect ten clips of under 30 seconds that represent your real use: your voice, your room, your names and product words.
  2. Write down what was said in each, carefully. This is your answer key.
  3. Install the comparison extras: pip install "cactus-needle[mic,compare]".
  4. Run needle whistle compare clip.wav for each clip. It prints Whistle, Whisper tiny, Whisper base and Moonshine tiny v2 side by side, with each one's timing.
  5. Count the wrong words per model against your answer key, then run Whistle again with your keywords list and count once more.
  6. Decide on the whole picture: accuracy on your audio, speed on your machine, the languages you need, and whether anything may leave the building.

What the cloud costs, for comparison

Whistle costs nothing per hour of audio. For scale, here are Amazon Transcribe's on-demand prices in US East (N. Virginia), from AWS's official price list published September 11, 2026: standard batch transcription is $0.0001 per second, which is $0.36 per hour of audio, and streaming is $0.0001667 per second, $0.60 per hour. Microsoft's MAI-Transcribe-2 is $0.10 an hour for recorded audio and $0.54 for streaming through the end of 2026. Jake's five hours of voice notes a month would cost under $2 on either, so price is not the argument for Whistle. Privacy, no account, no internet and the ability to run on a $50 board are. If you do want a managed service on AWS, our explainer on Hugging Face models on AWS covers hosting open models there as well.

Cactus Whistle: frequently asked questions

What is Cactus Whistle?

Whistle is a free, open speech-to-text model from Cactus Compute, released October 2, 2026 under Apache 2.0. It is a single 16.9 MB file that transcribes English, German, French, Spanish, Italian, Dutch and Polish on any CPU, with word timestamps and keyword biasing.

Is Whistle better than Whisper?

On five of eight shared tests, yes against Whisper base: both LibriSpeech sets, SPGISpeech, Earnings-22 and FLEURS. Whisper base is ahead on TED-LIUM, AMI meetings and MLS. Larger Whisper models are more accurate than base and far larger than Whistle.

How do I install Whistle?

Create a Python virtual environment, run pip install cactus-needle, then call needle.transcribe("clip.wav")["text"]. The first call downloads the engine and the 16.9 MB model. Add the [mic] extra for microphone input and other sample rates.

Does Whistle need a GPU?

No. It runs on the processor alone, at 11.1 ms to the first word and 1,319 tokens per second on an Apple M4 Pro. It also runs on Raspberry Pi boards, phones, watches and in the browser.

What languages does Whistle support?

English, German, French, Spanish, Italian, Dutch and Polish. It detects the language unless you set one. For other languages, use Whisper or a cloud service.

How long a recording can Whistle transcribe?

Up to 30 seconds in one pass. For longer audio, split it into pieces with ffmpeg, transcribe each piece in order and join the text. Splitting on silence keeps words whole.

Can Whistle transcribe audio to text offline?

Yes, completely. After the first download, nothing is sent anywhere. Set HF_HUB_OFFLINE=1 and copy the cache folder for a machine with no internet at all.

What audio format does Whistle need?

16 kHz mono. Convert phone recordings with ffmpeg -i note.m4a -ac 1 -ar 16000 note.wav, or install the [mic] extra, which handles other sample rates.

Does Whistle give word timestamps?

Yes. Pass word_timestamps=True and each word returns with its start time, end time and probability, read from the decoder's own attention.

What is keyword biasing in Whistle?

A list of names, places and product words you pass with keywords=[...]. Whistle favors them while decoding, so proper nouns that small models usually mangle come out right.

Is Whistle free for commercial use?

Yes. It is released under the Apache 2.0 license, with no revenue limit.

How do I run Whistle locally on Windows 11?

Install Python, create a virtual environment with python -m venv venv and venv\Scripts\activate, run pip install cactus-needle, convert your recording to 16 kHz mono WAV with ffmpeg, and call needle.transcribe("note.wav"). The 16.9 MB model downloads once; everything after that is offline.

Can Whistle run on Windows?

Yes. The cactus-needle package supports Windows x64 and ARM, and the standalone engine is needle.exe in the windows-x86_64 or windows-arm64 folder.

Can Whistle run on Kali Linux or a Raspberry Pi?

Yes. On Kali, install cactus-needle inside a virtual environment. For a Raspberry Pi, download the linux-arm64 engine folder and run the needle binary with whistle.cact.

Does Whistle separate speakers?

No. It transcribes one stream of speech without speaker labels, and its meeting-room accuracy is about 74%. For meetings with several speakers, use a larger Whisper or a cloud service with diarization.

What is Needle, and how does it work with Whistle?

Needle is Cactus's 8 to 29 MB tool-calling model. Loaded with Whistle in the same engine, a voice clip goes in and a function call with arguments comes out, plus the transcript, making an offline voice assistant.

How does Whistle compare with Amazon Transcribe?

Whistle is free and offline, seven languages, no speaker labels. Amazon Transcribe is a managed service at $0.36 per hour for batch and $0.60 for streaming in US East, with many languages, speaker labels and redaction.

By the end of the week Jake had stopped dreading his own voice notes. Each evening a Markdown file appeared with the day's reminders as checkboxes, Mrs. Patel's name spelled correctly in every one, and nothing he had said in the car had gone further than the laptop under the counter. "It's smaller than a photo," he said, looking at the 16.9 MB file, as if it might be a trick. It is not a trick. It is what happens when a company builds for watches and cameras first, and laptops turn out to be the easy case.

 If you keep one line from this page

Record close and clear, pass the names that matter as keywords, and split anything over 30 seconds.

Seven languages, one clear voice, nothing uploaded. For meetings and other languages, reach for Whisper.

Revision note. Written October 9, 2026, while Whistle sits near the top of Hugging Face's trending list. If your phone is full of voice notes you never play back, tonight is a good night to let a very small model listen to them for you.

Related