Run Qwen3.8-27B Locally on Windows 11, Mac and Kali Linux with an 8 GB GPU: Underdog Saluki 27B GGUF (7.89 GB), Ollama and llama.cpp Install, Real RAM Requirements, GSQ-RCO Quants and AWS Costs
Underdog Saluki 27B is Alibaba's Qwen3.8-27B, the 54 GB open model that tops most "best local AI" lists, squeezed into a single 7.89 GB file that runs in ordinary llama.cpp with no fork and no patch. It appeared on Hugging Face on October 8, 2026 from a small outfit called ConwayResearch, and within two days it had 15,000 downloads. Its maker's claim is narrow and interesting: the shrunken model keeps tool calling intact, the thing that lets a model actually do work (look up a record, call a function, run a command) rather than just chat, and on the maker's own 120-task test it scores 88 against the full-size model's 84. Here is the shock that the file size hides: a 7.89 GB file does not run on an 8 GB graphics card. The honest requirement for any 2-bit cut of this model is 9 to 11 GB of memory, graphics card and system RAM combined, and the math below shows where every gigabyte goes. The second surprise is that Saluki's secret is not Saluki's at all: it is built on a quantization recipe from a university lab in Austria that you can download directly, in four sizes, with better-documented numbers.
Jake runs a phone repair shop, and his laptop has an 8 GB graphics card and 32 GB of RAM, which is a very common combination and a very awkward one for local AI. The 18 GB Ollama download of Qwen3.8-27B crawls on it. His friend Ethan, a developer, has been trying to give Jake an assistant that can look things up in the shop's own stock list and ticket system instead of making answers up, which means a model that can call tools reliably, on hardware Jake already owns. This page is what they found: what Saluki is and what it is built on, the real memory requirement, how to install it on Windows 11, Mac and Kali Linux, how to turn tool calling on and test it, what the benchmarks say and hide, how it compares with Ternary Bonsai 2 and the full model, what it costs to run on AWS instead, and the honest list of things it is worse at.
New to running AI on your own machine? Our plain-English series on running AI locally explains models, parameters, GGUF files and quantization in about ten minutes. You can follow this page without it; every step starts from zero. If you want the full-size model and have the hardware for it, our Qwen3.8-27B install guide is the page for that.
What Underdog Saluki 27B is, in plain English
Start with the model underneath. Qwen3.8-27B is Alibaba's dense 27-billion-parameter model from August 2026: a single network rather than a mixture of experts, 64 layers in a hybrid design that mixes a fast linear-attention block called Gated DeltaNet with ordinary attention, a 262,144-token context window, native vision (it reads images and video through a separate projector file), tool calling, a thinking mode that is on by default, and an Apache 2.0 license. In full precision it is 54 GB of weights, which is why most people meet it as an 18 GB 4-bit download in Ollama and still need a 24 GB graphics card to enjoy it.
Saluki is that model stored in roughly 2.3 bits per weight instead of 16, which is how 54 GB becomes 7.89 GB. Squeezing a model that hard usually ruins it. What makes 2-bit survivable in 2026 is non-uniform quantization: instead of storing every tensor at the same precision, a method decides, tensor by tensor, which parts of the network can tolerate 2 bits and which need 3 or 4, under a total size budget. The method Saluki is built on is GSQ-RCO from ISTA-DASLab, the machine-learning lab at the Institute of Science and Technology Austria, which published its own GGUF files of Qwen3.8-27B on August 28, 2026; those files have been downloaded about 1.5 million times since.
ConwayResearch took that work and did two things. It produced a single "IQ2-mix" file that comes in under 8 GB, smaller than the lab's smallest 8.4 GB file, and it says the result was "tuned to keep tool calling intact." The model card does not say how: there is no description of the tuning method, the data, or which tensors were changed. What it offers instead is a benchmark table, which we go through below, and a NOTICE file crediting Qwen and ISTA-DASLab. The organization itself has no biography on Hugging Face beyond the line "Built by Underdog." That is not a reason to distrust the file; it is a reason to read the numbers as a maker's numbers and test the model on your own tasks before relying on it.
Jake: "So who are these people?"
Ethan: "Honestly, nobody knows yet. But the file is a standard GGUF, the license is Apache, the base model is Alibaba's and the compression is a university's. The only part that is theirs is a tune they have not explained. We can measure whether it works; we just can't read how."
Three practical facts about the file before anything else:
- It is text-only on its own. Vision is a separate projector file,
mmproj-Underdog-Saluki-27B-1.0-F16.gguf(0.93 GB) or the Q8_0 version (0.63 GB), loaded alongside it. - Thinking is on by default, as in Qwen3.8, and can be switched off per request. For tool calling, the maker recommends switching it off.
- It runs in stock llama.cpp. That is the practical difference from Ternary Bonsai 2, the other tiny Qwen3.8-27B, which needs a patched llama.cpp from its maker; our Bonsai 2 guide explains that fork and why LM Studio and Ollama cannot load it.
Qwen3.8-27B on an 8 GB graphics card: the real memory math
This is the section the file size tempts you to skip. A GGUF file's size is the space its weights take on disk. Running the model needs that much memory plus room for the context (the conversation and documents the model holds in mind), plus the vision projector if you load it, plus a working buffer. Unsloth, which publishes its own Qwen3.8-27B files and has tested the family more than anyone, gives these totals for RAM and VRAM combined:
| Precision | Memory needed (VRAM + RAM, or unified) | Files at this level |
|---|---|---|
| 1-bit | 7 to 8 GB | Unsloth IQ1 cuts; Bonsai 2 PTQ1_0 (fork needed) |
| 2-bit | 9 to 11 GB | Saluki 7.89 GB; GSQ-RCO IQ2_XS 8.4 GB, IQ2_S 9.3 GB |
| 3-bit | 12 to 14 GB | GSQ-RCO IQ3_XXS 10.1 GB, IQ3_S 11.8 GB |
| 4-bit | 16 to 19 GB | Ollama's qwen3.8:27b (18 GB); Unsloth UD-Q4_K_XL |
| 8-bit | 31 GB | Q8_0 |
| BF16 (original) | 56 GB | The 54 GB safetensors |
Unsloth's rule of thumb is that VRAM plus RAM should roughly match the file size for a comfortable run, and its table adds 1 to 2 GB if you use the multi-token-prediction head for speed, and more for a long context. llama.cpp keeps the model running when the graphics card is full by holding the remaining layers in system RAM, which works and is slower, and it will even page to disk if RAM runs out too, which Unsloth describes as "much slower."
So, for Jake's laptop and the three other common setups:
| Your machine | What to expect with Saluki | Setting that matters |
|---|---|---|
| 8 GB graphics card + 16 GB or more RAM (Jake) | Runs. Most layers on the card, the rest in RAM. Noticeably slower than a card that holds it all, fine for an assistant that answers in sentences, slow for long thinking. | Lower -ngl until it loads without an out-of-memory error; start at 40 of the 64 layers and raise it. Keep the context at 8K to 16K. |
| 12 GB card | The sweet spot. Whole model on the card with a 32K context. | -ngl 99, the maker's default command. |
| 16 GB card or 16 GB Mac | Comfortable, with room for the vision projector and a long context. Consider the lab's 10.1 GB IQ3_XXS instead for better math. | Add --mmproj if you want images. |
| No graphics card, 16 GB RAM | Loads and runs on the processor alone, at a few tokens per second. Usable for short tool calls with thinking off; tedious for anything long. | Thinking off, short context, -t set to your physical core count. |
| 24 GB card | You do not need Saluki. Run the 4-bit full model from our Qwen3.8-27B guide, or the lab's IQ3_S for a smaller file with the same scores. | See the GSQ-RCO section. |
Two more numbers. Disk: 7.89 GB for the model, 0.63 to 0.93 GB for vision, and a few hundred megabytes for llama.cpp. Context: the hybrid design of Qwen3.8 keeps per-token memory low compared with older 27B models, which is why a 32K context fits beside the weights on a 12 GB card at all; with the 262K maximum you would be back in big-GPU territory. If you are choosing a laptop for this kind of work, our laptop guide for local AI explains why the graphics card's memory, not the processor, is the number to shop for.
Jake: "So 'under 8 GB' was never going to mean 'fits an 8 GB card.'"
Ethan: "It means the file fits. The model also needs a desk to work at. On your laptop the desk is in system RAM, which is why it will think out loud a bit slower than the videos you've seen."
How to run Underdog Saluki 27B on Windows 11 (llama.cpp, step by step)
The maker's route is llama.cpp, and it is also the fastest to set up, because llama.cpp ships prebuilt Windows packages. If you would rather have an app with a chat window, skip to the Ollama and LM Studio section; both can load this file.
- Get llama.cpp. Open the llama.cpp releases page on GitHub and download the latest Windows package for your hardware: the CUDA build for an NVIDIA card (it also needs the matching CUDA runtime package from the same release page), the Vulkan build for AMD or Intel graphics, or the plain CPU build. Unzip it to a folder such as
C:\llama. - Download the model. Either click the file on the Hugging Face page, or in PowerShell:
The maker's card uses the olderpip install -U "huggingface_hub[cli]" hf download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir C:\llama\modelshuggingface-cli downloadspelling; both commands work. - Start the server. This is the maker's command, with the model path filled in:
cd C:\llama .\llama-server.exe -m models\Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768--jinjaturns on the Qwen3.8 chat template, which is what makes tool calls and thinking work;-ngl 99puts every layer on the graphics card;-fa onenables flash attention;-c 32768sets a 32K context. On an 8 GB card, change-ngl 99to-ngl 40and-c 32768to-c 8192, then raise them while it still loads. - Talk to it. Open
http://127.0.0.1:8080in a browser for llama.cpp's built-in chat page, or point any app that speaks the OpenAI chat API at that address. - Add vision later, if you want it. Download
mmproj-Underdog-Saluki-27B-1.0-Q8_0.gguf(0.63 GB) to the same folder and add--mmproj models\mmproj-Underdog-Saluki-27B-1.0-Q8_0.ggufto the command. The F16 projector (0.93 GB) is the one in the maker's example; the Q8_0 version is smaller and the usual choice on tight memory.
If the first message takes minutes and the graphics card's memory stays empty, llama.cpp has fallen back to the processor, usually because the CPU build was downloaded or the driver is old. Our guide to why local AI ignores your GPU walks through the checks, which are the same for every runner.
Thinking on or off, and the settings that go with each
Qwen3.8 reasons before it answers unless told not to, and with reasoning effort set to "xhigh" by default it can reason at length. The maker recommends two configurations:
- General use and reasoning: thinking on, temperature 0.6, top_p 0.95, top_k 20. (Alibaba's own card suggests temperature 1.0 for thinking mode; the maker's lower figure is what its tests used.)
- Fast, direct tool calls: thinking off, temperature 0.
Thinking is switched off per request by adding "chat_template_kwargs": {"enable_thinking": false} to the request body, or for the whole server by starting it with --chat-template-kwargs '{"enable_thinking": false}'. The same switch takes "reasoning_effort": "medium" or "low" if you want thinking kept but shorter. In PowerShell the quotes need escaping: "{\"enable_thinking\": false}".
Tool calling with Saluki: turning the claim into a test
Tool calling (also called function calling) is the feature Saluki is built around, so here is what it means and how to see it work. You describe a function to the model, with a name, a plain-English description and its parameters. When a question needs that function, the model does not answer; it replies with a structured request to call the function with specific arguments. Your code runs the function, hands the result back, and the model writes the final answer from real data. That is the difference between an assistant that guesses Jake's stock levels and one that reads them.
With llama.cpp's server running and --jinja on, this is a complete test, using the OpenAI-compatible endpoint from Python:
import json, requests
STOCK = {"iphone 15 screen": 4, "pixel 9 battery": 0, "usb-c port s24": 7}
def check_stock(part):
return {"part": part, "in_stock": STOCK.get(part.lower(), 0)}
tools = [{
"type": "function",
"function": {
"name": "check_stock",
"description": "Look up how many of a repair part the shop has on the shelf.",
"parameters": {"type": "object",
"properties": {"part": {"type": "string", "description": "Part name, e.g. 'iPhone 15 screen'"}},
"required": ["part"]}}}]
messages = [{"role": "user", "content": "Do we have a Pixel 9 battery in stock?"}]
body = {"model": "saluki", "messages": messages, "tools": tools, "temperature": 0,
"chat_template_kwargs": {"enable_thinking": False}}
r = requests.post("http://127.0.0.1:8080/v1/chat/completions", json=body).json()
msg = r["choices"][0]["message"]
if msg.get("tool_calls"):
call = msg["tool_calls"][0]
args = json.loads(call["function"]["arguments"])
result = check_stock(**args) # the real lookup happens here
messages += [msg, {"role": "tool", "tool_call_id": call["id"], "content": json.dumps(result)}]
final = requests.post("http://127.0.0.1:8080/v1/chat/completions",
json={**body, "messages": messages}).json()
print(final["choices"][0]["message"]["content"])
else:
print("No tool call:", msg["content"])
What a correct run looks like: the first reply contains a tool_calls entry naming check_stock with {"part": "Pixel 9 battery"}, your function returns zero, and the final answer says the shop is out of Pixel 9 batteries. What a failure looks like: the model answers "Yes, we have it" without calling anything, or produces a call with the arguments in the wrong shape. Run the test ten times with different parts; that is a smaller version of the maker's own benchmark, and the only one that matters for your shop.
Two settings matter here. Thinking off, because reasoning before a tool call adds seconds and, on the maker's tests, did not help. Temperature 0, so the call comes out the same way every time. The maker notes that about a fifth of its parallel tool-call replies had "small formatting slips," so your code should tolerate an extra space or a trailing comma in arguments, and should retry once rather than crash. For a complete agent rather than a test, a coding agent such as OpenCode can point at this same server and use the model's tool calls to read and edit files.
Underdog Saluki 27B benchmarks: what 96 percent retention means and hides
The maker publishes one table, and it is a more honest table than most, because it says where its numbers came from. "Underdog Bench" is 120 tool-calling tasks taken from the Berkeley Function Calling Leaderboard (BFCL v4), frozen before any model was tested, run with thinking off at temperature 0. Where the full-size model was run by ConwayResearch, it used the same setup. Rows marked "public" are the full-size model's published scores from a different harness, which makes those comparisons softer.
| Test | Saluki (7.89 GB) | Qwen3.8-27B full (54 GB) | Reading |
|---|---|---|---|
| Underdog Bench, tool calling (of 120) | 88 | 84 (same harness) | Level; the maker itself says a few tasks is within run-to-run noise. Ternary Bonsai 2 scored 70. |
| BFCL v4 parallel tool calls (of 100) | 42 | 35 (same harness) | Better, and both are low: parallel calls are hard for every model this size. |
| SWE-bench Verified, 50 issues fixed | 30 | 33 (same harness) | Slightly behind on real coding tasks |
| IFEval, prompt-loose | 93.5 | 91.5 (public) | Follows instructions as well as the original |
| IFBench, prompt-loose | 72.7 | 71.0 (public) | Level |
| MBPP+, coding | 78.0 | 83.9 (public) | Six points down on short coding puzzles |
| MuSR, reasoning | 67.5 | 79.6 (public) | Twelve points down on multi-step reasoning stories |
| AIME 2025, avg of 4 | 79.2 | 96.7 (public) | Competition math is where 2-bit hurts most |
| AIME 2026, avg of 4 | 80.0 | 94.6 (public) | Same story |
The maker summarizes this as "96 percent average retention across nine benchmarks," and the arithmetic is fair, but averages flatter. Read it as three groups. On tool calling and instruction following, Saluki is level with or ahead of the original on the maker's harness, which is the claim the model exists to make. On coding, it is a few points behind. On math and long reasoning, it keeps about 82 to 85 percent of the original, which the maker states plainly in its limitations list, along with weakness on "letter-level instruction puzzles" (palindromes, vowel counting, alphabetical ordering), the formatting slips on parallel calls, and the tendency to reason at length with thinking on.
Two cautions of our own. First, 120 tasks is a small test, and the four-task gap on it is inside the noise the maker admits to; treat "beats the full model at tool calling" as "matches it." Second, the full-size model's "public" rows were produced by other people on other setups, so a two-point lead on IFEval is not evidence of anything. The honest headline is that a 7.89 GB file is level with a 54 GB one at the job it was tuned for and clearly worse at competition math, and for Jake's stock lookups that is exactly the right trade.
GSQ-RCO explained: the lab files Saluki is built on, and when to use them instead
This is the part most coverage of Saluki misses, and it may change which file you download. ISTA-DASLab's own GGUF files of Qwen3.8-27B are the source material, they come in four sizes, and the lab published far more measurement than ConwayResearch did.
The two methods in the name: GSQ (Gumbel-Softmax Quantization) is the lab's way of rounding each tensor's weights to few bits while keeping the result as close as possible to the original; RCO (Riemannian Constrained Optimization) is the allocator that decides which tensor gets which bit width under a total size budget. The result is that two files with the same average bits per weight can differ a lot in quality depending on where the bits went. The lab's table, measured against the full BF16 model and against Unsloth's dynamic quantizations of the same model:
| File | Bits / size | Wiki perplexity (lower = closer) | AIME25 | GPQA-Diamond | LiveCodeBench v6 |
|---|---|---|---|---|---|
| BF16 original | 16 / 53.8 GB | 7.05 | 100.00 | 89.90 | 85.71 |
| GSQ-RCO IQ3_S (lab's pick) | 3.50 / 11.8 GB | 7.07 | 100.00 | 89.39 | 85.71 |
| GSQ-RCO IQ3_XXS | 3.00 / 10.1 GB | 7.20 | 100.00 | 88.89 | 84.57 |
| GSQ-RCO IQ2_S | 2.75 / 9.3 GB | 7.39 | 100.00 | 86.36 | 82.29 |
| GSQ-RCO IQ2_XS | 2.50 / 8.4 GB | 7.69 | 96.67 | 84.85 | 76.57 |
| Unsloth UD-IQ2_S, for comparison | 2.49 / 8.4 GB | 8.02 | 86.67 | 76.26 | 72.00 |
Read the two 8.4 GB rows together: at the same size, the lab's allocation beats the standard dynamic quantization by ten points on AIME25 and more than eight on GPQA-Diamond. That is the whole argument for non-uniform quantization in one line, and it is why Saluki can be as small as it is without falling apart. The lab calls IQ3_S "task-lossless": the same AIME25 and LiveCodeBench scores as the original in a file one fifth the size. Every file also comes in an -mtp version, about 0.35 GB larger, carrying the multi-token-prediction head that llama.cpp can use for speculative decoding (predicting several tokens at once and checking them), with no change in quality.
Note that these are the lab's measurements of its own files with a different set of tests from ConwayResearch's; nobody has yet run Saluki and the lab's IQ2_XS through the same harness, and Saluki's AIME25 score (79.2, averaged over four runs) is not directly comparable with the lab's single-run 96.67. What the two tables agree on is the direction: math drops first and furthest as the bits go down.
So, which to download:
- Saluki (7.89 GB) when memory is the constraint and tool calling is the job: an 8 GB card with RAM behind it, a 12 GB card with a long context, a 16 GB Mac. It is the only one under 8 GB.
- GSQ-RCO IQ3_XXS (10.1 GB) or IQ3_S (11.8 GB) when you have a 16 GB card or a 24 GB Mac and want the model that matches the original on math and code. The lab's commands are the same shape as Saluki's:
hf download ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf --local-dir .and thenllama-serveras above. Ollama can pull them directly withollama run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUFand LM Studio finds them by searching the repository name. - Ternary Bonsai 2 (5.95 GB) only if you are willing to run PrismML's patched llama.cpp; it scored 70 of 120 on the maker's tool test against Saluki's 88.
Underdog Saluki 27B on a Mac
Apple silicon is the easy case here, because the graphics card and the RAM are the same pool: a 16 GB Mac has 16 GB for everything, which clears the 9 to 11 GB requirement with room for a modest context, and a 24 GB or 32 GB Mac runs the lab's bigger files instead. Install llama.cpp with Homebrew, then use the same commands as Windows with forward slashes:
brew install llama.cpp
pip install -U "huggingface_hub[cli]"
hf download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir ~/models
llama-server -m ~/models/Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768
Homebrew's llama.cpp is built with Metal, so -ngl 99 puts the model on the GPU cores. On a 16 GB Mac, close the browser tabs you are not using before the first load and drop -c to 16384 if the system starts swapping. Speed is set by memory bandwidth, so a MacBook Air is noticeably slower than a Pro of the same memory, and a Mac mini with 16 GB is a surprisingly good home for a model like this.
How to install Underdog Saluki 27B on Kali Linux
Kali is Debian underneath, and llama.cpp builds cleanly on it. Two Kali habits first: install Python tools inside a virtual environment, because Kali refuses system-wide pip install with the "externally-managed-environment" error (our guide to that error covers the clean fix), and check nvidia-smi works before building with CUDA.
sudo apt update && sudo apt install -y git cmake build-essential python3-venv
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON # NVIDIA; use -DGGML_VULKAN=ON for AMD/Intel, or nothing for CPU only
cmake --build build --config Release -j
python3 -m venv ~/hf && source ~/hf/bin/activate
pip install -U "huggingface_hub[cli]"
hf download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir ~/models
./build/bin/llama-server -m ~/models/Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768
The CUDA build needs the NVIDIA CUDA toolkit installed; on Kali the simplest path is the nvidia-cuda-toolkit package from the Kali repositories, matched to the driver you already have. If the build fails on the CUDA step, build without it first to prove the model works on the processor, then fix the toolkit. For a laptop with no discrete GPU, the CPU build runs Saluki at a few tokens per second with thinking off, which is enough to test tool calls.
Kali users tend to ask one more question: can this thing be the brain of a local security-tooling assistant that calls nmap or reads a scan file through tools? Technically yes, through exactly the function-calling pattern above, and it stays entirely on your machine, which is the point. Treat its output as a junior analyst's draft, not a finding, especially on anything numeric.
Running Saluki in Ollama or LM Studio instead of llama.cpp
Both apps can load this file, with one trap each.
Ollama. The Hugging Face page's auto-generated Ollama snippet reads ollama run hf.co/ConwayResearch/Underdog-Saluki-27B-1.0:F16. There is no F16 file in the repository; the only model file is the IQ2-mix GGUF, and Hugging Face's snippet simply assumed a tag. The reliable way is to download the file as above and import it with a Modelfile, which is Ollama's documented route for any GGUF:
# Modelfile, saved next to the downloaded .gguf
FROM ./Underdog-Saluki-27B-1.0-IQ2-mix.gguf
PARAMETER num_ctx 16384
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
ollama create saluki -f Modelfile
ollama run saluki
Ollama reads the chat template from the GGUF, so tool calling works through its OpenAI-compatible endpoint at http://127.0.0.1:11434/v1 with the same Python test as above, changing the address and the model name to saluki. Thinking is controlled with Ollama's think option or the /set nothink command in the chat.
LM Studio. Search for the repository name in the model downloader, or use the Hugging Face page's one-click link, download the IQ2-mix file, and load it. LM Studio shows you the layers-on-GPU slider and the context length, which makes it the gentlest way to find the settings an 8 GB card tolerates: slide GPU offload down until the load succeeds. Its local server exposes the same OpenAI-style endpoint for tool calling. If you are still deciding which runner to live with, our comparison of Ollama, LM Studio and llama.cpp sets out the trade-offs.
Qwen3.8-27B on AWS: the managed route, and what renting a GPU for Saluki costs
Sometimes the right machine is one you rent for an afternoon. There are two AWS stories here, and they are easy to confuse; our Bedrock vs SageMaker AI explainer sorts them out if the names are new.
- Amazon Bedrock does not offer Qwen3.8-27B. Its Qwen catalog is the earlier Qwen3 generation (Qwen3 32B, 235B, VL, Next and Coder). You can bring your own weights through Bedrock Custom Model Import, but not a GGUF; it takes full-precision safetensors.
- Amazon SageMaker JumpStart has offered Qwen 3.8-27B since August 27, 2026, deployable from the SageMaker console or the Python SDK. AWS's announcement describes it as a dense 27B vision-language model with a 262K context and notes it runs "at about 17 GB when quantized." That is a managed endpoint of the real model, billed by the instance hour, with no GGUF involved.
- Self-hosting on EC2 is where Saluki or the lab's files make sense: a small GPU instance running llama.cpp exactly as on your laptop, switched off when you are done.
These are AWS's on-demand Linux prices in US East (N. Virginia), from the official EC2 price list published October 1, 2026, for the single-GPU instances that fit this model:
| Instance | GPU | Per hour | Fits |
|---|---|---|---|
| g6.xlarge | 1x NVIDIA L4, 24 GB | $0.8048 | Saluki or any GSQ-RCO file with a long context; the 18 GB 4-bit too |
| g6.2xlarge | 1x NVIDIA L4, 24 GB (more CPU and RAM) | $0.9776 | Same GPU; only if you also run other things |
| g5.xlarge | 1x NVIDIA A10G, 24 GB | $1.006 | The older equivalent; pick g6 unless g5 is what your account has quota for |
| g6e.xlarge | 1x NVIDIA L40S, 48 GB | $1.861 | The 8-bit model, or several models at once |
The arithmetic that matters: at $0.8048 an hour, an evening of testing Saluki on a g6.xlarge costs about the price of a coffee, and a month of 24-hour uptime costs about $590 plus storage, which is more than a used 12 GB graphics card. Rent to test and for bursts; buy for a shop assistant that runs all day. Whatever you rent, keep the llama.cpp port closed to the internet (it has no authentication) and reach it through SSH or inside your VPC. Our guide to Hugging Face models on AWS covers the instance setup, storage and the patterns for keeping an endpoint private.
Saluki vs Bonsai 2 vs full Qwen3.8-27B vs a 9B model: which to run
Four reasonable choices for someone with modest hardware, side by side:
| Option | File | Memory | Runs in | Choose it when |
|---|---|---|---|---|
| Underdog Saluki 27B | 7.89 GB | 9 to 11 GB | Stock llama.cpp, Ollama, LM Studio | Tool-calling agents on the smallest machine; you accept weaker math |
| GSQ-RCO IQ3_S | 11.8 GB | 12 to 14 GB | Stock llama.cpp, Ollama, LM Studio | 16 GB card or 24 GB Mac; you want the original's scores |
| Ternary Bonsai 2 27B | 5.95 GB | about 6 GB plus context | PrismML's llama.cpp fork only | 6 GB cards and you are happy to build the fork |
| Full Qwen3.8-27B, 4-bit | 18 GB | 16 to 19 GB | Everything, incl. Ollama's library | 24 GB card or 32 GB Mac; the reference experience |
| A good 9B model (MiMo-V2.6 9B, Qwen3.5 9B) | about 6 GB at 4-bit | 8 GB | Everything | 8 GB card with no RAM to spare; fastest responses |
The honest question for an 8 GB card is the last row. A 9B model at 4-bit fits entirely on the card and answers quickly; Saluki is a much larger brain running partly from system RAM, more capable per answer and slower per token. For an assistant that answers short questions from tools, the 9B is often the better daily driver; for one that has to reason through a messy ticket, Saluki earns its extra seconds. Our MiMo-V2.6 9B guide is the 9B we would try first.
Saluki not working? Troubleshooting by symptom
"Out of memory" or a crash while loading
The whole model did not fit on the graphics card. Lower -ngl (try 40, then 32), reduce -c to 8192, and do not load the vision projector until the text model runs. In LM Studio, drag the GPU offload slider down. If system RAM is also short, close the browser.
The model answers instead of calling the tool
Check that the server was started with --jinja; without it the chat template is not applied and tool calls never happen. Set temperature 0 and thinking off for the call. Make the function description specific: "look up how many of a repair part the shop has" works better than "stock function."
Tool arguments come back slightly malformed
Expected now and then: the maker reports formatting slips in about a fifth of parallel-call replies. Parse leniently, validate the arguments, and retry the request once if parsing fails. Avoid asking for several tool calls in one turn when one at a time will do.
It thinks for a very long time
Reasoning effort defaults to "xhigh." Pass "reasoning_effort": "medium" or "low" in chat_template_kwargs, or switch thinking off for tasks that do not need it.
Ollama says it cannot find the :F16 tag
The tag in the Hugging Face snippet does not exist in this repository. Download the GGUF and import it with a Modelfile as shown above.
Wrong answers on arithmetic or letter puzzles
Known weak spots of this 2-bit cut. For anything numeric, give the model a calculator tool and let it call that, which is exactly the pattern it is good at; or use the lab's IQ3_S if you have the memory.
Images are ignored
The main file is text-only. Add --mmproj with one of the two projector files, and make sure your client actually sends the image in the message.
Underdog Saluki 27B and Qwen3.8-27B: frequently asked questions
What is Underdog Saluki 27B?
A 7.89 GB GGUF version of Alibaba's Qwen3.8-27B, published October 8, 2026 by ConwayResearch. It is a 2-bit "mix" built on ISTA-DASLab's GSQ-RCO quantization and tuned, by an undisclosed method, to keep tool calling intact. It runs in stock llama.cpp, Ollama and LM Studio under Apache 2.0.
Can I run Qwen3.8-27B on an 8 GB graphics card?
With Saluki, yes, if you have system RAM behind the card: the 2-bit cuts need 9 to 11 GB of VRAM plus RAM in total, so an 8 GB card holds most layers and the rest sit in RAM, at reduced speed. A 12 GB card runs it entirely on the GPU.
How much RAM does Underdog Saluki 27B need?
About 9 to 11 GB of VRAM and RAM combined for the model and a working context, by Unsloth's figures for 2-bit Qwen3.8-27B, plus 0.63 to 0.93 GB if you load the vision projector and 1 to 2 GB more if you use multi-token prediction.
Is Underdog Saluki 27B as good as Qwen3.8-27B?
On its maker's tests it is level with the full model on tool calling and instruction following, a few points behind on coding, and keeps about 82 to 85 percent of the full model's competition-math scores. Those are the maker's numbers on a 120-task test; verify on your own tasks.
What is GSQ-RCO quantization?
A method from ISTA-DASLab that stores each tensor of a model at a different precision, chosen to fit a size budget while losing as little quality as possible. Its 11.8 GB IQ3_S file of Qwen3.8-27B matches the original on AIME25 and LiveCodeBench, and its 8.4 GB IQ2_XS beats a same-size standard quantization by ten points on AIME25.
Should I download Saluki or the ISTA-DASLab GSQ-RCO files?
Saluki when you need the smallest file for a tool-calling agent on a tight machine. The lab's IQ3_XXS (10.1 GB) or IQ3_S (11.8 GB) when you have a 16 GB card or a 24 GB Mac and want scores that match the full model, especially on math.
How do I run Underdog Saluki 27B on Windows 11?
Download the llama.cpp release for your GPU, download the GGUF with hf download, and run llama-server with --jinja -ngl 99 -fa on -c 32768. On an 8 GB card lower -ngl to about 40 and the context to 8192. Open http://127.0.0.1:8080 to chat.
Does Underdog Saluki 27B work with Ollama?
Yes, by importing the downloaded GGUF with a Modelfile (FROM ./Underdog-Saluki-27B-1.0-IQ2-mix.gguf) and ollama create. The :F16 tag shown in the Hugging Face snippet does not exist in the repository.
Does Underdog Saluki 27B support vision?
Yes, through a separate projector file, mmproj-Underdog-Saluki-27B-1.0-F16.gguf (0.93 GB) or the Q8_0 version (0.63 GB), passed to llama-server with --mmproj. The main file alone is text-only.
How do I turn off thinking in Qwen3.8 and Saluki?
Per request, add "chat_template_kwargs": {"enable_thinking": false} to the body; for the whole server, start llama-server with --chat-template-kwargs '{"enable_thinking": false}'. The maker recommends thinking off and temperature 0 for tool calls.
Is Underdog Saluki 27B free for commercial use?
Yes. It is Apache 2.0, as are Qwen3.8-27B and the ISTA-DASLab files it is built on. The Underdog Bench tasks come from the Berkeley Function Calling Leaderboard.
How fast is Underdog Saluki 27B?
The maker publishes no speed figures. On a 12 GB or larger card with all layers on the GPU, expect fluent reading speed; on an 8 GB card with layers in RAM, noticeably slower; on a processor alone, a few tokens per second. Thinking off is the biggest speed lever.
Is Qwen3.8-27B on Amazon Bedrock?
No. Bedrock's Qwen catalog is the earlier Qwen3 generation. Qwen 3.8-27B has been in Amazon SageMaker JumpStart since August 27, 2026, as a managed endpoint, and you can self-host Saluki or the GSQ-RCO files on an EC2 GPU instance.
What does it cost to run Qwen3.8-27B on AWS?
A g6.xlarge (one 24 GB NVIDIA L4) is $0.8048 an hour on demand in US East, by AWS's October 2026 price list, and runs Saluki, the GSQ-RCO files or the 18 GB 4-bit model. A g6e.xlarge with a 48 GB L40S is $1.861 an hour.
Saluki vs Ternary Bonsai 2 27B: which is better?
Saluki runs in stock llama.cpp, Ollama and LM Studio and scored 88 of 120 on the maker's tool-calling test; Bonsai 2 needs PrismML's patched llama.cpp and scored 70 on the same test. Bonsai 2 is smaller, 5.95 GB against 7.89 GB.
Is Qwen3.8-27B dense or a mixture of experts?
Dense: all 27 billion parameters work on every token, in a 64-layer hybrid of Gated DeltaNet and Gated Attention blocks. That is why there is no "active parameters" shortcut and why the file has to be quantized this hard to get small.
How do I install Underdog Saluki 27B on Kali Linux?
Install git, cmake and build-essential, clone llama.cpp, build with -DGGML_CUDA=ON (or Vulkan, or CPU only), download the GGUF with hf download inside a Python venv, and run build/bin/llama-server with --jinja -ngl 99 -fa on -c 32768.
By the end of the evening, Jake's laptop was answering "do we have a Pixel 9 battery" by reading the stock list instead of guessing, in a few seconds with thinking off, and Ethan had stopped apologizing for the speed. The 54 GB model had become a 7.89 GB file, the 8 GB card had become enough with some help from RAM, and the only thing Jake lost was a model that could do competition math, which has never once come up at the counter. The spreadsheet stayed on the laptop. So did the questions.
If you keep one line from this page
A 7.89 GB file needs 9 to 11 GB to run, and it keeps the tool calling while giving up the math.
Start llama-server with --jinja, test ten tool calls with thinking off at temperature 0, and reach for the lab's IQ3_S the day you have 16 GB.
Revision note. Written October 10, 2026, two days after Underdog Saluki 27B appeared. "Cut your coat according to your cloth," says the old proverb; this page is the measuring tape.
