A 27-billion-parameter model squeezed down to 2-bit precision can still do real development work on a single 16GB graphics card. Qwen3.8-27B-Escha-W2, a 2-bit build of Qwen3.8 27B, needs only about 10GB for its weights. In a nine-part test it found a hidden passkey in all 15 trials across a 256k-token context, passed 75 of 100 coding problems, and built working web apps and games with little or no help. The same tests also showed where 2-bit quantization costs you: polish on 3D assets and accuracy on framework-specific game code. Here is what the results mean if you want to run a large model locally on modest hardware.
What a 2-bit 27B model looks like
Quantization lowers the numeric precision of a model’s weights so the model takes less memory. Qwen3.8-27B-Escha-W2 is a 2-bit quantized fine-tune of Qwen3.8 27B, and its original weights take 10.15GB. That leaves plenty of room on a 16GB or 24GB consumer GPU for the context, the text the model keeps in view while it works.
The model is available on Hugging Face in GGUF, the file format used by the llama.cpp inference engine, in two builds:
| Build | File size | Notes |
|---|---|---|
| Q8_0 embed and head | 10.31GB | 2.38GB smaller and 3.4% faster at generation than F16 |
| F16 embed and head | 12.69GB | Weight rounding differs from Q8_0 by only about 0.02% |
Because the precision difference is negligible, the tests used the Q8_0 build. An optional 2.93GB MTP (multi-token prediction) draft model can speed up generation through speculative decoding, where a helper model guesses several tokens ahead. It was left out to keep as much of the 16GB as possible free for context.
The test machine and the nine tasks
The model ran under llama.cpp on Ubuntu Server 24.04.4 with an AMD Ryzen 5700X CPU and 32GB of DDR4 memory. The target GPU was an NVIDIA RTX 2000 Ada with 16GB of VRAM. A second card, an RTX 5090, was installed only to speed up the test automation.
The nine tasks fall into four groups:
- Performance: raw prompt processing and generation speed
- Memory: retrieval from a very long context
- Problem solving: a tiered reasoning suite and 100 coding problems
- Development: a Kanban board web app, a falling-sand physics sandbox, a dungeon crawler game, and two agent tasks that drive Blender and Godot through MCP (Model Context Protocol, a standard way for AI models to use outside tools)

▲ Nine tests for a local model
Speed depends mostly on the card
Prefill is how fast the model reads the prompt. Decode is how fast it writes the answer.
| Measurement | Result |
|---|---|
| Uncached prefill | 235.1 tokens per second at 512 tokens, 295.1 at 8k, 271.2 at 32k |
| Cached prefill | 1,904.7 tokens per second at 512 tokens, 115,252 at 32k |
| Decode | 11.6 tokens per second on short outputs, 11.4 on average up to 16k |
A decode speed of about 11 tokens per second is usable but not fast, and the GPU appears to be the main reason. The RTX 2000 Ada is a 70-watt workstation card with memory bandwidth of 225 to 250GB per second. Standard 16GB desktop cards from the RTX 30, 40 or 50 series, such as the RTX 4080 or RTX 4070 Ti Super, have higher power limits and faster memory, so they should generate tokens much faster with the same model.
Long-context memory, reasoning and coding scores
The memory test used a Needle-in-a-Haystack setup. The context was expanded to 256,608 tokens, and a secret passkey was hidden at 0%, 25%, 50%, 75% and 100% of the way through the text. Each depth got three trials, 15 in total, over 56 minutes and 36 seconds. The model found the passkey every time, a 100% pass rate. Its answers did come wrapped in extra thinking and chatter, and they grew a little longer and noisier at deeper positions, but the extracted answer was always correct.
| Test | Result |
|---|---|
| Reasoning, 48 questions | 38 passed (79%), 21,047 ms average per question |
| By difficulty | Easy 12 of 12, Medium 12 of 12, Hard 10 of 12, Expert 4 of 12 |
| HumanEval Remix, 100 problems | 75 passed (75%), 5 left unanswered |
| Answered problems only | 75 of 95 passed (79%) |
The reasoning suite covers logic, planning and probability and took 16 minutes and 50 seconds. HumanEval Remix is a set of 100 Python challenges adapted from OpenAI’s HumanEval benchmark; it ran for 1 hour and 8 minutes with an average response time of 41,067 ms. For a model of this size compressed to 2 bits, 79% on reasoning and 75% on coding look like solid results. The sharp drop at Expert difficulty is worth keeping in mind before you hand it hard logic problems.
Development tasks: strong on apps and games, weaker on 3D
| Task | Extra user prompts | Context used | Outcome |
|---|---|---|---|
| Kanban board web app | 1 | 57.3% | One missing button handler fixed |
| Falling-sand sandbox | 0 | 30.6% | Worked on the first try |
| Dungeon crawler | 0 | 34.2% | Worked on the first try |
| Blender lantern | 1 | 89.7% | Overexposed glow, misaligned hook |
| Godot 3D platformer | 3 | 64.0% | Inverted controls and collision errors fixed |
Web app and simulation
The Kanban board supported drag-and-drop between To Do, In Progress and Done, column reordering, card editing in a dialog, priority labels, and dark and light themes. The only flaw in the first build was that the Add Column button had no click listener attached. After one message saying the button did nothing, the model spotted the omission and patched its JavaScript file. The default look was decent rather than striking, but the functional code was high quality.
The falling-sand sandbox is a cellular automaton, a grid where each cell updates by simple rules. Sand piled up, water ran down slopes, and acid dissolved both walls and sand. Brush sizing, an eraser and a clear button all worked, with no corrections needed.
Procedural game
The dungeon crawler tested algorithmic skill. The model chose binary space partitioning, which splits the map into rectangles again and again, to lay out connected rooms and hallways. It used Bresenham’s line algorithm to cast sight lines up to 9 tiles from the player, revealing the map as the player moves and shading unexplored, remembered and visible tiles differently. It even caught and fixed its own bug where the spawn point was picked twice. No extra prompting was needed.

▲ Game logic versus 3D asset quality
Agent tasks through MCP
The MCP tasks showed the limits of 2-bit precision more clearly. In Blender, the model built a lantern with a hexagonal frame, a pointed cap, a hanging ring and an internal light, then rendered four views. The interior glowed too brightly, and the top hook clipped into the mounting ball instead of looping around it. The model also got stuck repeating attempts to save the scene and run Blender headlessly, and moved on only after it was asked directly for the screenshots.
The Godot platformer, the hardest task, asked for collectible keys, moving hazards and a goal. The first version had inverted controls on both axes and a pit cover with no collision. The model also used a collision property that does not exist in Godot 4, which broke the crumbling platform script, then fixed it by switching the collider’s disabled flag. After three correction prompts and two context compactions, the game was fully playable.
What to take away before running a big model on 16GB
These results suggest that a 2-bit build is a realistic way to run a 27B model on a single 16GB GPU, with no spill into system memory. For long-document retrieval and everyday web development, the quality loss is hard to notice. For visually polished 3D work or code that must follow a specific framework’s syntax exactly, Q4 or Q5 quantization is noticeably better if your hardware allows it.
If you plan to run a model like this locally:
- Pick the Q8_0 embed and head build over F16 on a 16GB card to save more than 2GB of VRAM.
- Skip the MTP draft model when memory is tight and give that space to context instead.
- Check your card’s memory bandwidth first, because it largely sets generation speed.
- When a generated app has a small bug, describe the exact symptom, such as which button does not respond.
- If the model loops on the same command during an MCP task, interrupt it and ask for the output files directly.
- Review code for version-specific tools like Godot 4 for properties that do not exist.