Allocation transfer
versus the field.
Two Pi-derived Qwen3.8-27B files were run through the same broad-use suite as the previous eleven quants: 24 applied Hermes tasks, 164 HumanEval+ problems per model, long-context probes, speed, and VRAM measurements on NVIDIA RTX 5060 Ti 16GB GPUs.
Did allocation transfer help?
The transfer model outscored its Pi control by 6.06 balanced points.
It changed HumanEval+ by +0.61 points, ran 10.73 seconds slower, and left 198 MiB less free VRAM in the representative configuration.
This is the cleanest comparison inside the new repository, but it is not a pure quantizer-only comparison against the prior eleven: these files use Pi-derived weights, while the earlier providers use their own source/quantization pipelines.
Head-to-head details
| Model | Balanced | HumanEval+ | Adaptive time | Recorded VRAM use | 64K target |
|---|---|---|---|---|---|
| Qwen3.8-27B Pi RCO allocation-transfer IQ3_XXS | 86.12 | 86.59% | 83.04s | 11.61 GiB | Passed |
| Qwen3.8-27B Pi IQ2_M control | 80.06 | 85.98% | 72.31s | 11.42 GiB | Passed |
All 13 tested models
Sorted by the personalized balanced index. The two new Pi rows are highlighted. *Maximum context is an allocation probe, not proof of recall across the whole window.
| Rank / model | Balanced | HumanEval+ | Adaptive workload | VRAM used | Max allocation |
|---|---|---|---|---|---|
| #1 Qwen3.8-27B Pi RCO allocation-transfer IQ3_XXSPi RCO transfer · IQ3_XXS | 86.12 | 86.59% | 83.04s | 11.61 GiB | 131K* |
| #2 Swift-Qwen3.8-27B IQ2_SSwift · IQ2_S | 85.05 | 77.44% | 67.66s | 10.71 GiB | 131K* |
| #3 Qwen3.8-27B GSQ-RCO IQ3_S MTPGSQ-RCO · IQ3_S · MTP | 83.19 | 89.63% | 53.49s | 14.29 GiB | 131K* |
| #4 Qwen3.8-27B Unsloth UD-Q2_K_XLUnsloth · UD-Q2_K_XL | 82.45 | 86.59% | 82.79s | 11.20 GiB | 131K* |
| #5 Qwen3.8-27B GSQ-RCO IQ3_SGSQ-RCO · IQ3_S | 81.34 | 87.80% | 105.79s | 13.43 GiB | 131K* |
| #6 Qwen3.8-27B GSQ-RCO IQ3_XXSGSQ-RCO · IQ3_XXS | 81.25 | 90.24% | 98.51s | 11.87 GiB | 131K* |
| #7 Qwen3.8-27B GSQ-RCO IQ2_XS MTPGSQ-RCO · IQ2_XS · MTP | 80.80 | 90.24% | 48.29s | 11.46 GiB | 131K* |
| #8 Qwen3.8-27B GSQ-RCO IQ2_S MTPGSQ-RCO · IQ2_S · MTP | 80.66 | 85.98% | 49.62s | 12.29 GiB | 131K* |
| #9 Qwen3.8-27B Pi IQ2_M controlPi control · IQ2_M | 80.06 | 85.98% | 72.31s | 11.42 GiB | 131K* |
| #10 Qwen3.8-27B GSQ-RCO IQ2_SGSQ-RCO · IQ2_S | 80.04 | 84.76% | 65.45s | 11.09 GiB | 131K* |
| #11 Qwen3.8-27B GSQ-RCO IQ2_XSGSQ-RCO · IQ2_XS | 79.12 | 85.98% | 71.73s | 10.43 GiB | 131K* |
| #12 Qwen3.8-27B GSQ-RCO IQ3_XXS MTPGSQ-RCO · IQ3_XXS · MTP | 78.68 | 87.80% | 49.76s | 13.21 GiB | 131K* |
| #13 Qwen3.8-27B Unsloth UD-IQ2_SUnsloth · UD-IQ2_S | 76.11 | 72.56% | 84.44s | 10.25 GiB | 131K* |
Recommendations by Hermes role
HumanEval+ remains 10% of the balanced score. The other 90% reflects business/C-Suite analysis, research, knowledge work, agent tools, practical sysadmin, and executive content—the broader operating layer this benchmark is meant to select for.
64K context on 16GB
The deployment target remains 65,536 tokens with Q4_0 K/V cache; the representative benchmark allocates 73,728 tokens and exercises a roughly 57K-token prompt. Both new files passed the 64K allocation target. Recorded representative VRAM use was 11.61 GiB for the transfer model and 11.42 GiB for the control.
131K remains the next real context test.
Allocation alone does not establish retrieval quality, prompt-processing practicality, or long-session stability. A full 131K workload should test all three.
Speed, MTP, and quality
These two new files contain no MTP head, so their speed should be compared with the non-MTP entries. The prior matched GSQ pairs remain the evidence for MTP: they delivered substantial speed gains with roughly 1.0–1.4 GiB of extra VRAM, while quality changes varied by quant. Across all thirteen models, the speed leader is Qwen3.8-27B GSQ-RCO IQ2_XS MTP at 48.29s, and the HumanEval+ leader is Qwen3.8-27B GSQ-RCO IQ2_XS MTP at 90.24%.
Methodology and cost
One GPU per shard.
Pinned commit c82967099.
24 applied prompts and 164 generated programs per new model.
Previous eleven-model testing cost $5.31; cumulative benchmark rental cost is $5.696. The current number includes failed/retried rentals associated with this experiment, while internet-transfer charges may settle separately in Vast billing.
Reproducibility and limitations
- Applied tasks use deterministic scoring rubrics, but small differences should be checked against the source outputs.
- HumanEval+ measures Python correctness; it does not directly measure research synthesis, financial analysis, tool reliability, or document quality.
- GPU class and runtime are controlled, but separate rental hosts can introduce CPU and storage variance.
- The Pi transfer and control quant types differ (IQ3_XXS versus IQ2_M), so this measures each published artifact as a deployable choice—not the allocation-transfer technique in isolation.
- Generated code was executed only in the local network-disabled, read-only, capability-dropped Docker scorer.
New source repository: https://huggingface.co/firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF