Sonny's Labs · personalized Hermes benchmark

Allocation transfer
versus the field.

Two Pi-derived Qwen3.8-27B files were run through the same broad-use suite as the previous eleven quants: 24 applied Hermes tasks, 164 HumanEval+ problems per model, long-context probes, speed, and VRAM measurements on NVIDIA RTX 5060 Ti 16GB GPUs.

Best of all 13Qwen3.8-27B Pi RCO allocation-transfer IQ3_XXS86.12
Pi transfer overall rankQwen3.8-27B Pi RCO allocation-transfer IQ3_XXS#1 of 13
Pi control overall rankQwen3.8-27B Pi IQ2_M control#9 of 13

Did allocation transfer help?

The transfer model outscored its Pi control by 6.06 balanced points.

It changed HumanEval+ by +0.61 points, ran 10.73 seconds slower, and left 198 MiB less free VRAM in the representative configuration.

Qwen3.8-27B Pi IQ2_M control80.06 · #9
Qwen3.8-27B Pi RCO allocation-transfer IQ3_XXS86.12 · #1

This is the cleanest comparison inside the new repository, but it is not a pure quantizer-only comparison against the prior eleven: these files use Pi-derived weights, while the earlier providers use their own source/quantization pipelines.

Head-to-head details

ModelBalancedHumanEval+Adaptive timeRecorded VRAM use64K target
Qwen3.8-27B Pi RCO allocation-transfer IQ3_XXS86.1286.59%83.04s11.61 GiBPassed
Qwen3.8-27B Pi IQ2_M control80.0685.98%72.31s11.42 GiBPassed

All 13 tested models

Sorted by the personalized balanced index. The two new Pi rows are highlighted. *Maximum context is an allocation probe, not proof of recall across the whole window.

Rank / modelBalancedHumanEval+Adaptive workloadVRAM usedMax allocation
#1 Qwen3.8-27B Pi RCO allocation-transfer IQ3_XXSPi RCO transfer · IQ3_XXS86.1286.59%83.04s11.61 GiB131K*
#2 Swift-Qwen3.8-27B IQ2_SSwift · IQ2_S85.0577.44%67.66s10.71 GiB131K*
#3 Qwen3.8-27B GSQ-RCO IQ3_S MTPGSQ-RCO · IQ3_S · MTP83.1989.63%53.49s14.29 GiB131K*
#4 Qwen3.8-27B Unsloth UD-Q2_K_XLUnsloth · UD-Q2_K_XL82.4586.59%82.79s11.20 GiB131K*
#5 Qwen3.8-27B GSQ-RCO IQ3_SGSQ-RCO · IQ3_S81.3487.80%105.79s13.43 GiB131K*
#6 Qwen3.8-27B GSQ-RCO IQ3_XXSGSQ-RCO · IQ3_XXS81.2590.24%98.51s11.87 GiB131K*
#7 Qwen3.8-27B GSQ-RCO IQ2_XS MTPGSQ-RCO · IQ2_XS · MTP80.8090.24%48.29s11.46 GiB131K*
#8 Qwen3.8-27B GSQ-RCO IQ2_S MTPGSQ-RCO · IQ2_S · MTP80.6685.98%49.62s12.29 GiB131K*
#9 Qwen3.8-27B Pi IQ2_M controlPi control · IQ2_M80.0685.98%72.31s11.42 GiB131K*
#10 Qwen3.8-27B GSQ-RCO IQ2_SGSQ-RCO · IQ2_S80.0484.76%65.45s11.09 GiB131K*
#11 Qwen3.8-27B GSQ-RCO IQ2_XSGSQ-RCO · IQ2_XS79.1285.98%71.73s10.43 GiB131K*
#12 Qwen3.8-27B GSQ-RCO IQ3_XXS MTPGSQ-RCO · IQ3_XXS · MTP78.6887.80%49.76s13.21 GiB131K*
#13 Qwen3.8-27B Unsloth UD-IQ2_SUnsloth · UD-IQ2_S76.1172.56%84.44s10.25 GiB131K*

Recommendations by Hermes role

C-Suite AnalysisSwift-Qwen3.8-27B IQ2_S78.80
Researcher / ScoutQwen3.8-27B Pi RCO allocation-transfer IQ3_XXS91.93
Hermes CoordinatorQwen3.8-27B Pi RCO allocation-transfer IQ3_XXS93.61
Librarian / KnowledgeQwen3.8-27B Pi RCO allocation-transfer IQ3_XXS92.45
Forge / CoderQwen3.8-27B Pi RCO allocation-transfer IQ3_XXS91.36

HumanEval+ remains 10% of the balanced score. The other 90% reflects business/C-Suite analysis, research, knowledge work, agent tools, practical sysadmin, and executive content—the broader operating layer this benchmark is meant to select for.

64K context on 16GB

The deployment target remains 65,536 tokens with Q4_0 K/V cache; the representative benchmark allocates 73,728 tokens and exercises a roughly 57K-token prompt. Both new files passed the 64K allocation target. Recorded representative VRAM use was 11.61 GiB for the transfer model and 11.42 GiB for the control.

131K remains the next real context test.

Allocation alone does not establish retrieval quality, prompt-processing practicality, or long-session stability. A full 131K workload should test all three.

Speed, MTP, and quality

These two new files contain no MTP head, so their speed should be compared with the non-MTP entries. The prior matched GSQ pairs remain the evidence for MTP: they delivered substantial speed gains with roughly 1.0–1.4 GiB of extra VRAM, while quality changes varied by quant. Across all thirteen models, the speed leader is Qwen3.8-27B GSQ-RCO IQ2_XS MTP at 48.29s, and the HumanEval+ leader is Qwen3.8-27B GSQ-RCO IQ2_XS MTP at 90.24%.

Methodology and cost

HardwareNVIDIA RTX 5060 Ti 16GB

One GPU per shard.

Runtimellama.cpp 0.5.0-dev

Pinned commit c82967099.

New quality work48 applied + 328 HumanEval+

24 applied prompts and 164 generated programs per new model.

This experimentActual Vast rental$0.386

Previous eleven-model testing cost $5.31; cumulative benchmark rental cost is $5.696. The current number includes failed/retried rentals associated with this experiment, while internet-transfer charges may settle separately in Vast billing.

Reproducibility and limitations

New source repository: https://huggingface.co/firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF