TL;DR: Indie developers report that Gemma 4 26B (QAT) feels more intelligent and coherent than Qwen 3.6 35B, despite Qwen's superior benchmark scores.
Summary: A developer noted that Qwen 3.6 35B, despite higher benchmarks, performs less intelligently than Gemma 4 26B (QAT) in terms of prompt adherence and output coherence. This observation suggests a potential discrepancy between benchmark scores and real-world perceived intelligence, possibly due to quantization-aware training (QAT) in Gemma.
Why it matters: AI builders should be aware that benchmark scores don't always reflect practical model performance, especially with quantized models. Experiment with both models to validate real-world utility for specific applications, and consider the impact of quantization methods.
Source: reddit