I’m just here for the moral superiority.🌱
Mainly interested in FOSS
Currently in uni and working part-time as a developer and system administrator.

PC Specs
CPU: 7800X3D
GPU: 7900XTX
Memory: 64GB
System: Arch

  • 6 Posts
  • 16 Comments
Joined 9 months ago
cake
Cake day: January 15th, 2026

help-circle












  • I’ve been using it for the past few days and the output quality seems to be on par or slightly better than 3.5 27b. The biggest issue is the token usage that has exploded with this revision. It can easily reason for 20k-25k tokens on a question where the qwen3.5 models used 10k. Since it runs more than 3 times faster, it still finished earlier than the 27b, but I won’t have any context/vram left to ask multiple questions.

    Artificial Analysis has similar findings. image







  • Unfortunately, the AI community prefers rushed buggy development over proper, tested releases, so the quants and maybe the PR weren’t fully working.

    As of 3 hours ago, unsloth was still updating their quants and guide. I don’t have time to test now but I wouldn’t judge the base model performance in the first few days when the bugs are still being worked out.

    They also recommend some unconventional parameters in the Unsloth guide.

    It could also be that the model is truly shit of course.

    Edit I just took a look at the llama.cpp repo and there are still issues with the implementation as well.