xista science ventures hat dies direkt geteilt
Qwen3.5-9B, 82% smaller and still capable. The core question in LLM compression is simple: how far can you compress it before the model becomes useless? Our answer is called OraQuant. The idea is intuitive: not every part of a model deserves the same resources. OraQuant finds where the "intelligence" actually lives and spends its budget there: generous where it counts, frugal where it doesn't. We shrank Qwen3.5-9B from ~18GB to 3.2GB and compared it to the best quantized open-source models. The smaller we went, the further ahead our smart bit allocation stack pulled: at the tiniest size, ours scored 40 on AIME-25 while the competition scored a flat 0. Compression isn't about size. It's about knowing what to keep and what to throw away. At Ora, we're pushing that boundary: making models smaller without sacrificing capabilities. Links to the blog post and models are in the comments! #AI #MachineLearning #OpenSource #LLM