Today Multiverse Computing is releasing Quasar 1.1 438B, a rebuild of our flagship coding model. We took our flagship model back through the full CompactifAI pipeline: re-healed on broader data, tuned to answer in fewer tokens, and stripped of the political refusal behavior it inherited from its base model. Part of that healing set was generated by a hybrid quantum large language model, with circuits run on IBM Quantum System Two in Donostia-San Sebastian, a 156-qubit IBM Heron processor. It is the first time quantum-generated data has entered the CompactifAI pipeline. Quantum computing is not a label on this release, it is part of how the model was built.
More than compression
Quasar 1.1 438B is built from GLM-5.2, the open-weights model from Z.ai, transformed with CompactifAI. Quantum-inspired expert selection prunes each layer from 256 experts down to 148, which is where the reduction in size and serving cost comes from.
Compression on its own produces a smaller model that is worse at its job. Every capability the pruning disturbs has to be put back deliberately, and the traits that were never suited to production (such as verbosity, inherited refusals, precision the workload does not need) are worth fixing while the model is open on the table. That is what CompactifAI is: a pipeline that takes an open-weights checkpoint and returns something improved and deployable, not merely something smaller.
What changed in 1.1
Healing on more data.
After pruning, we retrain against a broader dataset: reasoning traces, tool-call sequences and general knowledge. The new model improved in general knowledge and reasoning (HLE, +6.2%; and GPQA, +4.3%) and on instruction-following (IFBench, +4.6%). After healing, average output also dropped from 3,322.9 to 2,074.3 tokens (37.6% fewer) across SciCode, HumanEval, GSM8K, TriviaQA and BBH.
Quantum synthetic data.
A part of the healing set was generated by a hybrid quantum large language model (1/6 of Qwen3-30B-A3B’s layers replaced with a QNN) with circuits run on IBM Quantum System Two in Donostia-San Sebastian, a 156-qubit IBM Heron processor, and on a noise model calibrated from that device. This is the first time quantum-generated data has entered the CompactifAI pipeline, and one more step in a long-running line of work at the intersection of language models and quantum computing.
This is the first time quantum-generated data has entered the CompactifAI pipeline, and it sits in a line of work we have been running for years at the intersection of language models and quantum computing. The quantum-inspired tensor network methods behind CompactifAI’s expert selection came out of the same line. What is new in 1.1 is that the quantum side is no longer only an inspiration for classical algorithms: part of the data that shaped this model came off a quantum processor in the city where we build these models.
Removing censorship.
The topic-level restrictions inherited from the base model are gone. The standard safety filters are untouched. Refusal on politically sensitive prompts drops from 71.18% (GLM-5.2) and 63.75% (Quasar 1.0) to 41.00%. Refusal on harmful prompts holds at 93.00% on JailbreakBench, above Quasar 1.0’s 92.00%.
GLM-5.2 carries restrictions such as hard-coded refusals on politically sensitive subjects and state-aligned framing on history and politics, none of which compression changes. It rarely surfaces in agentic coding, but for research or policy analysis a model that produces a single approved narrative is not a reliable analytical tool. The new model removes the behavior at inference time using the process described in Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics.
An LLM-as-a-judge scores refusal confidence and a ridge-regularized steering vector isolates the refusal-compliance direction, so the edit moves refusal behavior without dragging capability along. The safety envelope is untouched: Quasar 1.1 438B still refuses weapons instructions, targeted harassment and the rest of what any responsibly deployed model refuses.
Built for agentic work and European deployment
Quasar 1.1 438B is aimed at one thing: running as the engine inside an agent. That means tool calls that parse, long context that stays coherent across a multi-step task, code that compiles, and answers that stop when the work is done. The healing set, the verbosity tuning and the quantization target were all chosen against that profile rather than against a general chat leaderboard. Where a general-purpose model is tuned to be a good conversational partner.
The new model is tuned to be a reliable component and one European organizations can actually deploy: served by a European company incorporated under EU law, developed in line with EU regulatory requirements, including the transparency expectations of the EU AI Act.
More updates are coming for Quasar. Stay tuned.
Get started
CompactifAI API. Sign up here: dashboard.compactif.ai
Benchmarks. Quasar 438B on Artificial Analysis
The research. Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
Contact. business@multiversecomputing.com
