THE COMPACTIFAI PIPELINE

How we build our AI models

Four connected stages, one shared methodology, and the path that takes a base model to a finished release.

The AI model landscape is changing quickly. New models arrive every few weeks, capabilities that once required massive infrastructure are becoming available at a fraction of the size, and organizations are asking different questions than they did a year ago: not only how capable a model is, but how much it costs to run, where it can be deployed and who stays in control of it.

Multiverse Computing sits at the intersection of those questions. As an AI solutions and models provider, we build models designed for efficiency and security โ€” models that organizations can deploy and control in the environments where they need them.

Behind our compressed models there is a consistent process. We take a strong open model as a starting point and shape it, stage by stage, into a Multiverse Computing model with its own size, behavior and benchmark profile. The technology that drives that process is CompactifAI, and the sequence we follow from base model to release is what we call the CompactifAI pipeline.

We publish compressed models regularly, both open-source and closed-source. HyperNova-60B, Pulsar-16B, and LittleLamb are all on Hugging Face under Apache 2.0, each built from a different base model and each accompanied by its own model card with benchmark tables. What they share is the methodology behind them.

  • HyperNova-60B

    60B total / 4.8B active MoE, from gpt-oss-120B

  • Pulsar-16B

    16.15B total / 3.1B active MoE, BF16 and NVFP4

  • LittleLamb

    290M parameters, tool-calling and mobile variants

  • Quasar v2

    Flagship model, from GLM-5.3, available through our API

This page describes that methodology: how each stage of the pipeline works and how the stages fit together. CompactifAI is our proprietary framework, and while we share part of how it works here, there is more work that cannot be published. Think of this page as a standing reference, not a paper. We update it when the pipeline changes, not when a new model ships. Per-release numbers live in the model cards; the story of how those models are built lives here.

Overview

The pipeline at a glance

CompactifAI runs in four stages: profiling, structured pruning, healing, and post-training. A compressed model passes through all four, but the stages are not independent steps applied in isolation.

At the profiling step, the pipeline generates several pruning candidate configurations and evaluates them in parallel. Each candidate gets a short healing run, and the one with the best result in that run, its predicted recoverability, proceeds to full healing. Accuracy immediately after pruning is not the selection criterion.

  1. Stage 01

    Profiling

    Finds where parameters carry redundant information

    • Tensor-network analysis
    • MPO ยท MPS ยท TT
    • MERA ยท PEPS
    Where to prune
  2. Stage 02

    Structured pruning

    Removes whole structures, never individual weights

    • Channel pruning
    • Layer & block trimming
    • Expert trimming & merging

    Candidate selection

    Several pruning configurations are evaluated at once, not one after another.

    Candidates

    • Config A
    • Config B (selected)
    • Config C

    Short healing

    One short run per candidate

    Recovery

    • Lower
    • Best
    • Lower
    Smaller Dense Model
  3. Stage 03

    Healing

    Distils the original model into the pruned one

    • Offline KL distillation
    • Cached top-100 logits
    • Fused chunked KL loss
    Recovered capabilities
  4. Stage 04

    Post-training

    Refines behaviour and lowers precision

    • Supervised fine-tuning
    • Reinforcement learning
    • Quantization
    Further size & latency gains
  5. Deployed Model

    Standard dense transformer. Runs on stock inference stacks, no custom kernels.

We select by recoverability because the two measures often disagree. Aggressive candidates can look strong right after pruning and still finish behind more moderate ones once healing completes. A toolchain that optimises each stage against its own metric cannot see this, because no stage knows what the next one will be able to repair.

This coupling is also why we report end-to-end results rather than per-technique attribution. Profiling determines which layers, channels or experts are candidates for pruning. Those choices determine the structure that healing has to repair, and healing is calibrated to that specific configuration. Because each stage's effect depends on the others, a breakdown by technique would not describe how the system behaves. We optimise for the model the user deploys, not for an intermediate score that would not survive healing.

Stage 01

Profiling with tensor networks.

Decompositions used in physics for three decades, applied at the scale of frontier language models โ€” and kept out of the model we ship.

The first stage analyses the model to understand where its parameters carry redundant or low-impact information. This profile defines the candidate pruning configurations that the pipeline generates and evaluates. The profiler is based on a broad family of tensor-network methods โ€” Matrix Product Operator (MPO), Matrix Product State (MPS) and Multi-scale Entanglement Renormalization Ansatz (MERA), among others, used in algorithms such as the Density Matrix Renormalization Group (DMRG). The mathematics is established; our contribution is applying this analysis to models in the 0.6B to 800B parameter range and integrating its signal into the compression decisions that follow.

Our tensor-network profiler is an analysis instrument that informs which parameters are safe to remove; it is not part of the architecture of the model we deliver. The shipped model keeps a standard transformer structure, loadable on any inference framework without custom layers or patched libraries. We describe this method as tensor-network-aware pruning.

A related technique in our portfolio is layer tensorization, which replaces a layer with a tensor-network factorisation, keeping the tensor-network structure inside the model rather than only at profiling time. This is a genuine technique that yields real gains in convolutional architectures, where the factorised structure maps well to CPU computation. On modern LLMs running on NVIDIA GPUs, it does not beat a dense layer of equivalent size, because GPUs are optimised for dense vector-matrix operations. For that reason, layer tensorization still stays a research line for LLMs and is kept out of the shipped LLM artifact. It is part of the portfolio, applied where it pays off and held back where it does not.

Stage 02

Structured pruning

The compression stage removes parameters from the model, working from the candidate configurations that profiling defines.

We use structured removal at channel, layer, and expert granularity, never unstructured weight sparsity. Our speed-ups are algorithm-level. We do not write custom CUDA kernels, build custom compilers, or operate at the hardware-firmware layer. The speed-up the user sees on a standard GPU is the direct effect of a smaller model running on the same vector-matrix kernels: fewer parameters, fewer activations, less wall-clock work per token.

Four structured-compression techniques are in the production pipeline:

  • Channel pruning

    Removes channels from the model's weight matrices, selected using the tensor-network-aware profiling signal described above. This is the stage where the profiling analysis has its most direct effect on the final model. The output is a narrower model with the same depth but fewer parameters per layer.

  • Layer and block trimming

    Removes entire transformer blocks identified as marginal contributors during profiling. We formulate the selection of which blocks to remove as a constrained binary optimization problem, equivalent to finding low-energy states of an Ising glass. Read the paper | Try the code

  • Expert trimming and merging.

    Profiling mixture-of-experts architectures finds rarely used or overlapping experts. These are removed or merged to combine their functions. Few experts activate per token, so removing experts cuts total parameters more than active ones. Parameter removal impacts MoE models differently than dense models, affecting quality variably. Thus, profiling and multi-candidate evaluation are crucial for MoE.

  • Layer tensorization

    As described above, this technique is in the portfolio and applied where it is effective, which today means convolutional models, not production LLMs on GPU.

Stage 03

Healing: distillation at production scale

Pruning removes parameters. Healing recovers the capabilities that were lost. This is the stage where the bulk of our engineering investment has gone over the past year, and it is where the pipeline's quality is ultimately decided.

Frozen

Teacher

Original uncompressed model

output distribution

KL divergenceย ยท gradient

Updated

Student

The pruned model

Unlabelled text ยท teacher format

We implement healing as knowledge distillation. The loss function is the Kullback-Leibler divergence between the teacher's and the student's output distributions, rather than cross-entropy on hard labels. We use a proprietary variant of this objective that focuses the gradient on the most informative regions of the distribution.

The dataset supplies the text on which the two distributions are compared. It provides no labels. The text must match the teacher's format โ€” its tokenizer and prompt format โ€” and public corpora do not always do so. To close that gap we built an in-house synthetic data pipeline that generates teacher-formatted text at scale.

Healing runs in two phases. The first is a pretrain-style stage, using pretraining-quality data to recover knowledge and patterns lost to pruning, using the original model's pretraining data where available. The second is an SFT-style stage, where instruct-format data is used and domain-specific data can optionally be mixed in to steer the model toward a customer's area.

A full healing run typically uses on the order of 100 billion tokens. The figure grows with model size, the gap between the base and target domains, and the gap to the teacher the customer is willing to accept. Smaller models, or customers who accept a wider gap, need materially less

  • ~100B

    tokens in a full healing run, growing with model size, with the gap between base and target domains, and with the gap to the teacher the customer accepts

  • 2 phases

    pretrain-style recovery, then SFT-style steering

  • 0 labels

    the corpus supplies text, the teacher supplies the target

Engineering the distillation stack

The distillation step is what decides most of the final model quality, and it is also the most expensive part of the pipeline. The standard setup, online distillation, keeps both teacher and student loaded at the same time. At each training step, the teacher runs a full forward pass to produce its output distribution, and the student is trained to match it. This is the most expressive setup, since the full teacher distribution is available, but it is also the most memory-intensive: two full-vocabulary tensors must be held per token position, and the teacher must be recomputed on every step even though its behaviour does not change across a training run.

We built our healing stack on Megatron Bridge and NVIDIA NeMo, with substantial in-house extensions that re-engineer the distillation workflow for the compression-recovery use case. The stack is not a fork of NVIDIA's tooling; it is a re-engineered version targeted at our specific scenario. The key systems changes are:

  • Offline, top-K logits

    We run the teacher once over the healing corpus, store its 100 most likely tokens per position, and train the student against that cache. The teacher is never in memory during training, and the cache is reused across ablations and across pruning candidates, which share a teacher.

    ~5ร— faster in that phase
  • Fused, chunked KL loss

    Computed naively, the loss builds a full vocabulary-by-sequence-length grid before it can produce a single number: for gpt-oss-120b โ€” a vocabulary of 201,088 tokens at a sequence length of 32K and batch size 4 โ€” the teacher-probability tensor alone is about 50 GB in bfloat16, and a single iteration peaks at roughly 250 GB once gradients, activations, weights and optimizer states are added, more than an H200 or B200 can provide. Ours processes the sequence one slice at a time, fusing the output projection directly into the loss so the full grid is never materialised. Peak memory grows linearly with sequence length.

    Linear in sequence length
  • Proprietary KL variant

    Tuned for the healing case, with faster convergence focused on the informative regions of the distribution. We state its existence; implementation details are not disclosed.

    Not disclosed

Offline distillation with a fused chunked KL loss

PEAK MEMORY (GB) AT 8K CONTEXT ยท SINGLE H200

H200 limit ยท 141 GB

Online distillation

102.8 GB

Offline dense KL

78.3 GB

Offline forward-chunked

61.8 GB

Offline fused chunked KL

58.3 GB

AT 32K CONTEXT ยท GPT-OSS 20B

  • Step time (s)
  • Throughput (TFLOP/s)

4 GPU nodes standard setup

57.0 s
74.2

1 GPU fused chunked KL

12.2 s
345.7
Offline distillation with the fused chunked KL loss. Left: peak VRAM across four distillation methods at 8K context on a single H200. The fused chunked variant uses 58.3 GB against 102.8 GB for online distillation. Right: at 32K context on GPT-OSS 20B, the same loss shrinks the setup from four GPU nodes to one, cutting step time from 57.0 to 12.23 seconds (about 5x faster). Source: paper Figures 1 and 2.

The result is documented in Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss and its companion blog post, Making Knowledge Distillation Cheap Enough to Run at Scale. The chunked-loss implementation is open-sourced at CompactifAI/Full-Chunked-KL-Loss, and we have contributed to logits-precomputation handling in SGLang and vLLM.

That affordability is what makes the multi-candidate pipeline practical: we can evaluate several compression configurations, heal each one and compare the final results, because healing each candidate is no longer prohibitively expensive.

  • 4 โ†’ 1

    GPU nodes to a single GPU, at 32K context

  • 57.0 โ†’ 12.2 s

    step time, about 5ร— faster

  • 102.8 โ†’ 58.3 GB

    peak VRAM at 8K context on one H200

Quantization-aware healing

Quantization adds a second source of damage on top of pruning, and the standard recovery recipes were designed for models that were only quantized. Quantization-aware training (QAT) inserts fake-quantization operators into the forward pass and keeps fine-tuning on a task loss. A variant, quantization-aware distillation, replaces the task loss with a KL divergence to a full-precision teacher. Both rely on a full-precision version of the same model, which exists when quantization is the only change. This works when the only change is quantization, because a genuine full-precision version of the same model exists to serve as teacher. After a model has gone through structural pruning, not just fewer bits but fewer layers, heads, and neurons, there is no independently trained full-precision version of the smaller architecture. The only candidate teacher is the recovered bfloat16 checkpoint, which is itself a distilled approximation of the original. Distilling from it anchors the quantized student to a degraded target.

Our approach, which we call quantization-aware healing, removes that ceiling with one change: it distills directly from the original, pre-compression model rather than from the recovered checkpoint. Teacher and student do not share an architecture. The teacher is full-size and full-precision; the student is at least half the size in parameters and running in 4-bit precision. Because a teacher's output distribution is architecture-agnostic, nothing about the size or shape mismatch prevents the transfer.

  • Quantization-aware training

    Objective

    Task loss, with fake-quantization operators inserted into the forward pass

    Teacher

    None โ€” it fine-tunes against hard labels

  • Quantization-aware distillation

    Objective

    KL divergence to a full-precision teacher

    Teacher

    The same model at full precision โ€” which only exists if quantization was the only change

  • Quantization-aware healing

    Objective

    KL divergence, with teacher and student not sharing an architecture

    Teacher

    The original, pre-compression model โ€” full size, full precision

Applied to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4, this recipe produces a 4-bit model that beats its own bfloat16 source on 7 of 9 benchmarks. The largest gains land on exactly the capabilities compression usually damages most: long-context reasoning and math. The method also reaches its peak in roughly 100 steps, about 7 times faster than quantization-aware training, and then stays stable, because KL distillation against a frozen teacher gives the student no incentive to drift once it has matched the teacher. A cross-entropy task loss, by contrast, keeps pushing toward hard labels and eventually erodes inherited capabilities. Our blog post on quantization-aware healing describes the full results, including the head-to-head comparison against QAT.

  • 7 of 9

    benchmarks where the 4-bit model beats its own bfloat16 source

  • ~100 steps

    to reach its peak, about 7ร— faster than QAT

  • Stable

    no drift once it matches the frozen teacher

Stage 04

Post-training and stacking quantization

Healing recovers most of a compressed model's knowledge, and it is where we have concentrated most of our effort. Post-training targets what healing leaves behind. This area has been especially strengthened in the past months from different research groups and it is where the weight is now.

Quantization-aware healing

Reinforcement learning has driven much of the recent progress in LLM reasoning. Just to name a few, Group Relative Policy Optimization (GRPO), on-policy distillation (OPD) and multi-teacher OPD (MOPD). All these methods complement healing.

  • GRPO

    Group Relative Policy Optimization trains a model on its own sampled responses, scoring each against the others sampled for the same prompt.

    Reward from the group
  • On-policy distillation

    Keeps that loop but replaces the reward with a teacher signal: the student samples responses and the teacher scores every sampled token, pulling the student toward the teacher's distribution on its own outputs.

    Reward from the teacher
  • Multi-teacher OPD

    Routes each prompt to a domain-specialised teacher and distills them into a single student.

    One student, many teachers

We are focused now in this area, driving new experiments to improve our pipeline, and soon we will publish a few articles and blog posts about this stage.

Quantization compounds with compression.

Quantization is complementary to our compression, not a substitute for it. Quantization reduces bit-width without changing the architecture; structural pruning removes parameters entirely. Both reduce memory and compute, but they operate on different axes and their gains compound when stacked.

Pulsar-16B: compression and quantization stacked

WEIGHT MEMORY (GB) ยท EACH LAYER COMPOUNDS

Nemotron 3 Nano 30B

59 GB

Pulsar-16B BF16โˆ’50%

30 GB

Pulsar-16B NVFP4โˆ’67%

10 GB

INFERENCE PERFORMANCE ยท FASTER AT EVERY LAYER

  • Throughput (tok/s)
  • TTFT (s)

Nemotron 3 Nano 30B

3,363
2.18 s

Pulsar-16B BF16

3,760
1.80 s

Pulsar-16B NVFP4

4,735
1.25 s
Figure 2. Pulsar-16B: compression and quantization stacked. Compressed from NVIDIA's Nemotron 3 Nano 30B-A3B at roughly 50% parameter reduction, shipped in BF16 and NVFP4. NVIDIA independently reproduced the full evaluation suite on its own hardware.

Quantization-aware healing, described in the previous section, is the technique that bridges compression and quantization. Rather than treating quantization as a lossy postprocessing step applied after healing is finished, it uses the quantization stage as a second pass of distillation against the original teacher. The 4-bit student is not compensating for information lost to quantization; it is picking up information the earlier recovery stage did not transfer.

custom compression

Three ways a compression project is scoped.

Engagements typically follow one of three scenarios, which differ in what the compression targets.

  1. For a use case or domain

    Without a defined purpose, the pipeline preserves the model's broad knowledge and skills, typically at a compression ratio of 50% or more. LLMs hold knowledge across many topics and are trained for tasks a given customer may never need.

    When the customer defines what the model must keep โ€” a topic, an area, a use case or something broader โ€” the pipeline prunes and heals toward that goal.

    Because the capabilities to preserve are narrower, more can be removed, up to 95% of the original size. CompactifAI works for general-purpose compression, but its advantage is greatest when the target is specific.

    Up to 95% of the original size
  2. For a memory footprint

    The customer specifies the hardware the model must fit on, and we find the most accurate compressed model that fits within it.

    Fits the target accelerator
  3. Of a customer's fine-tuned model

    The fine-tuned model itself becomes the teacher, and the student is derived from it. We need the model weights, not the training data: the student learns to reproduce the fine-tuned model's output distribution, and the upstream labels that produced the fine-tune are irrelevant for healing.

    No proprietary data leaves the perimeter

DEPLOYMENT

What the delivered model looks like

The model that comes out of the pipeline is a standard transformer. No custom layers, no tensor-network structures in the weights, no patched libraries. Swapping it for the original model on an existing inference stack is a drop-in operation.

  • vLLM
  • TensorRT
  • plain CUDA
  • llama.cpp
  • HF transformers

This was a deliberate design choice. An earlier iteration of our work, from the Llama 2 era, kept the MPO decomposition inside the layer at deployment time. That produced a model with non-standard layer structure, which required patching every third-party library to accept the new layer type. For production-scale deployment we switched to using tensor-network methods only at profiling, so the deployed artifact is a stock transformer that loads on any framework without modification. The tensor-network-in-model line remains active in our research team and would re-enter production if hardware optimised for tensor operations becomes available.

The speed-ups we deliver come from the model being physically smaller: fewer parameters, fewer activations, less wall-clock work per token, running on the same vector-matrix kernels as the original. We can also adapt model dimensions to hardware constraints โ€” if a target accelerator requires a specific block size or weight format, the compressed model can be configured to meet it.

RESULTS

Evidence from public releases.

Across more than 25 model families compressed end-to-end, these four show the pipeline at different scales. Per-release benchmark numbers, inference configurations and hardware-specific results live in each model's card.

Quasar v2

Flagship ยท from GLM-5.3 ยท Available through the Multiverse Computing API

Our flagship closed-source model, released together with this report. It applies the full CompactifAI pipeline to GLM-5.3: structural compression, knowledge-distillation healing, and quantization-aware healing stacked on top. The model is available through our API and is not distributed as open weights.

The most common description of a model like this in the field is a pruned model, but that does not capture what Quasar is or what producing it involved. Compression removes parameters, but the model the user gets is the one that emerged from healing, quantization-aware healing, and the synthetic data pipeline that feeds both. Those stages are where the engineering investment landed, and they are what shape the model's behaviour. The base checkpoint is the starting point. The profiling analysis, compression configuration, distillation loss, data tools, quantization bridge, and the systems work to operate them at this scale are ours. Producing Quasar was not a free win; it was a directed engineering investment.

Beyond the standard pipeline, Quasar v2 received targeted work specific to this release. The GLM base model carries content restrictions on Chinese political topics that reduce its usefulness for general-purpose deployments. We removed those restrictions. We also applied security improvements, whose details will be documented as they are finalised.

Artificial Analysis independently benchmarked Quasar v2. The full report and charts will appear in this section once published.

  • GLM-5.3

    base model

  • 3 stages

    compression, healing and quantization-aware healing

  • API only

    not distributed as open weights

  • Pending

    Artificial Analysis Report

HyperNova-60B 2605

60B total / 4.8B active MoE ยท from OpenAI's gpt-oss-120B

A 60B total / 4.8B active MoE model compressed from OpenAI's gpt-oss-120B at roughly 50% parameter reduction. It was independently evaluated by Artificial Analysis and ranked the most efficient model in the 40B to 150B category, with an Intelligence Index of 29.3, and the only European-origin model in the ideal quadrant of intelligence versus parameter count.

It beats its uncompressed 120B teacher on IFBench (instruction following) and LiveCodeBench (coding), and lands within 2 to 4 points of the teacher on most other benchmarks. At inference, it delivers 36% higher throughput and 31% lower time-to-first-token than the teacher on an H200, at half the weight memory.

Model card: MultiverseComputingCAI/Hypernova-60B-2605.

  • +36%

    throughput against the teacher on an H200

  • โˆ’31%

    time to first token

  • ยฝ

    of the teacher's weight memory

  • 29.3

    Artificial Analysis Intelligence Index

HyperNova-60B 2605 vs its 120B teacher

  • gpt-oss-120B (teacher)
  • HyperNova-60B 2605 (Compressed)

IFBench

67.0
68.0

LiveCodeBench

62.8
68.7

AIME25

93.7
90.0

GPQA-D

74.6
71.9

MMLU-Pro

79.6
76.8

ฯ„ยฒ-Bench

63.7
61.7

HLE

18.5
15.0

SciCode

41.5
36.0

T-Bench Hard

24.2
15.9
020406080100

Score

Pulsar-16B

16.15B total / 3.1B active MoE ยท from NVIDIA's Nemotron 3 Nano 30B-A3B

A 16.15B total / 3.1B active MoE model compressed from NVIDIA's Nemotron 3 Nano 30B-A3B at roughly 50% parameter reduction, shipped in BF16 and NVFP4 variants. NVIDIA independently reproduced the full evaluation suite on B200 and L40S hardware and confirmed the results.

The NVFP4 variant is one-sixth the weight memory of the uncompressed teacher while staying within 3 to 5 accuracy points on hard benchmarks, and long-context behaviour is preserved: needle-in-a-haystack at 100K tokens is at parity with the original.

Model cards: Pulsar-16B-BF16, -NVFP4.

  • 59 โ†’ 10 GB

    weight memory, teacher to NVFP4

  • โ…™

    of the teacher's weight memory

  • 3โ€“5 pts

    of the teacher on hard benchmarks

  • 100K

    tokens, needle-in-a-haystack at parity

LittleLamb

290M parameters ยท from Qwen3-0.6B ยท two variants

Demonstrates the pipeline at edge scale. Two variants, both compressed from Qwen3-0.6B at roughly 50% reduction to 290M parameters, are published openly.

The tool-calling variant scores 51.55 on BFCL v4 (the de facto industry standard for function calling), roughly double Google's functiongemma-270m-it at 27.03. The mobile variant achieves 86.7% accuracy on Mobile Actions, comparable to or better than functiongemma at 85.0%, while delivering 28% lower time-to-first-token and half the on-disk footprint of its Qwen3-0.6B base.

Model cards: LittleLamb-ToolCalling, LittleLamb-Mobile.

  • 51.55

    BFCL v4, against 27.03 for functiongemma-270m-it

  • 86.7%

    Mobile Actions accuracy, against 85.0%

  • โˆ’28%

    time to first token

  • ยฝ

    the on-disk footprint of its base

REFERENCES

Where to go deeper.