SAN SEBASTIÁN, Spain - September 21st 2026. Multiverse Computing today launched Luminary, a platform that tells enterprises which large language model performs best for their specific use case, including use cases built on multi-turn conversation. Luminary is available through early access now.
Companies choosing an LLM today face roughly 2.9 million models on public repositories and almost no reliable way to compare them for their own work. Public leaderboards measure academic tasks, not real workflows. Cost-per-token says nothing about cost-per-completed-task. The result is predictable: decision-makers default to the most hyped frontier model and pay a premium for performance their use case never needed.
Single prompts do not predict conversations
The gap is sharpest exactly where enterprise AI spend is now concentrated: chatbots, virtual agents and agentic systems. Model evaluation today is overwhelmingly single-turn. A prompt goes in, a response comes out, the response gets a score. That tells a business nothing about the interactions it actually sells and supports, where a customer changes their mind at turn four, contradicts themselves at turn seven, asks something off-policy at turn nine, and the model has to hold context, call the right tools and still close the task.
Luminary evaluates the whole conversation. Each candidate model runs inside a digital twin that plays the customer across a full multi-turn journey and simulates the backend systems, tool calls and function executions it has to work with. Failure modes that only appear over several turns, such as lost context, tool misuse, policy drift and conversations that never reach resolution, surface before deployment rather than after it.
Because Luminary scores complete journeys, it can compare models for agentic and conversational systems, a class of use case that single-turn evaluation cannot assess at all.
Luminary is vendor-neutral, allowing enterprises to evaluate Multiverse models alongside third-party models from any OpenAI-compliant inference provider.
How Luminary works
The workflow runs in five automated steps: describe the use case, generate journeys, simulate, evaluate, and select the champion model. Four capabilities sit underneath it.
Intelligent test generation. Luminary identifies missing information in the brief and auto-generates multi-turn user journeys covering the happy path, edge cases and adversarial scenarios, all compliant with the customer’s own policies.
Digital twin simulation. Each test runs inside a digital twin that simulates the full interaction turn by turn, including the customer, backend systems, tool calls and function executions, rather than a single isolated prompt.
Business-relevant metrics. Luminary measures cost per completed task, success rate, latency per job and policy compliance across the entire conversation, in place of cost per million tokens.
Full auditability. Every test case is reproducible end to end, turn by turn, and compliance teams get an exportable audit trail. This gives technical, business and compliance stakeholders a shared body of evidence behind the model-selection process.
Luminary is designed to help enterprises reduce model spend by identifying when smaller, less expensive models can meet their requirements, cut deployment risk by catching failure modes before production, and move from brief to champion model selection in days.
Model selection also isn’t a one-time decision. Enterprises can rerun evaluations as new models and provider versions become available, using the same framework for regression testing, model upgrade evaluation, cost optimization and vendor comparison.
“Every enterprise AI programme starts with the same question, and almost nobody answers it with evidence: which model should we actually run? Teams pick the model with the loudest launch and discover the cost six months later,” said Enrique Lizaso, co-founder & CEO of Multiverse Computing. “And if you are building a chatbot or an agent, scoring one prompt at a time tells you nothing, because your product is the conversation, not the reply. Luminary treats model selection the way aerospace treats airframe design. You put the thing in a wind tunnel before you fly it, across the whole flight envelope. We simulate the customer over every turn, the systems and the tool calls, then hand the business a ranked answer on cost per completed task, not a benchmark score.”
Availability
Luminary is available through early access now. Check out Luminary’s page.
About Multiverse Computing
Multiverse Computing is a leader in sovereign and efficient AI. The company develops fast, efficient, and highly specialised AI models that enable organizations to deploy advanced artificial intelligence securely within their own infrastructure, ensuring full control over data, governance, and compliance. Headquartered in Donostia-San Sebastián, Spain, with offices in the United States, Canada, and across Europe, Multiverse serves more than 100 global customers, including Iberdrola, Bosch, and the Bank of Canada.
Contact
business@multiversecomputing.com