On August 5, 2026, Multiverse Computing and Qualcomm Technologies announced a collaboration to optimize AI models for Qualcomm's Dragonfly AI200 and AI250 data center accelerators. The partnership combines Multiverse Computing's model compression and optimization technology with Qualcomm's AI acceleration hardware to reduce compute and memory footprint, enabling higher throughput and lower power consumption without requiring additional hardware
.
Demonstrated Performance Improvements
The efficiency gains were first demonstrated live at Mobile World Congress in March 2026 on a Qualcomm Cloud AI100 Ultra accelerator (the same optimization techniques apply to the AI200/AI250 family). Two use cases were showcased
:
Real-time emergency medical reporting (compressed LLM):
- Up to 93% faster response times
- Up to 44% higher throughput
- Up to 45% reduction in memory usage
- Up to 21% reduction in power consumption
- No loss in accuracy
Confidential financial document RAG chatbot:
- Up to 35% faster response times
- Up to 54% higher throughput
- Up to 45% reduction in memory usage
- Up to 14% reduction in power consumption
- No loss in accuracy
For the medical use case, the results were independently corroborated by multiple outlets, confirming the 93% faster response time, 45% memory savings, and 21% power reduction
. Multiverse Computing noted that running on the newer Dragonfly AI200/AI250 accelerators directly is expected to yield even better results .