7.81 tokens per second on an Intel Xeon 6737P marked a 94.1% throughput gain for Multiverse Computing’s CompactifAI-compressed Llama 3.3 70B versus 4.02 tokens per second uncompressed.
2,598.22 seconds to handle a 1,024-input, 1,024-output request cut latency 48.6% from 5,056.34 seconds, with inter-token latency and time per output token both improving by about 48%.
256 concurrent users produced the biggest gains, with throughput up 107.0% and latency down 51.7%, while the compressed model’s size fell roughly 50% to 65 GiB from 130 GiB.
97% of baseline accuracy was retained across standard benchmarks, the company said, despite declines including 2.48% on MMLU; WinoGrande improved 6.86% after targeted retraining.
PyTorch and Hugging Face compatibility positions CompactifAI as a way to run large models on standard server hardware with lower storage, energy and deployment demands.
With CPUs now running large AI models, is the era of absolute GPU dominance in artificial intelligence nearing an end?
Does making AI cheaper to run risk a massive surge in usage, negating the promised energy efficiency gains?
Llama 3.3 Runs Twice as Fast on Intel Xeon 6 Thanks to Multiverse Computing’s CompactifAI
Overview
Multiverse Computing has announced a major breakthrough with its CompactifAI technology, nearly doubling the performance of Llama 3.3 on Intel Xeon 6 processors. This advancement makes powerful AI models more accessible and cost-effective for enterprises, allowing sophisticated AI to run efficiently on standard hardware. As a result, businesses can integrate advanced AI functionalities like Llama 3.3 into their existing infrastructure without the need for expensive hardware upgrades. This marks a pivotal moment for AI deployment, addressing key challenges in the industry and enabling faster, more efficient AI adoption across a wider range of applications.