Researchers make smaller GPT-OSS model outperform its full-size version

Photo: Logan Wallace / The Ohio State University/Logan

Researchers make smaller GPT-OSS model outperform its full-size version

Multiverse Computing cut the model from 120 billion to 60 billion parameters, and the compressed version beat the larger model on seven of nine tests.

Researchers at Multiverse Computing developed a technique that made a smaller version of OpenAI’s open GPT-OSS model outperform the full-size version on most tests, Decrypt reported.

Advertisement

The team reduced the model from 120 billion parameters to 60 billion and compressed its memory to 4-bit. The smaller model beat the full-quality model on seven of nine benchmarks.

The method, called Quantization-Aware Healing, teaches the compressed model directly from the original uncompressed model instead of from an intermediate version that has already lost accuracy.

The approach could reduce the memory, electricity, and computing resources needed to run AI models while preserving or improving performance, according to the researchers’ paper published on the Hugging Face blog.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
Researchers make smaller GPT-OSS model outperform its full-size version
Researchers make smaller GPT-OSS model outperform its full-size version

Multiverse Computing cut the model from 120 billion to 60 billion parameters, and the compressed version beat the larger model on seven of nine tests.

Share

Add us on Google

Photo: Logan Wallace / The Ohio State University/Logan

Researchers at Multiverse Computing developed a technique that made a smaller version of OpenAI’s open GPT-OSS model outperform the full-size version on most tests, Decrypt reported.

Advertisement

The team reduced the model from 120 billion parameters to 60 billion and compressed its memory to 4-bit. The smaller model beat the full-quality model on seven of nine benchmarks.

The method, called Quantization-Aware Healing, teaches the compressed model directly from the original uncompressed model instead of from an intermediate version that has already lost accuracy.

The approach could reduce the memory, electricity, and computing resources needed to run AI models while preserving or improving performance, according to the researchers’ paper published on the Hugging Face blog.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.