Pruna AI makes genAI models smaller and faster. And it’s open source | Large language models course | Llm meaning software | Hacker news chatgpt voice github | Turtles AI

Pruna AI makes genAI models smaller and faster. And it’s open source
Pruna AI opens its optimization framework to the community: a new standard for AI model compression
Editorial Team20 March 2025

 

Pruna AI, a European startup specializing in AI compression algorithms, has recently open-sourced its optimization framework, providing the community with advanced tools to improve the efficiency of AI models.

Key points:

  • Open-source release of the Pruna AI optimization framework.
  • Application of techniques such as caching, pruning, quantization and distillation.
  • Automated post-compression performance evaluation.
  • Availability of a compression agent for custom optimizations.


The Pruna AI framework integrates methodologies such as caching, pruning, quantization and distillation to reduce the size and increase the speed of AI models. This platform standardizes the saving and loading processes of compressed models, allowing flexible combinations of compression techniques and accurate performance evaluation after optimization.

A distinctive aspect of the framework is the ability to analyze the impact of compression on model quality, determining whether accuracy losses are significant and what performance improvements have been achieved. This approach is reminiscent of the standardization introduced by platforms like Hugging Face for transformers and diffusers, but applied to efficiency methodologies.

While large AI labs have already implemented compression techniques, such as the distillation used by OpenAI to create faster versions of its flagship models, Pruna AI offers an integrated solution that combines several methodologies in a single tool. This aggregated approach makes it easy to use and combine compression techniques, offering significant added value compared to tools focused on single methods.

Pruna AI’s framework is designed to support a wide range of models, including natural language processing, diffusion, speech-to-text, and computer vision. The company is currently focusing its efforts on image and video generation models, collaborating with companies such as Scenario and PhotoRoom.

In addition to the open source version, Pruna AI offers an enterprise solution with advanced features, including an optimization agent. This agent allows developers to specify performance and accuracy requirements, automating the optimization process to meet those criteria. The pricing model for the pro version is based on hourly rates, similar to renting GPUs on cloud services, and aims to ensure a return on investment through savings on inference costs.

Pruna AI recently raised $6.5 million in seed funding, with investors including EQT Ventures, Daphni, Motier Ventures, and Kima Ventures. This capital injection underscores the company’s confidence in the potential to revolutionize the efficiency and accessibility of AI models through its compression framework.

Pruna AI’s initiative represents a significant step towards democratizing AI model optimization techniques, offering advanced tools to both the open source community and enterprises to develop more efficient and sustainable solutions.