Reka Flash 3: a 21B reasoning model with great performance | Build a large language model from scratch pdf | What is generative ai | Hackers guide to machine learning | Turtles AI
Reka recently open-sourced a preliminary version of Reka Flash 3, a 21 billion parameter multimodal language model designed to excel at tasks like general chat, encoding, instruction, and function calls. This compact model offers competitive performance compared to proprietary solutions like OpenAI o1-mini, making it ideal for applications that require low latency or local device deployments. Its ability to handle contexts of up to 32,000 tokens makes it particularly versatile.
Key Points:
- 21 billion parameter open source model
- Competitive performance with proprietary solutions
- Support for extended contexts of up to 32,000 tokens
- Suitable for low latency and local device applications
Reka Flash 3’s training process was meticulous, starting with pre-training on a heterogeneous set of synthetic and publicly accessible data. Next, the model was tuned with instructions based on high-quality, curated data to improve its performance. In the final step, reinforcement learning was applied using REINFORCE Leave One-Out (RLOO), leveraging both model-based and rule-based rewards to further enhance its capabilities. Unlike models that specialize in mathematics or coding, Reka Flash 3 aims for general improvements through reinforcement learning.
Reka Flash 3’s performance is remarkable: on AIME-2024, with a budget of 16 reasoning steps, the model demonstrated significant efficiency. Additionally, on WMT’23, it achieved a COMET score of 83.2, showing improvements over previous versions in multilingual understanding. However, as a more compact model, it may not be the ideal option for tasks that require extensive knowledge, as indicated by its MMLU-Pro score of 65.0. Therefore, it is recommended to integrate it with web search for tasks that require in-depth knowledge.
A distinctive aspect of Reka Flash 3 is its ability to "think" before generating an answer, using tags to delimit the reasoning process. This mechanism allows the model to stop reasoning after a certain number of steps, while still ensuring reasonable outputs. This feature allows processing time to be managed according to predefined budgets, providing flexibility to developers.
In terms of deployment, Reka Flash 3 is optimized for low-cost applications that require low latency or execution on local devices. At full precision, the model occupies 39 GB (fp16), but can be compressed up to 11 GB while maintaining high performance thanks to 4-bit quantization. This makes it more efficient than larger models such as the QwQ-32B, which requires 64 GB at bf16 and 18 GB with 4-bit quantization.
It is important to note that while Reka Flash 3 was designed primarily for English, it has demonstrated some understanding of other languages. However, in some cases, the model may process English reasoning even when questions are asked in other languages, impacting the quality of the output. Additionally, the model has not undergone extensive alignment or training, suggesting room for improvement in the future.
For those who want to test Reka Flash 3, a trial version is available on Reka Space. Additionally, the model weights are downloadable and modifiable under the Apache 2.0 license, providing developers and researchers with a powerful yet lightweight foundation on which to build custom applications.
Reka Flash 3 represents a significant step forward in the field of multimodal language models, combining efficiency, versatility, and accessibility, opening up new opportunities for innovative applications in the AI field.
