When the chips learn to think: the artificial mind that pulsates like a brain | Is chatgpt a large language model | Llm meaning death | Llm machine learning tutorial for beginners | Turtles AI
Spikingbrain 1.0, a linguistic model developed in China, uses neural networks impulsive impulsively inspired by the human brain: it activates only relevant neurons, requires less data (<2 %), works on Chinese Metax chips and manages megalong sequences with speed up to 100 ×.
Key points:
- Selective activation of neurons only on relevant inputs
- Training with about 2 % of conventional data
- Performance up to 100 times faster on huge sequences
- Independent operation from Nvidia chip, on Domestic Metax GPU
Spikingbrain 1.0 is a new architecture conceived at the Institute of Automation of the Chinese Academy of Sciences in Beijing which re -elaborates the concept of artificial language: instead of activating every node simultaneously as in transformer, it uses neuronal spiking networks that respond only when necessary, reducing consumption and accelerating elaboration times.The result is an almost linear inference process compared to the length of the input, with a speed of Time -to -First -Token (TTFT) greater than 100 times on contexts of 4 million token compared to standard models.
During the pre -training phase, Spikingbrain 7b and the 76b variant with MOE (Mixth of Experts) architecture they reach performance comparable to models such as Llama 2 70b, Gemma2 27B or Mixtral, but using just 2 % of the typical corpus data, around 150 billion token. In addition, the team made the 7 billion parameters model Open Source and published a web demo for version 76b, allowing public tests.
The system was developed and tested entirely on a Made in China GPU infrastructure, in particular Metax C550 clusters, avoiding any dependence on NVIDIA hardware and ensuring technological autonomy in the context of commercial restrictions imposed by the United States. This step reflects a national strategy to create a self -sufficient ecosystem, potentially useful in scenarios where access to advanced GPUs is limited.
The combination of Sparseness at the single neuron level (spiking with dynamic thresholds) and macro modularity (MOE) offers an efficient and biologically plausible multi-scala architecture, with over 69 % of inactive neurons in most of the steps. In addition, in tests on compressed CPU mobile devices (1 billion parameter model), Spikingbrain has reached speed 4-15 times higher than Llama 3.2 when it processes long sequences up to 256k token.
The approach is particularly suitable for managing texts or ultra -moving data: from the verification of legal contracts or medical records to the modeling of scientific data such as genomic sequences, molecular dynamics or complex physical simulations.The researchers underline that the model could inspire the design of future neuromorphic chips with consumption close to those of the human brain, which works with about 20 total watts.
Spikingbrain 1.0 thus proposes an alternative to the dominance of transformer, focusing on the principle of "emerging intelligence" from only selective activation in neurons, rather than on always active dense networks. The model also maintains stability on extended training for weeks on hundreds of Metax GPUs, with good use of hardware and balance control of parallel communication and optimization of operators.
An elegant synergy between neuroscience and computational engineering that lays the foundations for new efficient and sustainable calculation models.


