Centaur: The Model That Mimics Human Thought on a Large Scale | Hands-on large language models pdf | Large language models introduction pdf | Llm model | Turtles AI
Centaur is an advanced computational model, based on LLaMA 3.1 refined with QLoRA on over 10 million human choices in 160 Psych‑101 experiments. It generalizes to novel tasks and reflects human neural structures with astonishing accuracy.
Key points:
- Psych‑101 dataset: 60,000 participants, 160 experiments, 10M+ choices.
- Efficient fine‑tuning on LLaMA 70B with QLoRA.
- Ability to predict human choices even in new or modified contexts.
- Alignment of internal representations with human fMRI data.
In the cognitive AI landscape, Centaur represents a remarkable leap towards an integrated view of the mind. Developed by a team at the Institute for Human-Centered AI at the Helmholtz Center in Munich led by Marcel Binz and Eric Schulz, the model was born by refining LLaMA3.1 (70 billion parameters) with a single training epoch on Psych-101, an extraordinarily large corpus: over 10 million decisions collected from more than 60,000 people across 160 trial-by-trial paradigms.
This dataset, uniformly transcribed in natural language, includes decision-making, supervised learning, multi-armed bandits, logic puzzles, memory, and Markov decision processes. The QLoRA tuning strategy allowed the model to be adapted parsimoniously, changing only 0.15% of the parameters, thus ensuring the preservation of the original linguistic knowledge.
The results are surprising: Centaur outperforms not only the basic LLaMA, but also specialized cognitive models on unseen participants, tasks with modified narrative plots, and even entire new domains, such as LSAT-like logic tests. It not only predicts choices, but also reaction times, demonstrating a deeper understanding of human behavior.
One aspect of particular interest comes from the analysis of internal representations: without having been trained on neural data, Centaur shows a surprising convergence with human fMRI activities, suggesting that it is capturing real cognitive structures.
In theoretical terms, Centaur meets many of Newell’s criteria for a unified theory of cognition: responsiveness to the environment, real-time operation, use of extensive symbolic knowledge, and robustness to the unknown. The potential applications are broad: in-silico prototyping of psychological studies, simulation in clinical contexts (anxiety, depression), optimization of experimental design, and even uses in automated computational psychology.
The project is also open-source: the Psych-101 dataset and the code for Centaur (repository “Llama-3.1-Centaur-70B”) are publicly available and already accessible to the scientific community. The team plans to expand the corpus to include additional domains – language psychology, social psychology, economics – and individual variables such as age, personality and socioeconomic status, to make the model even more representative and powerful.
Rewriting cognitive theories in a computational key becomes concrete today, with Centaur inaugurating a new era of shared understanding between mind and machine.


