Towards an AI that learns by living: the era of experiential flows | Hacker news chatgpt voice app | A compact guide to large language models pdf github | What is llm | Turtles AI
AI is undergoing a transition from simple text responses to more experiential and continuous learning. DeepMind proposes a new paradigm, based on prolonged interaction with the environment and autonomous formulation of goals through dynamic reward signals.
Key points:
- DeepMind proposes an AI based on prolonged experiences and reinforcement learning.
- Current models lack continuity and adaptation between interactions.
- Reward signals should come directly from the environment, not just from human data.
- Experiential agents could develop superior and autonomous capabilities over time.
A new scenario for AI is emerging in the technological landscape, according to researchers David Silver and Richard Sutton of DeepMind. Their analysis calls into question the very foundations on which modern large-scale language models, such as those used in generative chatbots, are based. Current AIs, they argue, operate in a fragmented episodic system: they answer isolated questions, with no memory between sessions, and are guided by human evaluation criteria that limit their ability to innovate autonomously. Silver and Sutton instead propose a model that breaks with this logic, giving shape to a vision based on “flows of experience”. The AI agent, in this approach, does not limit itself to producing answers based on isolated prompts, but builds a progressive understanding of the world, through continuous interactions, receiving rewards based on the effects of its actions. These rewards are not predetermined by a static set of data, but arise from real and changing environmental signals: biometric parameters, economic metrics, behavioral feedback and other indicators that represent a direct reflection of the agent’s effectiveness in pursuing its goals. The starting point, according to scholars, could be a simulated model of the world, useful for making agents take their first steps, which would then evolve into increasingly complex real environments. This process is based on reinforcement learning, the same principle underlying AlphaZero, the system that has surpassed the best human players in games such as chess and Go. Silver and Sutton point out how generative models have taken over in recent years, but have abandoned the potential of autonomous learning for more general and conversational purposes. This transition, although it has broadened the scope of AI applications, has led to the loss of a fundamental element: the agent’s ability to discover new strategies without direct supervision. The proposal of flows therefore represents a synthesis between the agility of linguistic models and the cognitive depth offered by a dynamic interaction with the context. Experiential agents would be able to pursue complex objectives over time, such as improving an individual’s health or carrying out complex scientific research, adapting to changes and autonomously correcting their own errors. The implications of this evolution go beyond the technical sphere: the debate opens on the role of humans in defining objectives and on the possibility that these systems can operate autonomously for extended periods, raising questions about supervision and safety. However, according to Silver and Sutton, an agent equipped with experiential adaptation could also learn to recognize and mitigate the discomfort or disapproval of humans, avoiding unwanted behaviors. They conclude that the potential of data collected through direct experiences will be enormously superior to the datasets currently used for training, suggesting that it is precisely from this new source of learning that previously unimaginable capabilities will emerge.
A new frontier is therefore opening up for AI, where knowledge is no longer just learned, but also lived.


