Preprint
Reinforcement Learning

Projective simulation for artificial intelligence

Hans J. Briegel(Universität Innsbruck), Gemma De las Cuevas(Universität Innsbruck)
May 15, 2012Scientific Reports171 citations

171

Citations

13

Influential Citations

Scientific Reports

Venue

2012

Year

Abstract

We propose a model of a learning agent whose interaction with the environment is governed by a simulation-based projection, which allows the agent to project itself into future situations before it takes real action. Projective simulation is based on a random walk through a network of clips, which are elementary patches of episodic memory. The network of clips changes dynamically, both due to new perceptual input and due to certain compositional principles of the simulation process. During simulation, the clips are screened for specific features which trigger factual action of the agent. The scheme is different from other, computational, notions of simulation, and it provides a new element in an embodied cognitive science approach to intelligent action and learning. Our model provides a natural route for generalization to quantum-mechanical operation and connects the fields of reinforcement learning and quantum computation.

Analysis

Why This Paper Matters

This paper introduces projective simulation (PS), a novel framework for building learning agents that fundamentally differs from traditional reinforcement learning (RL) approaches. Instead of relying on value functions or policy gradients, PS models the agent's decision-making as a random walk through a network of episodic memory clips. This allows the agent to mentally simulate future scenarios before acting, a capability that is central to human-like planning and deliberation.

The significance lies in its potential to unify concepts from cognitive science, RL, and quantum computation. By grounding the simulation process in a physical (quantum) framework, the authors open a path toward quantum-enhanced learning agents that could exploit superposition and interference to explore multiple action sequences simultaneously. This bridges two previously separate fields and offers a fresh perspective on how to design agents that learn efficiently from limited experience.

Technical Contributions

  • Projective simulation framework: The agent maintains a network of clips (elementary memory units) connected by transition probabilities. A random walk through this network simulates possible futures, and clips that match certain features trigger actual actions.
  • Dynamic clip network: The network evolves based on new perceptual inputs and compositional rules, allowing the agent to build structured memories and generalize across similar situations.
  • Quantum generalization: The authors show that the classical random walk can be replaced by a quantum walk, enabling the agent to explore multiple simulation paths in superposition. This provides a natural route to quantum speedups in learning.
  • Embodied cognitive grounding: The model is positioned within embodied cognitive science, emphasizing that simulation is not a detached computational process but is grounded in the agent's sensorimotor experience.

Results

The paper is primarily conceptual and does not present experimental results or quantitative comparisons with existing RL algorithms. No benchmarks, learning curves, or performance metrics are reported. The contributions are theoretical: a formal description of the PS model, its dynamics, and its quantum extension. The impact is therefore measured by the subsequent work it inspired (171 citations) rather than by immediate empirical validation.

Significance

Projective simulation has influenced research at the intersection of RL and quantum computing, inspiring further work on quantum reinforcement learning and quantum agents. Its emphasis on episodic memory and simulation aligns with trends in deep RL that use memory-augmented networks (e.g., differentiable neural computers). The framework also offers a principled way to incorporate planning into RL without requiring a separate world model. While the original paper lacks empirical results, its conceptual clarity and quantum connection have made it a touchstone for researchers exploring alternative foundations for intelligent agents.