Kavli Affiliate: Daeyeol Lee
| Authors: Joanna Aloor, Timothy PH Sit, Oliver M Gauld, Joseph Warren, Matthew Mower, Daeyeol Lee and Chunyu A. Duan
| Summary:
Adaptive behaviour usually requires exploiting regularities in the environment, but in competitive settings the opposite can be true: predictable choice patterns can be exploited by others, making unpredictability itself advantageous. How neural circuits generate such strategic variability remains poorly understood. Here, we trained mice to play a zero-sum game against an opponent that exploited statistical regularities in their choices and rewards, and tracked their behaviour and dorsal cortical dynamics across learning. Using a hidden Markov model, we found that mice transitioned from structured, predictable strategies towards a near-optimal stochastic strategy as they learned. Applying the same framework to monkeys playing the same game identified a shared stochastic strategy across species, despite differences in how it was deployed. Cortex-wide imaging revealed that stochastic choices were associated with reduced representation of reward history, while immediate reward signals remained robust. Critically, while the strength of cortical reward signals predicted subsequent choice during reward-guided behaviour, this relationship was abolished during stochastic behaviour. Thus, adaptive stochasticity does not simply arise from a loss of reward information, but from selectively decoupling reward from future choice. These results reveal a neural mechanism through which animals suppress otherwise useful reward-guided structure to generate adaptive unpredictability in competitive environments.