Dissociable roles of prefrontal plasticity in decision-making strategy and execution of habitual behavior
Article excerpt
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript.
Nature Communications volume 17 , Article number: 6822 ( 2026 ) Cite this article
Habits are essential for sustaining adaptive behaviors but can also lead to maladaptive behaviors such as compulsive and addictive disorders. They emerge through a shift in decision-making from motivation-driven strategy to an automatic one. During this, the level of habit execution, such as frequency and duration, is superficially maintained despite a decline in motivational drive, raising the question of how the amount of execution is maintained even when decision-making strategies undergo substantial changes. By developing a unique paradigm capable of inducing habit formation within a defined time window in male mice, we found that shift in decision-making and amount of habit execution are controlled by plasticity in distinct cortical pathways. Erasure of each plasticity selectively altered the decision-making strategy or the execution level without affecting the other. These findings reveal a dual regulatory model for habit, providing insights into the neurocircuit mechanisms underlying both adaptive and maladaptive habits.
Habits play a crucial role in maintaining beneficial but effortful behaviors, while maladaptive habits cause refractory, over-engaged behaviors shown in compulsive disorders and behavioral addictions 1 , 2 , 3 , 4 , 5 . Elucidating the regulatory mechanisms of habitual behavior is therefore essential for understanding both adaptive and maladaptive habits 6 , 7 , 8 . Habit formation has been conceptualized as a shift in decision-making, from a flexible, motivation-dependent goal-directed strategy to a more rigid, automatic one 2 , 5 , 7 . While this framework captures qualitative differences in behavior selection, it overlooks the quantitative dimension of habitual execution, such as duration and frequency of the behavior, that is critical for evaluating the real-world utility of adaptive habits and the pathological severity of maladaptive ones. Indeed, during the transition, despite the decline in motivational drive, the level of execution is superficially maintained; however, the mechanism that allows the execution level to be preserved despite significant changes in decision-making strategies remains unclear.
This motivated us to propose that the assessment of habitual behavior should incorporate not only the nature of the underlying decision-making strategy but also an additional dimension: regulation of execution level. While decision-making and execution levels in goal-directed behavior are regulated by evaluations of cost and benefit 9 , 10 , habitual actions, at least at the level of decision-making, proceed automatically without such value-based evaluations. Therefore, a distinct mode of control is likely to govern execution levels in habitual behavior. Based on this premise, we hypothesize the existence of a neural mechanism that stores and maintains the execution level of behavior, dissociable from the mechanisms that govern the transition from goal-directed to habitual strategies.
To address this, we developed a behavior task, which allows us to dissect the neuronal mechanisms responsible for the transitions in decision-making strategies together with determining the execution level in the same cohort of animals. Our analyses revealed that the transition in decision-making strategy and execution level are indeed separately controlled. To analyze the circuit mechanism, we combined this task with ex vivo slice recordings, in vivo Ca 2+ imaging and a recently-developed optogenetic method to erase long-term potentiation (LTP). We found that predominance of a goal-directed strategy and execution level are orthogonally represented as potentiation of synaptic transmission in different prefrontal cortical regions. By targeting region specific plastic changes, we were able to modify the decision-making strategy without affecting execution level or vice versa, supporting the notion that decision strategy and execution level are distinct processes that involve different circuits. Our findings on this dual regulatory model extend the current view of habit regulation and offer insights for understanding and treating maladaptive habits in disorders such as obsessive-compulsive disorder and addiction.
Previous studies have demonstrated that training with ratio-based and interval-based operant conditioning tasks have contrasting outcomes on animal behaviors 3 , 5 , 11 . With ratio-based tasks, where rewards are delivered after a certain number of lever presses, animals acquire goal-directed behavior. In contrast, with interval-based tasks, rewards are delivered on the first lever press after a specific amount of time has passed since the prior reward 12 , 13 . The number and frequency of lever pressing are less contingent on reward deliveries than in ratio-based tasks and the timing of reward deliveries is less affected by subjects’ lever pressing effort. Owing to these factors, interval-based task designs make it difficult for subjects to understand the association between behavioral effort and reward acquisition. Instead, subjects focus on the behavior that previously led to reward delivery; resulting in the predominance of a habitual over goal-directed strategies 3 , 14 .
To date, studies have used these two protocols separately to monitor the acquisition of goal-directed and habitual behaviors 3 , 5 . However, this approach makes identifying the timing of the transition from goal-directed to habitual behaviors difficult, and thus does not allow for the study of the neural mechanisms underlying this transition. Moreover, due to such unclear timing of habit formation, the quantitative changes in action execution accompanied by habit formation could not be adequately analyzed. To overcome these issues, we performed ratio- and interval-based tasks in a sequential fashion in the same cohort of mice and evaluated within-individual transition of decision-making strategy together with that of action execution (Fig. 1a ). First, goal-directed behavior was established with a ratio-based task, after which habitual behavior was invoked by switching to an interval-based task, allowing us to track within-subject transitions of decision-making strategy and action execution. Mice were first trained with a continuous reinforcement (CRF) task in an operant chamber for 3 days (days 1, 3), where each lever press was rewarded with a drop of sucrose solution from a reward port. Following this, they were trained with a ratio-based, variable ratio (VR) task for 6 days (days 4, 9), being rewarded after 10 presses on average on days 4, 5, and 20 presses on days 6, 9. The lever press rate during VR training increased, indicating successful learning of the action-outcome (lever-reward) association 3 , 5 (Fig. 1b ). On day 10, the task was seamlessly switched to an interval-based, variable interval (VI) task, where animals were rewarded on the first lever press after a variable time interval from the last reward (ranging from 30 to 90 s, average 60 s). When switched to the VI task, mice slightly but significantly decreased the frequency of lever pressing, compared with animals that continued the VR task (VR + VR group) (Fig. 1b ).
a Schematic diagram of the two-step operant training protocol. On day 10, the training was seamlessly switched from a variable ratio (VR) to a variable interval (VI) task (VR + VI). Mice without the VR → VI switching (VR + VR) are considered as the control group. b Transitions of lever pressing rate in the two-step training. VR + VR, N = 19; VR + VI, N = 79. c Mice underwent the first devaluation test after the 6th VR session. Then, after the 4th day of VI or VR training, the second devaluation test was performed. On each test day, mice were fed with regular chow, or sucrose solution. With free access to sucrose solution, its relative value declined (devalued condition). After 30-min free-access, mice were allowed to press the lever for 5 min, during which no sucrose was delivered. Numbers of lever presses on each test day in the VR + VR ( d ) and VR + VI ( e ) groups. f Within-subject differences in lever pressing between valued and devalued conditions, are shown as the devaluation index. d, f VR + VR, N = 13; VR + VI, N = 17. g Lever pressing activities before and after habit formation were compared in a within-subject manner. h Within-subject differences in the lever press rate between days 9 and 13 are plotted as the execution index. The VR + VI group was subdivided into the VR+VI High and VR+VI Low groups by the median value of the execution index. i Lever press rate in the VR+VI High and VR+VI Low groups. j Single session lever press rate on days 9 and 13 in the VR+VI High and VR+VI Low groups. k Comparison of the lever press rate on days 9 and 13, reveal a significant difference in the distribution only on day 13. l The number of lever press bouts on days 9 and 13. VR → VI switching reduced lever pressing bout in the VR+VI Low group. m , n Reward presentation and nose poke rates in the VR+VI High and VR+VI Low groups. o Relations between the execution index and within-subject differences in the nose poke rate between days 9 and 13 are plotted in the XY-plane. h, o VR+VI High N = 40; VR+VI Low N = 39. p , q Number of lever presses in each devaluation test in the VR + VI group ( e ) are separated into the VR+VI High and VR+VI Low groups. p N = 9, q N = 8. r Within-subject differences in lever pressing between the valued and devalued conditions are shown as the devaluation index. Relations between the execution index and the devaluation index in the first and second devaluation tests are plotted in XY-planes. N = 17. ** P < 0.01 and *** P < 0.001; N.S., not significant. For detailed statistics, see Supplementary Data 1 .
To confirm the VR + VI task design led to a transition from a goal-directed to habitual strategy, we conducted two tests: a devaluation test and an omission test 15 , 16 , 17 . In the devaluation test, mice were given free access to either sucrose solution (devalued condition), or regular rodent chow (valued condition) for 30 min. They were then subjected to a 5-min lever-press task during which the sucrose reward was not delivered, without signaling this to the mice (Fig. 1c ). When mice had already consumed sucrose solution, the sucrose reward is considered “devalued”. In contrast, when mice were pre-fed with rodent chow, the reward remains “valued”. On the following day, the same mice were tested in the opposite condition. When mice employ a goal-directed strategy, they reduce the number of lever presses in the devalued condition compared to the valued condition, as they are already satiated with the sucrose solution. Conversely, when they employ a habitual strategy, the number of lever presses is comparable between the valued and devalued conditions, because the reward value does not affect the decision. Because the absolute number of lever presses during the devaluation test is heavily influenced by the lever press rate during the preceding training, we defined the devaluation index for each mouse, which is the difference in lever pressing between the two conditions, normalized by the total number of lever presses (see Methods). This index takes a positive value when subjects rely on a goal-directed strategy and approaches zero when they employ a habitual strategy, irrespective of absolute lever press number 18 . When the devaluation test was conducted after VR training, there was a significant reduction in lever pressing in the devalued condition resulting in a positive devaluation index (1st test; Fig. 1d, f