Double XCSF on Target?

Connor Schönberner & Sven Tomforde

Evolutionary Computation2025https://doi.org/10.1162/evco.a.377article
AJG 3
Weight
0.50

What the paper says

The XCS Classifier System (XCS), the most prominent Learning Classifier System (LCS), originally focused on Reinforcement Learning (RL) problems. Over time, emphasis shifted heavily to supervised learning, with some applications in unsupervised learning. Following rekindled interest in LCSs for RL domains, we intend to capitalise on the close relationship between Q-learning and XCS. Except for Experience Replay, hardly any advances built on Q-learning have been investigated in XCS variants such as XCSF. Recognising this, we introduce three extensions inspired by Q-learning derivates: Target prediction inspired by DQN's target networks to improve the learning stability and double target prediction inspired by Double DQN as well as a Double Q-learning mechanism as countermeasures against overestimation. Addressing these two issues, aims to improve the performance of XCSF and the high variance between runs. We apply them to the Maze Problem, Frozen Lake, and Cart Pole. Our observations indicate mixed results: The Double Q-learning mechanism leads to no improvement. Target and double target prediction can lead to observable and also significantly improved performance and can provide variance reduction. This underscores that improving the RL capabilities of XCSF is non-trivial but indicates that adapting Deep Reinforcement Learning mechanisms for XCSF can be advantageous.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1162/evco.a.377

Or copy a formatted citation

@article{connor2025,
  title        = {{Double XCSF on Target?}},
  author       = {Connor Schönberner & Sven Tomforde},
  journal      = {Evolutionary Computation},
  year         = {2025},
  doi          = {https://doi.org/https://doi.org/10.1162/evco.a.377},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Double XCSF on Target?

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.