Evolutionary Content Generation via Multimodal LLM-based Fitness Evaluation

Yaqing Hou et al.

IEEE Transactions on Evolutionary Computation2026https://doi.org/10.1109/tevc.2026.3657774article
AJG 4
Weight
0.50

What the paper says

Evolutionary algorithms (EAs) have gained prominence as a powerful optimization tool inspired by biological evolution, excelling in various complex domains. In the context of Generative Artificial Intelligence (GAI), EAs have shown promise in generating diverse, high-quality solutions. However, traditional EAs heavily rely on human-designed fitness functions, which may often lack flexibility and comprehensiveness for different GAI scenarios. Recently, the emergence of Large Language Models (LLMs) has opened new avenues for enhancing the evolutionary process. This paper proposes a formal framework named ECG-LFit, which utilizes an LLM (e.g., GPT-4-Turbo) for multimodal fitness evaluations. We validate the effectiveness of our framework using EAs (CMA-ES, MAP-Elites, and CMA-ME) in our case study on Super Mario game level generation. The results show that our framework improves the quality and playability of the generated levels. Additionally, user studies indicate that participants prefer the levels generated by the ECG-LFit framework, particularly regarding aesthetics, challenge, and playability. Furthermore, to enhance evaluation efficiency, we design a distilled model to simulate the scoring process of the LLM, enabling rapid and effective content evaluation in resource-constrained environments.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1109/tevc.2026.3657774

Or copy a formatted citation

@article{yaqing2026,
  title        = {{Evolutionary Content Generation via Multimodal LLM-based Fitness Evaluation}},
  author       = {Yaqing Hou et al.},
  journal      = {IEEE Transactions on Evolutionary Computation},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1109/tevc.2026.3657774},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Evolutionary Content Generation via Multimodal LLM-based Fitness Evaluation

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.