Evolutionary Content Generation via Multimodal LLM-based Fitness Evaluation
Yaqing Hou et al.
What the paper says
Evolutionary algorithms (EAs) have gained prominence as a powerful optimization tool inspired by biological evolution, excelling in various complex domains. In the context of Generative Artificial Intelligence (GAI), EAs have shown promise in generating diverse, high-quality solutions. However, traditional EAs heavily rely on human-designed fitness functions, which may often lack flexibility and comprehensiveness for different GAI scenarios. Recently, the emergence of Large Language Models (LLMs) has opened new avenues for enhancing the evolutionary process. This paper proposes a formal framework named ECG-LFit, which utilizes an LLM (e.g., GPT-4-Turbo) for multimodal fitness evaluations. We validate the effectiveness of our framework using EAs (CMA-ES, MAP-Elites, and CMA-ME) in our case study on Super Mario game level generation. The results show that our framework improves the quality and playability of the generated levels. Additionally, user studies indicate that participants prefer the levels generated by the ECG-LFit framework, particularly regarding aesthetics, challenge, and playability. Furthermore, to enhance evaluation efficiency, we design a distilled model to simulate the scoring process of the LLM, enabling rapid and effective content evaluation in resource-constrained environments.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.