LLM-PDM: An LLM Persona-Driven Method for replicating personal mobility preferences at scale
Ioannis Tzachristas et al.
What the paper says
Traditional travel surveys are costly, time-consuming and face declining response rates, motivating the exploration of artificial data generation methods. In this research, we propose a novel persona-driven method for generating synthetic mobility survey data using Large Language Models (LLMs). The method defines representative <em>personas </em>- each characterized by specific sociodemographic attributes - and prompts an LLM to emulate survey respondents with these personas. A guided prompting strategy is introduced to calibrate the synthetic data distributions so that they closely match real-world population statistics. We evaluate the approach on the German <em>Mobilita</em><em>¨</em><em>t </em><em>in Deutschland 2017 </em>(MiD 2017) dataset. The quality of the LLM-PDM-generated synthetic data is assessed against ground truth using a comprehensive set of metrics, including mean absolute error (MAE), root mean square error (RMSE), Jensen-Shannon distance (JSD), entropy, conditional entropy and the Earth Mover’s Distance (EMD). Empirical results demonstrate that the LLM-PDM approach produces high-fidelity synthetic populations that preserve key distributions and relationships present in the real data. Across the case studies, the LLM-PDM method achieves low distributional errors (e.g. MAE <em>< </em>3%) and captures important joint patterns, significantly outperforming a number of LLM baselines.
1 citation
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.16 × 0.4 = 0.06 |
| M · momentum | 0.53 × 0.15 = 0.08 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.