Learning to ask and answer in specialized documents: Exemplifying through modular integrated construction regulatory documents
Yinyi Wei et al.
What the paper says
Large language models perform well in general question-answering tasks but face challenges in local contextual question-answering within specialized domains due to the high cost of domain-specific dataset curation and unstable model performance. To address these issues, this paper proposes a two-stage framework. In the first stage, “learning to ask,” fine-tuned LLMs generate question-answering pairs from contexts, guided by contextual relevance (i.e., question-context alignment) and answer fidelity (i.e., accuracy and faithfulness of the answer). The second stage, “learning to answer,” systematically compares different CQA paradigms, including fine-tuning, retrieval-augmented generation, and proprietary LLMs. The framework is demonstrated using modular integrated construction regulatory documents. Extensive experiments yield three main insights: (1) Synthetic data generation often mirrors training distributions, necessitating effective filtering; (2) Despite inherent biases, synthetic data retains 90–100% of the performance achieved with original data and appropriate models, demonstrating its practical utility; and (3) Domain-specific, fine-tuned models achieve the best performance, underscoring the importance of tailored adaptation. This work bridges gaps in synthetic data quality assurance and domain-aware language model customization, providing practical guidelines for applications in low-resource, expertise-driven, and privacy-sensitive domains. • Guidelines for local contextual question-answering systems in specialized domains. • Synthetic data generation for extracting question-answer pairs from contexts. • Introduction of quantitative metrics for assessing the quality of synthetic data . • Comparative analysis for establishing contextual question-answering paradigms.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.