Cognitive Bias and Reassignment: Who Can Contribute High Quality LLM Data
Abstract
In recent years, the rapid development of Large Language Models has highlighted the urgent need for large-scale, high-quality, and diverse data. We have launched an LLM data co-creation platform aimed at bringing together a wide range of participants to contribute data. Within six months, the platform has attracted over 10,000 participants who contributed more than 150,000 data entries across more than 200 tasks. An observable user cohort was constructed around the question, "Who is the best data contributor?" along with sub-questions concerning user preferences, task competence, and more. Through a detailed analysis of data contributors, this paper reveals several data collection patterns related to human factors. It reveals that contributors who provide high-quality data often do not meet initial expectations, as their behavior exhibits typical characteristics of the Dunning-Kruger effect. This paper examined the cognitive bias between users' self-assessment and actual abilities, where individuals tend to overestimate their capabilities in certain tasks, leading to a decreased willingness to continue contributing and a consequent waste of human resources. To address this issue, we propose a task reassignment method based on multi-task fine-tuning of small language models (SLMs) to better align user groups with appropriate task types. After the reallocation, we observed a significant increase in user engagement and platform benefits, along with improved overall platform efficiency. The versatility of this method makes it applicable to broader data collection scenarios.
Cite
Text
Gao et al. "Cognitive Bias and Reassignment: Who Can Contribute High Quality LLM Data." AAAI Conference on Artificial Intelligence, 2025. doi:10.1609/AAAI.V39I27.35018Markdown
[Gao et al. "Cognitive Bias and Reassignment: Who Can Contribute High Quality LLM Data." AAAI Conference on Artificial Intelligence, 2025.](https://mlanthology.org/aaai/2025/gao2025aaai-cognitive/) doi:10.1609/AAAI.V39I27.35018BibTeX
@inproceedings{gao2025aaai-cognitive,
title = {{Cognitive Bias and Reassignment: Who Can Contribute High Quality LLM Data}},
author = {Gao, Yunfan and Xiong, Yun and Hu, Zhongyuan and Zhang, Yiming and Wang, Meng and Wang, Haofen},
booktitle = {AAAI Conference on Artificial Intelligence},
year = {2025},
pages = {28007-28014},
doi = {10.1609/AAAI.V39I27.35018},
url = {https://mlanthology.org/aaai/2025/gao2025aaai-cognitive/}
}