Can AI Guess What You Know? Performance Comparison of Large Language Models for Human Domain Knowledge Estimation From Communication Logs
Quick Answer
This paper shows that Large Language Models (LLMs) like Gemini 2.5 Flash can estimate individual domain knowledge from Slack logs, achieving a low MAE of 21.13%.
Quick Take
In contrast, GPT models showed larger discrepancies, indicating that message volume alone does not enhance inference accuracy. This research underscores the potential and limitations of automated expertise mapping in organizations.
Key Points
- Gemini 2.5 Flash achieved the lowest mean absolute error (MAE) at 21.13%.
- GPT models exhibited significantly larger discrepancies in knowledge estimation.
- Estimation accuracy was weakly dependent on the volume of messages analyzed.
- Study analyzed 27,188 messages from 43 users over long-term Slack logs.
- Findings highlight the need for privacy-preserving methods in expertise mapping.
Paper Resources
Article Excerpt
From source RSS / original summaryarXiv:2605. 22971v1 Announce Type: new Abstract: Employees often struggle to identify ``who knows what,'' leading to organizational productivity losses. We investigate whether (LLMs) can infer individual domain knowledge directly from long-term Slack logs. Analyzing 27,188 messages from 43 users, we evaluated seven models (including Gemini, Claude, and GPT families) by comparing their zero-shot estimates against self-reported skill ratings from 27 participants. Gemini 2.
5 Flash achieved the lowest error (MAE 21. 13%), while GPT models showed significantly larger discrepancies. Notably, estimation accuracy depended only weakly on message volume, indicating that more text alone does not guarantee better inference. These findings demonstrate the feasibility and current limits of automated expertise mapping, highlighting the need for privacy-preserving deployments and richer, structure-aware representations of human knowledge.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.