HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule

arXiv cs.CL·Xi Xuan, Wenxin Zhang, Yufei Zhou, King-kui Sin, Chunyu Kit

3h ago

·~2 min·6/8/2026·en·0

Quick Answer

The HKJudge dataset is the first expert-annotated legal discourse corpus for Hong Kong judgments, featuring 290k sentences and 6.5 million tokens.

Quick Take

The HKJudge dataset is the first expert-annotated legal discourse corpus for Hong Kong judgments, featuring 290k sentences and 6.5 million tokens. It includes a two-tier discourse schema for rhetorical role classification and legal element extraction, benchmarked against four BERT-based models and various LLMs, demonstrating significant potential for legal judgment prediction.

Key Points

HKJudge contains approximately 290k sentences and 6.5 million tokens.
The dataset features a two-tier discourse schema with 26 rhetorical roles.
Annotations were produced by ten legal linguistics experts with a kappa of 0.8.
Benchmark evaluations include four BERT-based models and various commercial LLMs.
The dataset is available for further research in legal judgment prediction.

Article Content

From source RSS / original summary

arXiv:2606. 06679v1 Announce Type: new Abstract: Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora. We introduce the Hong Kong Judgment Discourse Dataset (HKJudge), the first sentence-level expert-annotated legal discourse corpus. HKJudge includes criminal judgments across all five levels of HK's court hierarchy, comprising $\sim$290k sentences and $\sim$6.

5 million tokens, fully annotated by legal linguistics experts. We design a two-tier discourse schema that captures what facts a court finds, how it reasons, and what it rules. At the sentence level, each sentence is assigned one of 26 rhetorical roles. At the span level, sentences are further annotated with three sentencing elements (charge, imprisonment term, fine). Ten legal linguistics annotators produced the annotations with an inter-annotator agreement of $\kappa = 0. 8$.

We formulate two tasks on HKJudge, termed rhetorical role classification and legal element extraction, and provide the first benchmark evaluation of four BERT-based models, two open-source LLMs under zero-shot and fine-tuning settings, and four commercial LLMs on both tasks. Our work demonstrates the value of sentence-level discourse annotation for modeling the structure of HK judgments and provides a rich data foundation for future work on legal judgment prediction.

The HKJudge dataset and code are available at https://github. com/xuanxixi/HKJudge.

Reader Mode unavailable (could not extract clean content).

Read on arxiv.org

Want this in your inbox every morning?

Daily brief at your local 8am — bilingual EN/中文, free.

Subscribe — it's free

More from arXiv cs.CL

See more →

arXiv cs.CL·Leyao Wang, Yanan He, Peng Chen, Asaf Yehudai, Yixin Liu, Rex Ying, Michal Shmueli-Scheuer, Arman Cohan

2w ago

FeaturedOriginal

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

AI Summary

The REFLECT benchmark reveals that current LLM judges are unreliable, achieving below 55% accuracy in evaluating reasoning and evidence use, highlighting the need for improved evaluation methods for deep research agents.

#LLM #Agent #Inference #Policy