
An AI system helped Pakistani judges clear massive backlogs at $38.50 return per dollar invested
Quick Answer
An AI system, JudgeGPT, helped Pakistani judges increase case resolution by 1,848 cases per year per district, yielding a $38.50 return on every dollar invested.
Quick Take
The study involved 1,559 judges across 118 courts, demonstrating that targeted training significantly enhanced AI usage and judicial productivity without compromising judgment quality.
Key Points
- JudgeGPT, based on GPT-4, utilized 129,235 documents for legal queries.
- Trained judges averaged nearly 60 logins and 200 prompts compared to 20 logins in the control group.
- Ruling quality improved, with trained judges rated better in 59% of comparisons.
- AI use did not increase bias in judicial language, maintaining fairness.
- Productivity gains are likely conservative; newer models could yield even better results.
DeepSignal Analysis
What happened
A study involving 1,559 judges across 118 Pakistani courts demonstrated that the AI tool JudgeGPT, combined with targeted training, led to an increase of 1,848 resolved cases per year per district. The investment yielded a return of $38.50 for every dollar spent. The training focused on appropriate AI usage, resulting in improved judgment quality without increasing bias.
Key evidence
- The study involved a randomized trial with 1,559 judges across 118 courts, representing roughly half of all trial court judges in Pakistan.
- Judges who received targeted training used JudgeGPT four times more than those who only attended a general seminar, averaging nearly 60 logins and over 200 prompts.
- The appeal rate per 1,000 resolved cases decreased slightly, indicating that judgment quality remained stable or improved while judges maintained their work-life balance.
Why it matters
This research highlights the potential of AI to enhance judicial productivity in Pakistan, suggesting that effective training is crucial for maximizing AI benefits. The significant return on investment indicates that integrating AI tools can be a cost-effective strategy for addressing judicial backlogs. However, the findings also emphasize that AI should not replace human judgment, as the final decisions remain with the judges.
📖 Reader Mode
~3 min readResearchers from ETH Zurich, the New Economic School, and Imperial College London ran a large-scale field experiment with Pakistan's judiciary. The randomized trial covered 1,559 judges across 118 courts, roughly half of all Pakistani trial court judges.
The tool was JudgeGPT, an AI assistant built on OpenAI's GPT-4 and designed for Pakistani trial courts. It uses retrieval augmented generation to search a database of 129,235 documents, including 128,292 court rulings and 943 Pakistani laws. When a judge enters a query, JudgeGPT picks the ten most relevant passages and generates a cited answer.
Trained judges resolve 1,848 more cases per year per district
The researchers split judges into three groups. One got JudgeGPT access plus targeted training: six 90-minute lectures over three weeks, taught by ETH Professor Elliott Ash after court hours. Judges learned which tasks suited the tool, where it fell short, and how to check its output.
A second group got the same AI access but only a general seminar on technology and law. The control group attended that seminar with no JudgeGPT access.
AI access alone did little. Judges with targeted training used JudgeGPT four times as much as those in the general seminar group. After 40 weeks, trained judges averaged nearly 60 logins and over 200 prompts. The comparison group averaged about 20 logins and fewer than 50 prompts.
Districts with more trained judges resolved more cases. At moderate exposure levels, that meant about 1,848 extra cases per year per district, a 6.3 percent bump. Even districts in the bottom quartile still cleared about 616 more cases.
Judgment quality ticks up without added bias
Ruling quality held steady or improved. The appeal rate per 1,000 resolved cases fell slightly, and judges worked the same hours with no change in work-life balance. The researchers estimate savings of about $38.50 per dollar invested, based on what it would cost to hire enough extra judges to match the same output. Even conservative estimates put the return at "at least" $10 per dollar.
A review of roughly 4,000 court judgments found more AI-flagged text, as expected. But readability, length, and the number of legal arguments held steady. An LLM-based quality check, validated by two Pakistani lawyers, showed a slight improvement. Rulings from trained judges were rated better in 59 percent of pairwise comparisons, up from 42 percent in the control group. The study found no evidence that AI use increased gender or religious bias in judicial language.
Hands-on training shapes how judges use AI
The researchers reviewed anonymized chat logs from about 1,500 judges. Legal research, text editing, and text generation were the most common tasks. Around 60 percent of queries sought information about laws, procedures, or legal concepts.
Trained judges used JudgeGPT more for editing and summarizing text, tasks where language models are more reliable. They asked fewer broad legal questions, where hallucination risk is higher. The researchers say the training steered judges toward limited support tasks while leaving the final decisions in their hands.

Only about a fifth of requests involved what the researchers call "substantive AI delegation," where judges asked JudgeGPT to evaluate decisions or draft reasoning on its own. Training made judges more likely to decide cases themselves and only use AI to write up their reasoning.
The researchers stress that their findings don't support replacing judges with AI. What the experiment shows, they say, is that AI can boost public sector productivity, but only when paired with training that steers users toward the right tasks. Without it, most of the gains disappear.
The productivity gains in this study are likely a floor, not a ceiling. JudgeGPT ran on GPT-4, a pre-reasoning model that was the best fit for the project at the time. Today's reasoning models write better, hallucinate less, and handle complex tasks more reliably.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

