DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
Quick Answer
This paper shows that The DeepLens Diagnosis Agent utilizes a five-stage workflow with the JSL Medical Small 7B v2 model, achieving 60.14% diagnostic accuracy on the DiagnosisArena benchmark, outperforming traditional models by over 36 points.
Quick Take
It operates at a cost of $0.0072 per case, significantly cheaper than competitors while providing structured outputs for better traceability in medical diagnostics.
Key Points
- Achieved 60.14% top-1 diagnostic accuracy on the 915-case DiagnosisArena benchmark.
- Outperformed traditional models by a 36-point increase in diagnostic reasoning accuracy.
- Costs $0.0072 per case, making it 35-45% cheaper than Claude Sonnet 4.5 and Gemini 3.1 Pro.
- Structured intermediate outputs enhance traceability and reproducibility in high-stakes medical settings.
- Demonstrated that workflow design can improve diagnostic reasoning beyond mere knowledge recall.
DeepSignal Analysis
What happened
The DeepLens Diagnosis Agent employs a structured five-stage workflow with the JSL Medical Small 7B v2 model, achieving a diagnostic accuracy of 60.14% on the DiagnosisArena benchmark. This performance surpasses traditional models by over 36 points. The agent operates at a cost of $0.0072 per case, making it more economical than competitors like Claude Sonnet 4.5 and Gemini 3.1 Pro.
Key evidence
- The DeepLens Diagnosis Agent achieved 60.14% diagnostic accuracy on the 915-case DiagnosisArena benchmark, outperforming traditional models by over 36 points.
- The JSL Medical Small 7B v2 model, when used without the agent workflow, only achieved 23.99% accuracy, indicating a significant improvement due to the structured workflow.
- The agent operates at a cost of $0.0072 per case, which is 35-45% cheaper than Claude Sonnet 4.5 and Gemini 3.1 Pro.
Why it matters
The DeepLens Diagnosis Agent's structured approach to medical diagnostics highlights the importance of workflow design in enhancing diagnostic accuracy. By achieving a significant improvement in performance while maintaining lower operational costs, it demonstrates that smaller models can effectively compete with larger, more expensive models. This could shift the landscape of medical AI, emphasizing the need for structured reasoning over sheer model size.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagnostic reasoning. We present the DeepLens Diagnosis Agent, a five-stage harnessing pipeline (combining model capabilities with disciplined process constraints) centered on a small medical reasoning model (JSL Medical Small 7B v2) and retrieval-augmented generation (RAG). The pipeline enforces structured clinical extraction, disciplined retrieval, constrained candidate generation, explicit evidence triangulation, and an auditable final decision. On the 915-case DiagnosisArena benchmark, the agent achieved 60.14% top-1 diagnostic accuracy, the highest among small and medium-sized models. The same model without the agent workflow achieved 23.99%, a +36-point gain from workflow design alone, despite 88.2% on standard medical benchmarks, showing that diagnostic reasoning under uncertainty requires more than knowledge recall. The agent costs USD 0.0072 per case (24K tokens on A100) with 24-second latency, 35-45% cheaper than Claude Sonnet 4.5 (USD 0.0110) and Gemini 3.1 Pro (USD 0.0128) while outperforming them by +9.70pp and +9.17pp. Harnessing can also correct frontier model failures; workflow constraints can outweigh parameter count or API cost.
Beyond aggregate accuracy, the pipeline produces structured intermediate artifacts that make each stage inspectable and support error localization. These properties support high-stakes settings where traceability, reproducibility, and auditable evidence matter alongside benchmark performance.
| Comments: | 20 pages, 5 figures, 5 tables. Technical report, John Snow Labs |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.22555 [cs.AI] |
| (or arXiv:2607.22555v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22555 arXiv-issued DOI via DataCite |
Submission history
From: Yigit Gul [view email]
[v1]
Tue, 19 May 2026 12:55:39 UTC (1,073 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.