Hand a language model a difficult clinical case and it will hand you a diagnosis. Quickly, fluently, with an air of total certainty. Quite often it will also be wrong, and the reason usually isn’t that the model lacked the knowledge. It’s that somewhere between reading the case and answering it, the model built an approximate mental copy of the patient, and the approximation leaked.
The interesting part is that the leaks have names, and they aren’t new ones. A single-shot model anchors on the most salient finding and never revises. It closes prematurely on the first coherent story. It drops pertinent negatives, which carry most of the discriminating power in a hard case and none of the narrative weight. Occasionally it confabulates a test nobody ordered. These are the classic cognitive traps of diagnostic error, reproduced faithfully by a system that has no metacognition at all, because nothing in a single forward pass forces it to separate what the chart says from what the chart suggests.
Clinicians are drilled out of these early, and the correction is structural rather than intellectual. You take the history and lay out the findings. You consult what you know. You build a differential. Then you argue the candidates against the evidence and commit. Four steps, with a deliberate pause between each, because the pauses are where errors get caught before they compound.
So we built the assistant that way instead. We call the approach harnessing: combining model capability with disciplined process constraints rather than reaching for more parameters. The result is the DeepLens Diagnosis Agent, a five-stage pipeline built around JSL Medical Small 7B, John Snow Labs’ small medical language model.
The gap that motivates all of this
Start with a number that should be uncomfortable.
JSL Medical Small 7B scores 89.4% on MedQA and averages 88.2% across nine standard medical benchmarks. By the usual measures it knows medicine at a specialist level.
Put the same model on DiagnosisArena, single-shot, and it scores 23.99%.
That is a 64-point cliff, and it isn’t a quirk of one checkpoint. The 32B model from the same family averages 85.0% on standard benchmarks and lands at 22.19% here. Whatever DiagnosisArena is measuring, medical knowledge recall isn’t it.
What it measures is the workflow. The benchmark’s 915 cases are long-form narratives with exam findings, labs and histology, and no answer options to eliminate against. Getting one right requires holding the negatives, keeping the timeline straight, building a real differential and discriminating between close contenders. Every one of those is a process step, and a single forward pass performs none of them explicitly.
Which sets up the actual question. If the knowledge is already in the model, can structure alone get at it?
The result
60.14%. The highest accuracy of any small or medium-sized model on this benchmark, and ahead of the frontier systems evaluated alongside it: +9.17 points over Gemini 3 Pro Preview, +9.70 over Claude Sonnet 4.5, +17.24 over deepseek-reasoner-speciale. GPT-5.2, which we benchmarked separately, remains the strongest single model on DiagnosisArena at 62.73%.
Scoring was deliberately conservative. Two independent judges, Claude Sonnet 4.5 and Gemini 2.5 Flash, scored every model against the reference label, and we report the average. An answer counts as correct if it matches, uses an accepted medical synonym, or names the reference diagnosis as primary within a multi-diagnosis response. Inter-judge agreement ran 86–87% with identical model rankings. The agent’s two judge scores were 61.86% and 58.43%, a spread consistent with the systematic offset we saw across every model.
Then there’s the comparison that isolates the effect:
Same weights, same temperature-0 sampling, same benchmark, same judges. The only variable is whether the model answers in one shot or runs through the agent. Thirty-six points, with no retraining and no additional parameters.
Where 36 points can possibly come from
The failure log answers this. When we read through the vanilla model’s wrong answers, most were not knowledge gaps. The model frequently knew the condition and still missed the case, because a pertinent negative disappeared somewhere between the narrative and the conclusion, or the differential collapsed into three phrasings of one disease, or the answer arrived buried in enough hedging that it couldn’t be scored cleanly.
Knowledge present, process absent. Which reframes what extra parameters would even buy. If a case is lost because a negative immunoblot evaporated between paragraph two and the conclusion, no amount of additional parametric recall recovers it. The fix is to make the negative impossible to lose.
That suggests a division of labour, and it’s the one the agent implements:
- the small medical model executes clinical steps under constraint, which is what it’s tuned for
- retrieval supplies the breadth of literature the model doesn’t need to memorise
- the workflow supervises the process across steps, catching failures no single forward pass can catch alone
Each of the five stages below closes one of those process failures. They’re worth reading as verification gates rather than as prompt engineering, because that’s the level the gains come from.
The five stages
Each stage maps onto something a clinician already does. The difference is that here the boundaries between them are enforced by the system rather than by training and discipline.
1 · Extraction, the structured history and physical. The case is rewritten exactly once into a fixed schema: timeline, demographics, positives, explicit negatives, diagnostics one test per line with units preserved. From that moment the table is the only patient truth in the system, and nothing downstream may reason from its own recollection of the narrative. Pertinent negatives are why this matters most. They carry the discriminating power in a hard case and decay first in prose, because a negative immunoblot is a fact about something that isn’t there. As table rows they stop being optional.
2 · Retrieval, on a short leash. One brief query, carrying no ages, years, names or institutions, runs a hybrid keyword and embedding search across Semantic Scholar plus John Snow Labs’ curated knowledge base, roughly 200M scholarly records, in milliseconds. The compiled note is stripped of diagnosis labels and citations before the model sees it, so retrieval may suggest what to check but never hand over the answer. The agent also pulls a few pattern triggers, reusable clinical rules of thumb that point the differential in a direction without naming a disease.
3 · The differential, kept cramped. Usually four candidates, each a label, a likelihood from 1 to 10 and one short reason. Give this step more room and it writes a second essay that contaminates everything after it. The gates are the ones any attending applies to a resident’s list: candidates mutually exclusive, coverage two-track (common plausible alongside rare high-fit), a practical discriminator per candidate, and a unifying-diagnosis check, because one condition explaining everything beats three that each explain a third.
4 · Evidence triangulation, where correctness gets decided. The agent writes a FOR/AGAINST memo under one rule: every bullet opens with a verbatim quote from the extracted facts, or says Unknown. A filter then walks the finished memo and deletes any bullet whose quote doesn’t exactly match a fact in the table. Paraphrased evidence never reaches the decision, invented evidence never reaches the decision, and Unknown is a first-class outcome rather than a silent omission.
5 · Commit. One label, exactly matching a generated candidate string, plus a confidence from 0 to 10 and reasoning that cites patient quotes. It cannot contradict an extracted negative, and when nothing fits it must return insufficient evidence and name the missing discriminator rather than reach for the nearest plausible disease.
Temperature 0 throughout. Support components fail soft, so a pattern-trigger timeout never kills a case, while the three core steps are validated and repaired.
One case, start to finish
A man in his fifties, two years of painful oral erosions, recurrent respiratory infections and chronic diarrhoea. Whitish lacy patches with erosions on the tongue and buccal mucosa, skin and nails spared. Lingual biopsy showing acanthosis, scattered necrotic keratinocytes, basal-layer vacuolation and a band-like lymphohistiocytic infiltrate, with DIF showing fibrinogen along the basement membrane. A 10-cm hypercapturing mediastinal mass on PET, resected and reported as type AB thymoma, Masaoka-Koga 2A.
The labs are where it gets interesting. IgG 465 mg/dL, normal IgA and IgM, decreased total B lymphocytes, raised CD8 count with an inverted CD4:CD8 ratio. And then a run of negatives: indirect immunofluorescence on monkey oesophagus, salt-split skin and rat bladder all negative, ELISA and immunoblot negative across multiple antibodies, hepatitis B and C negative, liver function normal.
An interface dermatitis with fibrinogen-only DIF is nonspecific, and the mass is the loudest thing in the chart. This is a case built to reward anchoring.
| Stage | What came out |
| Extraction | 7 positives, 3 negatives, 7 diagnostics locked in |
| RAG | 60 papers searched, compact note compiled, no labels |
| Triggers | 10 heuristics retrieved as scaffolding |
| Candidates | Good Syndrome 9 · Linear IgA Bullous Dermatosis 7 · Mucous Membrane Pemphigoid 6 · Type AB Thymoma 5 |
| Triangulation | Memo built over 50 more papers, every claim tied to a quoted fact |
| Decision | Good Syndrome, confidence 9/10 |
Look at what did the eliminating. Linear IgA bullous dermatosis went out because IgA deposition was never demonstrated. Mucous membrane pemphigoid went out for absent IgG and C3. The thymoma went out as a standalone answer because a thymoma does not by itself account for hypogammaglobulinaemia, B-cell lymphopenia and an inverted CD4:CD8 ratio.
Every one of those eliminations rests on a negative or on an explanatory gap rather than on a positive finding, which is precisely the evidence class a single-shot reader skims past. What survived is the unifying diagnosis: thymoma with immunodeficiency, the constellation that ties the mass, the recurrent sinopulmonary infections, the diarrhoea and the mucosal disease into one condition instead of three coincidences. Ten LLM calls, start to finish.
Cost: why small models plus harnessing wins on economics
The intuition says a five-stage agent must cost more than one API call. It doesn’t.
0.0072percase, 24K tokens across the full pipeline on self-hosted A100 infrastructure. Claude Sonnet 4.5 costs $0.0110 for a single call, 53% more. Gemini 3.1 Pro costs $0.0128, 78% more. Both are also nine to ten accuracy points behind.
That arithmetic is the whole argument for small models in agentic systems. A 7B’s per-token cost is low enough that you can spend ten calls on a case and still undercut one frontier call, so every reliability behaviour you’d want, retrieving context, re-checking grounding, verifying negatives survived, gating candidates, becomes the default instead of a budget conversation. On a frontier API the same ten-call workflow multiplies an already higher price by ten.
At scale it compounds: 100,000 cases a year costs 720against1,100–$1,280 on cloud APIs, a 35–45% reduction while delivering higher accuracy. Self-hosting also keeps patient data inside your network and takes vendor lock-in off your risk register.
When something goes wrong, you can see where
Accuracy gets the headline. Traceability is what makes a system like this deployable in a hospital.
Every run leaves its work behind: the fact table, the titles of the papers actually pulled, the scored candidates, the quote-anchored memo, the final label and its reasoning. When someone disputes an output, nobody reverse-engineers a paragraph of confident prose. You open the stage that broke. Retrieval drift, extraction miss, or a genuine reasoning error, and you can usually tell which inside a minute.
Smaller models help here too. A 7B model’s behaviour is easier to probe and explain than a 100B+ generalist, which matters when regulators and clinicians expect to see the reasoning rather than take it on trust.
Conclusion
Diagnostic reasoning is not the same capability as medical knowledge, and the 64-point gap between JSL Medical Small 7B’s benchmark scores and its single-shot DiagnosisArena result makes that concrete. The knowledge was already there. What was missing was process.
Harnessing supplies the process. Fact lock-in, disciplined retrieval, gated candidates and quote-anchored verification took the same 7B weights from 23.99% to 60.14%, the best result of any small or medium model on this benchmark, ahead of frontier systems costing 53–78% more per case.
For high-stakes reasoning, workflow constraints can matter as much as raw model capacity. And if you’re running a single prompt against a frontier API today, the cheapest accuracy available to you probably isn’t a bigger model.
These outputs are decision support. A clinician stays in the loop.
Talk to us
John Snow Labs builds medical language models and clinical agents for healthcare enterprises. If you want to evaluate JSL Medical LLMs or the DeepLens agent suite on your own cases, in your own environment, we’ll help you scope it.
- Pick the model that fits the job. JSL Medical LLMs span 7B and 9B through 27B, 30B and 70B, including multimodal variants for cases where the imaging and the document matter as much as the text. The 7B carries this agent, but the same workflow discipline applies at every size, and we’ll help you find the accuracy, latency and hardware balance your deployment actually needs.
- See it on your data. We’ll run the Diagnosis Agent against a de-identified sample from your setting and share the full per-stage artifacts, not just the labels.
- Deploy where your data lives. Self-hosted, on your infrastructure, with no patient data leaving your network and no per-token metering.
- Build beyond diagnosis. The agent slots into the wider DeepLens suite, so its structured differential and evidence memo feed medication safety, guideline retrieval and follow-up question generation directly.
Get in touch with John Snow Labs → · Explore JSL Medical LLMs →
Reference: Mahmood Bayeshi, Veysel Kocaman, Muhammed Ali Naqvi, Yigit Gul, David Talby (John Snow Labs). DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs. Full methods, ablations and a complete annotated end-to-end run: arxiv.org/html/2607.22555v1


































