MLCR v1.1: Calibrating LLM Judges for Long-Context Medical Reasoning

Earlier this summer, we released the Medical Long Context Reasoning benchmark, or MLCR, to measure how well large language models reason across long, fragmented medical records. Wisedocs is proud to announce we have now released MLCR with Artificial Analysis as MLCR-AA, an independently run evaluation of frontier closed- and open-weight models.

MLCR v1.1: Calibrating LLM Judges for Long-Context Medical Reasoning

MLCR v1.1 is primarily a judge-calibration release and this blog post focuses on the following: 

  1. How we calibrated conciseness, completeness, and accuracy of responses
  2. A quick overview of some interesting findings from MLCR-AA.

With regards to calibration, the following specific changes were made which we will discuss in detail in this blog post:

  1. Separated accuracy from completeness more explicitly. Accuracy now asks “is anything wrong?”, while completeness asks whether the answer includes the necessary reasoning, evidence, and requested components. This means the missing information is no longer considered an accuracy failure, but only a completeness failure.
  2. Grounded judges against the source documents, not just the reference answer. This was probably the biggest calibration change: the record became the factual authority, while the gold answer acts more as a guide for what a good answer may cover.
  3. Relaxed the conciseness gate from 3× → 5× the human reference length. 3× was too aggressive for difficult synthesis questions and penalized otherwise strong answers from frontier models; 5× still discourages unnecessary verbosity without over-penalizing detailed reasoning. 
  4. Explicitly treated negative findings as findings. In v1.1, judges are instructed that evidence such as “no provider documented X,” “no reason was given,” or the absence of an expected finding is clinically meaningful. 
  5. Prompt wording adjustments to boost judge agreement rates. In v1.0, while the judges were calibrated to agree on easier question tiers, disagreement persisted in more difficult tiers. Minor adjustments to language helped various judges coalesce around the correct definitions of accuracy and completeness even on more difficult questions. The goal here was not to manufacture consensus, but to ensure judge models correctly understood the intended rubric.
  6. Cleaned up the evaluation inputs/configuration. We clarified/corrected reference answers where needed.

A quick background

Earlier this summer, we released the Medical Long Context Reasoning benchmark, or MLCR, to measure how well large language models reason across long, fragmented medical records. 

A claims professional reviewing a record that spans dozens or hundreds of visits is not simply looking up an isolated diagnosis or date. They are reconstructing a narrative and synthesizing conclusions: what happened, in what order, how the claimant’s condition evolved, where providers agreed or disagreed, and what different pieces of evidence means for the claim.

MLCR was designed to test the performance of LLMs on this work that a claims professional typically engages in. The benchmark contains 10 synthetic, real-world-inspired medical cases ranging from approximately 25,000 to 64,000 tokens. Each case contains 50 to 150 medical summaries, with questions distributed across six difficulty tiers.

During the V1 release, we opensourced the cases and easier question tiers publicly, along with an open-source testing harness. The most difficult questions were kept private to reduce contamination and preserve their usefulness as a held-out evaluation set.

For our evaluation, we preserved the same general architecture of the scoring pipeline, as illustrated in the diagram below:

From MLCR to MLCR-AA

We have now released MLCR with Artificial Analysis as MLCR-AA, an independently run evaluation of frontier closed- and open-weight models.

Artificial Analysis evaluates the private set of 60 questions drawn from MLCR’s three difficult question tiers: expert synthesis. compound, and multi-part reasoning tiers. Each question is run three times, allowing results to account for some of the variation that occurs between model runs. 

This partnership gives MLCR a continuously updated home where models can be compared not only by their overall score, but also by accuracy, completeness, output length, cost, token usage, and speed.

Anchoring accuracy to the source record

One of the most important calibration changes in MLCR v1.1 was anchoring accuracy to the actual source documents rather than treating the human reference answer as the sole authority.

The reference answer remains valuable: it captures the expected conclusion and the key information a strong response should contain. However, it is ultimately a compressed representation of a much longer medical record. A model response should not fail on accuracy simply because it uses different wording, includes additional facts supported by the source, or reaches a reasonable inference that is consistent with the record.

We also drew a much clearer boundary between “wrong” and “missing.” In v1.0, accuracy and completeness were partially conflated within the same pass/fail rubric. For example, the failure criteria placed “Detail from the reference is missing” alongside factual errors such as “Any date, name, number … is altered.” This meant an omission could incorrectly cause an accuracy failure.

In v1.1, we made that distinction explicit:

“accurate=false rationale must cite a WRONG claim (not an omission).”
“Self-check: if you wrote ‘missing X’ as a rationale for accuracy, change to accurate=true.”

Missing information is now handled by the completeness rubric, while accuracy is reserved for claims that are unsupported, contradicted, or factually incorrect.

We also strengthened protections for answers that differ from the reference without being wrong. Equivalent medical terminology and abbreviations are accepted, additional facts supported by the source are permitted, and reasonable inferences that do not contradict the record are not treated as accuracy failures. The last point is especially important for difficult questions that explicitly require interpretation, for example:

“What does the lumbar RFA duration pattern … suggest about the nature of the facet-mediated pain component?”

There may not be a single sentence in the record that directly answers such a question. A strong response must synthesize the available evidence and make a supported inference.

At the same time, we wanted to avoid making the accuracy rubric overly permissive. We therefore added a final verification step:

“VERIFICATION: Before setting accurate=true, check each specific claim (visit numbers, dates, attributions, conclusions) individually against the source. Do not pass based on overall narrative coherence.”

This proved particularly important because some judge models like Claude Opus 4.8 would often classify an answer as accurate when its overall clinical narrative was directionally correct, even if individual claims cited were not. The additional verification step forces the judge to evaluate those claims independently, increasing the likelihood that subtle factual errors are caught rather than being overlooked.

Completeness is the most important differentiator

Accuracy asks whether anything in an answer is wrong. Completeness asks whether anything important is missing.

A model can produce a response containing several entirely correct statements while omitting a crucial fourth fact. That answer may be accurate, but it is incomplete. Similarly, a model may reach the right conclusion while failing to explain the treatment history, negative findings, provider disagreements, or other evidence required to support it. For MLCR, completeness therefore means covering the reasoning the question asks for, not just stating a plausible conclusion.

The v1.0 prompt treated completeness as a relatively shallow two-part check:

“Field coverage: every field the question asks for has a corresponding answer.”
“Detail completeness: all factual detail present in the reference answer is captured.”

This worked well for easier tier questions, but it was insufficient for MLCR’s more difficult tiers, where a fluent conclusion also required a sound chain of reasoning.

In v1.1, we defined completeness much more explicitly:

“Completeness = same reasoning chain + comparable specific evidence + all key qualifiers from the REFERENCE ANSWER.”

We also required that:

“Every intermediate analytical link in the REFERENCE ANSWER's reasoning chain must appear. If the reference argues A→B→C→D, all links must be present.”

This distinction matters because many of MLCR’s harder questions ask for a chain of reasoning, not simply a fact. Consider:

“What does the evolution of the surgical candidacy assessment across the four neurosurgical consultations suggest about what drove the progressive narrowing of the surgical option?”

A response that simply states the final interpretation is not necessarily complete. It must reconstruct the progression across the consultations and explain how that evidence supports the conclusion.

We also made an important medical-domain clarification in v1.1: negative findings are themselves meaningful findings.

The updated rubric states:

“NEGATIVE FINDINGS (‘no provider documented X,’ ‘no clinical rationale was given,’ ‘the record lacks Y’) are key analytical findings, just as important as positive findings. If the reference highlights an absence or gap, the model must include it.”

This is particularly important for MLCR’s trap questions, some of which are deliberately constructed around the distinction between something being implied and something actually being documented.

For example, if a question in part asks:

“Was a formal documented refusal of surgery recorded anywhere in the record?”

A complete answer must identify that no formal refusal was documented.

We also changed how completeness handles supporting evidence to allow for comparable, but not identical evidence.

The v1.0 requirement to capture “all factual detail” from the reference could make the judge overly sensitive to differences between an otherwise strong response. Two responses may support the same conclusion using different, but equally valid visits, dates, and findings from the source record.

In v1.1, we instead require comparable evidence depth:

“If the reference cites specific evidence (visits, dates, scores), the model must cite specific evidence at comparable depth—not vague summaries. It need not cite the EXACT same data points, but must provide specifics for ~80% of claims where the reference does.”

The goal is to ensure responses carry high quality evidence when the question requires it, without forcing the model to reproduce it verbatim to be considered complete. If the reference supports a conclusion using four clinically relevant visits, a model that provides no visits should fail completeness. On the other hand, if a model that cites a partially overlapping set of relevant visits at comparable depth, the answer should be considered complete. 

Across our calibration, we found evidence depth is one of the more difficult aspects of LLM-as-a-judge evaluation. Even with the improved rubric, judge models can struggle to determine when two sets of evidence are genuinely equivalent. In future versions of MLCR, we expect to explore this form of evidence equivalence further.

Together with the accuracy-calibration changes described in the previous section, these updates made the distinction between factual precision and answer coverage much more visible.

At the time of writing, GPT-5.6 Terra and GPT-5.6 Sol were the two highest-scoring models on accuracy among judged responses, at 93.7% and 92.5%, respectively. Their completeness scores, however, were substantially lower at 33.9% and 27.6%.

Claude Opus 5 excelled in completeness. The strongest Opus 5 configuration reached 86.1% completeness, leading the benchmark on that metric, while its accuracy score of 88.6% remained below the top GPT-5.6 configurations.

This gap illustrates why a model can look exceptional on accuracy while finishing lower on the overall benchmark. An answer can be accurate but incomplete, or complete but inaccurate. Even models that perform strongly on both metrics independently do not necessarily pass both on the same response. For example, the best Claude Opus 5 configuration reached an overall score of only 59.4%. This gap between individual metric performance and joint success is important: it shows that MLCR remains far from saturated, and that reliably producing answers that are simultaneously accurate, complete, and concise remains a difficult frontier capability. It also further shows that the judges are able to distinguish between the two.

Why the conciseness threshold moved from 3× to 5×

The original scoring configuration used a conciseness threshold of three times the length of the human reference answer.

During calibration, that proved too strict, especially for frontier models like Fable 5.

Human-validated references are written by someone who already understands the case and can express the answer succinctly. On the other hand, we found most frontier models are naturally verbose and a 3x clip sometimes penalized correct answers.

While most strong responses were still relatively compact (1-3x the human reference length), the distribution has a long tail. On some difficult questions, model answers expanded to as much as 75×.

We therefore moved the threshold to 5× the gold-standard answer. This threshold filters out responses that bury the result in pages of repetition or unnecessary background. At the same time, it avoids penalizing an otherwise strong answer.

Judges are less likely to agree on difficult question tiers

Using three judges does not mean that all three interpret an answer in the same way.

During calibration, we found that the judges exhibited different tendencies. Within this evaluation:

  • Claude Opus 4.6 judge tended to reward the overall clinical narrative at the expense of minor details
  • GPT 5.5 behaved more like a strict checklist, penalizing small omissions
  • Gemini 3.1 Pro generally landed between those two positions.

These should not be interpreted as universal characteristics of the underlying model families. They are behaviors observed in a specific judging task, with a specific rubric and set of answers. But they demonstrate why relying on a single judge can create a hidden source of benchmark bias. In other words, judges have something resembling personas.

We also found an important tension: judges were comparatively stable when rescoring the same answers, but they could still disagree substantially with one another on a particular prompt. 

Prompt enhancements materially improved judge agreement

The improvement from prompt iterations described in previous sections was measurable. After re-anchoring the judges to the source and consolidating the rubric, agreement increased most sharply in the judge pairs that had previously diverged the most. In the calibration sample, Opus-to-GPT accuracy agreement increased from 12.9% to 62.3%, while Gemini-to-GPT agreement increased from 12.5% to 47.9%.

What the new leaderboard shows

The updated MLCR-AA results reinforce several findings from the original release while adding greater detail about why models pass or fail.

Completeness separates the frontier

The strongest models are increasingly good at avoiding clear factual errors. The more difficult capability is covering the full answer: every requested component, relevant piece of evidence, and necessary reasoning step. 26 of the model configurations tested had an accuracy score of over 80%, however only 3 configurations (Opus 5 with high, xhigh, and max reasoning) had a completeness score of over 80%.

At the time of writing, Anthropic’s fifth-generation models occupied the leading positions on the overall MLCR-AA score, with Claude Fable 5 at 64.4% and Claude Opus 5 configurations following at 59.4% and 58.3%. Their largest advantage was in completeness rather than pure factual accuracy. Because the leaderboard is live, exact positions will evolve as Artificial Analysis evaluates additional models.

Reasoning effort can matter but not uniformly

Higher reasoning levels produced substantial improvements for some of the latest frontier models, particularly Anthropic’s fifth-generation configurations, but max reasoning did not necessarily translate to the best outcome.

Cost and performance remain closely connected

The release snapshot showed a broad relationship between model cost and MLCR-AA performance. More expensive frontier configurations generally produced stronger overall results, particularly on questions requiring extensive synthesis.

That relationship is not perfectly monotonic, however. The leaderboard includes models that achieve competitive results at meaningfully lower cost, and the best option depends on whether an application prioritizes maximum completeness, factual precision, latency, or price. For example, GLM 5.3 was the only model, proprietary or not, in the most attractive quadrant in a cost vs. score matrix.

Open-weight models are increasingly competitive

Open-weight models are also closing the gap more quickly than we expected.

Kimi K3 was one of the most notable results in the release snapshot, outperforming several frontier proprietary configurations and even finishing ahead of a tested GPT-5.6 Sol configuration on the overall MLCR-AA score.

Nevertheless, the successful open source models like Kimi K3 and GLM 5.3 often took twice the time per task as their similarly performing proprietary counterparts.

Continuing the open release

The 10 synthetic MLCR cases and the public question tiers remain available through the Wisedocs MLCR dataset on Hugging Face and the reference answers have been updated where necessary.

The open-source MLCR evaluation harness supports configurable model providers, reasoning levels, modalities, and deterministic insertion of irrelevant medical and administrative context. We are updating the public scoring pipeline to reflect the v1.1 judge calibration described here.

The difficult held-out set will remain private and will continue to power the independently run Artificial Analysis MLCR-AA leaderboard.

August 31, 2026

Aryan Dhar

Author

Aryan Dhar is a Machine Learning Engineer at Wisedocs based in Toronto. Prior to Wisedocs, he was worked in a variety of other companies in the AI/ML space such as Cerebras Systems and Equinix. In his spare time, he enjoys hiking, canoeing, running, cooking, and reading.

Soft blue and white abstract blurred gradient background.

Stay ahead of the (AI) curve

How is AI changing the way insurance, legal, and medical professionals work across claims? 
Get analysis and best practices from our team of experts. Sent every other week.