We updated our flagship virology troubleshooting benchmark, the Virology Capabilities Test (VCT), to be more shortcut-resistant. Auditing VCT confirmed that the majority of the questions are scientifically valid and that the benchmark is not saturated; we expect the updated VCT-v2 to saturate later due to an increase in dynamic range.
Benchmarks measure AI capabilities, until they don’t
Since the release of VCT in April 2025, the accuracy of the top-performing model on the benchmark has increased by ~12 percentage points (o3: 42.0%; GPT-5.6 Sol: 53.9%). The steady improvement of models on VCT raises one concern: when does the benchmark stop being informative?
Once frontier models reach the ‘effective ceiling’ of a benchmark’s dynamic range, the benchmark is saturated and is no longer sensitive to further increases in model capability. While the effective ceiling of a benchmark should ideally be at a score of 100%, imperfect benchmarks may saturate earlier. The developers of FrontierMath identified fatal errors that rendered the answer key incorrect for about a third of problems, which lowered the effective ceiling to below 70%. BixBench was found to contain problematic questions with incorrect ground truth and ambiguous or underspecified context. The effective ceiling of VCT is likely below 100% for several reasons:
- Biology is not black and white. Some answer choices may fall into a gray area and cannot be unambiguously marked as correct/incorrect even by experts.
- Some questions do not have a single ground truth answer. A failed biology experiment can be troubleshot in several ways; the success of one approach does not necessarily invalidate another, so the answer given in the key may not be the only correct answer.
- Information provided in the question may not be truly sufficient. VCT questions are derived from real-life scenarios, and the question authors may have omitted key information without themselves realizing its importance.
Figure 1. Benchmarks are restricted by the reliability limit of the answer key. This sets the “effective ceiling” of benchmarks and causes benchmark accuracy to stop scaling linearly with the scientific capability of the model beyond a certain threshold, resulting in lost dynamic range. Note: This schematic is only an illustrative representation.
In addition to the issue of the effective ceiling compromising benchmarks, models can ‘game’ a benchmark by exploiting heuristic shortcuts (e.g., picking the longest answer statement). Models exploited shortcuts in over 30% of WMDP-Bio questions, answering correctly without access to the question text. While the multiple-response format1 of VCT may reduce its vulnerability to such shortcuts, further scrutiny is warranted.
Patching potential issues in VCT
We set out to identify and refine imperfect questions in VCT to increase the benchmark’s effective ceiling and to improve our confidence that VCT accuracy correlates strongly with practical virology troubleshooting capability. Reasoning that problematic questions would be more likely to result in specific answer patterns, we developed a set of deterministic criteria to enrich for questions likely to be invalid, be misleading, or contain unnecessary information. For example, several models converging on the same incorrect answer might indicate a faulty answer key or misleading question; a question that neither models nor experts can solve may be flawed, or merely difficult; and models arriving at the correct answer without access to the question text or image suggests the question does not demand genuine scientific reasoning.
Since we did not know in advance which signals indicated genuinely problematic questions, the flagging criteria favored recall over precision, to capture a fuller picture of how questions can fail. Flagged questions were manually reviewed by an expert virologist with LLM assistance and were either approved, removed, or edited. This process revealed several common failure modes of VCT questions:
- Over-revealing text: question text and/or statements are too detailed, so models could answer without genuine reasoning.
- Scientific errors and ambiguity: questions have objectively incorrect answers, ambiguous distractors, imprecise or misleading wording, or content that lacks clear consensus.
- Unnecessary image: image-containing questions could be answered without the image, indicating that image interpretation capabilities are not assessed.
Beyond this manual refinement, we removed 25 easy (non-discriminating) questions, as they inflate headline accuracy and compress VCT’s dynamic range. As a result, 43 (13.4%) questions were removed and 163 (50.6%) questions were edited to generate a revised benchmark, VCT-v2, with 279 questions.
Figure 2. Generation of the updated 279-question VCT-v2. Of the 322 VCT questions, 43 were removed, including 25 easy, non-discriminating questions; 163 were edited to improve their scientific accuracy, clarity, and/or resistance to heuristic shortcuts; and 116 remained unchanged.
Presenting our updated benchmark — VCT-v2
Model performance differs only slightly between VCT and VCT-v2
Model accuracies2 dropped by 1.0–7.1 percentage points on VCT-v2 relative to VCT (the original benchmark) across the models tested, which reflects the removal of easy questions and the closing of shortcuts. Importantly, the difference between the strongest and weakest models tested was preserved (26.6 → 26.1%), suggesting that VCT-v2’s discriminatory power is unchanged. The human expert baseline accuracy on VCT-v23 remains around ~21%.
Figure 3. Model performances on VCT and VCT-v2. Human expert baseline accuracy on VCT-v2, which is derived from unedited or minimally edited questions, is also plotted alongside that for VCT. Bars represent mean accuracies.
The primary goal of revisiting VCT is to assess and improve its validity. Our auditing efforts meaningfully changed model performance in 42 questions, of which model(s) performed worse on 22 questions and better on 20 questions. Model accuracy remained unchanged on two-thirds of the edited questions, indicating that VCT results remain largely valid. The changes in per-question model performance, in both directions, resulted from the different motivations and outcomes of question edits. Closing shortcuts and removing easy, non-discriminating questions made the benchmark more difficult. Conversely, improving question clarity might have eliminated some incorrect answers that were given due to question ambiguity, thus increasing model performance. Rectifying invalid answer keys means that model answers are now scored differently.
Figure 4. Volcano plot of per-question accuracy changes from VCT to VCT-v2. Each point represents one question. Edited questions are colored by direction of accuracy change and significance at uncorrected α = 0.05 (dashed horizontal line). Unchanged questions (gray) serve as background noise reference. The mean accuracy delta was derived from 9 paired model runs. Statistical significance was calculated using a one-sample t-test, treating the 9 models as replicates.
VCT-v2 is substantially more shortcut-resistant than VCT
The flagging pipeline identified 87/322 (27.0%) VCT questions as shortcut-exploitable, that is, questions that models could answer correctly (and in some cases, better) without the image and/or question text. Fixing these questions by removing unnecessary information represents a step-improvement in benchmark quality: models’ ability to deduce the correct answer independent of genuine knowledge by exploiting shortcuts is no longer rewarded. In other words, VCT-v2 more faithfully measures what it claims to — capability in practical virology troubleshooting.
Figure 5. Shortcut exploitation by models on VCT and VCT-v2. Accuracies of three selected models on VCT and VCT-v2 in either image-omitted (images removed) or question-omitted (question texts and images removed) mode. Dotted lines represent model accuracies in standard run configuration on the respective question subsets. Only image-containing questions were included in image-omitted runs to assess image requirement.
While the changes in VCT-v2 markedly reduced model accuracies in image- and question-omitted runs, a meaningful number of questions slipped through and remained answerable by frontier models in these ablated runs. Three factors may explain this: first, the flagging criteria treat questions with an accuracy below 0.7 across ≥10 epochs as incorrect, so questions that are unreliably answered by models stay under the radar; second, models are getting better at exploiting heuristic shortcuts, so current frontier models may find shortcuts that the best models available at the time of flagging did not; and third, inter-run variance prevents unambiguous flagging of questions that are close to the threshold.
As models continue to improve, we expect further rounds of AI-assisted review to surface new caveats or flaws in VCT-v2 and other existing benchmarks.
Monitoring VCT-v2 for saturation
While VCT is not saturated at the time of this writing, the rapid improvements in model performance (see our AI benchmarks dashboard for model performance over time) mean that we may be only a few model generations away from VCT’s effective ceiling. Assessing future models on both VCT and VCT-v2 will reveal whether our reviewing efforts have increased the effective ceiling of VCT-v2, and whether the benchmark’s accuracy–capability curve (Figure 1) is bent upwards relative to VCT.
What we can establish now is a lower bound: the effective ceiling of VCT-v2 sits at 79% or higher. While frontier models achieve comparable accuracies on VCT-v2, they succeed on different question subsets with developer-family clustering. Aggregating highest per-question accuracy across models gives a theoretical best of 79%, or ~30 percentage points above the top-performing model we assessed (GPT-5.6 Sol: 49.2%). The same calculation on VCT gives a headroom of only 23 percentage points. VCT-v2 thus offers more room to track model improvements, where 18/279 (6.5%) questions are not answerable by any model and the top-performing model cannot reliably answer 52/279 (18.6%) questions (i.e., scores below 50% across ≥10 epochs).
Figure 6. Heatmap showing per-question performances of models and human experts on VCT-v2. Each cell represents the mean question × model accuracy, with darker colors indicating higher accuracy. Mean model accuracy on all questions is displayed to the right of the heatmap. Yellow cells represent data missing due to model refusals or missing expert baseline entries. The theoretical best model is taken to be the top-performing model on a per-question basis.
Lessons learned from revisiting VCT
Our efforts in analyzing, flagging, and editing VCT taught us several things:
- Benchmark maintenance should be an ongoing effort.
Benchmark creation should be treated not as a one-off project but rather as an ongoing effort with a defined re-audit cadence. Model and expert performance data exist only once a benchmark has been built, and they carry per-question information about its quality.
- Benchmarks require shortcut resistance checks.
Shortcut exploitation allows models to achieve inflated accuracy without relying on genuine domain-specific capabilities. Ablation runs, such as image- or question-omitted runs, can help identify exploitable questions to be excluded or fixed. Converting benchmarks into cloze-style (e.g., WMDP-Bio Verified Cloze) and running them in open-ended response format, which our benchmarks are equipped for, could further strengthen shortcut resistance.
- Flagging is cheap; manual review is expensive.
Applying deterministic question flags using run data we already have was straightforward, but having a human expert review every single flagged question was resource-intensive, especially when flagging was designed for recall, not precision. Since AI capability has improved substantially since we began this project, automation with human-in-the-loop could be explored.
- Accuracy is not the only relevant metric.
Biosecurity-relevant benchmark performances are often presented as a single accuracy value. Yet, two models achieving the same accuracy on a benchmark are not necessarily identical, and reporting accuracy alone conceals a lot of information. We are working on applying item-response theory to better understand model performances on our benchmarks.
Our benchmarks have taught us a lot about model capabilities in biosecurity-relevant domains, and they continue to track improvements with new model releases. Revisiting VCT, our flagship virology benchmark, showed us how to build better benchmarks and how to extend the shelf-life of the ones we already have. Keeping evaluations relevant is an ongoing commitment — one we intend to keep making.
To get a question correct, the solver must select all, and only, the correct statements from 4–10 different answer statements, where at least one statement is correct. In this format, the expected accuracy achieved by random guessing is 2.4%, much lower than that for a four-option multiple-choice question (25%).
All model accuracies reported here refer to refusal-corrected accuracy, which is the accuracy on the non-refused subset of questions over ≥ 10 epochs.
The human baseline for VCT was established by having expert virologists with matched areas of expertise answer the questions. To derive a baseline for VCT-v2, existing entries for edited questions were retained, rescored, or removed, depending on the extent of the edit. Entries were rescored for 11 questions, by grading the original expert answer against the updated answer key. This was only done where editing removed answer statements or flipped their polarity (e.g., true to false) without changing their content. No additional baselining was performed, so the VCT-v2 baseline is based on on fewer entries than VCT’s (VCT: 841 entries across 305 questions; VCT-v2: 593 entries across 216 questions; 1–3 entries per question and a total of 36 experts in both cases).