|
|
|
|
|
by mofeien
8 days ago
|
|
From the Fable 5 System card: > Results > On the VCT multimodal virology evaluation, Mythos 5 scored 0.56, well above the expert
baseline of 0.221 and nearly matching that of Mythos Preview (0.57). This represents an
improvement over both Opus 4.7 (0.50) and Opus 4.8 (0.47). > On the DNA synthesis screening evasion evaluation, Mythos 5’s performance was mixed
across screening criteria. Mythos 5 designed viable plasmids for 2 of 10 target pathogens on at least one screening method, not meeting the low-concern threshold (all 10 pathogens). > [...] we view the results of this
evaluation as indicating that the evaluated models are capable of designing viable plasmids that evade certain screening criteria, though their reliable success at this task is not guaranteed. Do you believe that this is fake, "AI company propaganda"? Or that the models are not going to improve further within months? Or that these results are not concerning? |
|
> VCT consists of 322 multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories.
So it's just question answering? Do you think that scoring well on this test is equivalent to synthesizing a virus?
[1] https://securebio.org/virologytest/