Keryxeval traces

The eval traces behind Keryx's published 6 of 10

All 9 traces here are control scenarios — alerts injected with no discoverable local cause, where the correct answer is “insufficient evidence.” They are the hardest thing to get right, because the failure mode they test for is an investigator that would rather name something than name nothing.

Keryx got it right 1 time. The other 8 are on this page because publishing failures is the point: a benchmark you can only read when it flatters you measures nothing. Each trace is complete — every tool call, every excerpt, every replayable query the investigator used to reach a conclusion that was wrong.

What is missing from this page

There is no passing S-scenario trace here — no example of Keryx correctly naming a root cause — and that is a gap in the archive, not in the result. Six golden scenarios did pass (S1, S2, S3, S4, S5, S6), which is where the 6 of 10 comes from. Their traces no longer exist: snapshot-writing was added to the scoring tool after that batch ran, and the eval cluster is ephemeral, so teardown destroyed the findings, evidence and trace steps unrecoverably. The scored rows survive in the published artifact; the evidence behind them does not.

Showing you the failures we still have is better than showing you a reconstruction of the successes we lost.

9 traces