[Leaderboard Update] Permute EQ - Claude Opus 5 - 90.58% Pass@1 - #95
[Leaderboard Update] Permute EQ - Claude Opus 5 - 90.58% Pass@1#95ericmillsio wants to merge 2 commits into
Conversation
|
Hi @ericmillsio! Thanks for opening the PR. We can accept the GITHUB_REPOS Q2 rerun, but not the crmarenapro Q3, Q7, and Q9 ones, because replacing only the trials that failed removes the very failures Pass@1 is meant to count. This puts you at 0.8922 for now. We can put that on the leaderboard if you are okay with it. Alternatively, we are happy to re-audit if you rerun all 5 trials of those three queries under the same frozen setup and send whatever comes out, including the failures. |
|
Okay thank you clarification. Please hold off merging for now, I will be attempting reruns. Thanks! |
|
We’ve completed five reruns each for CRMArenaPro Q3, Q7, and Q9, with all 15 passing the current validators. We also reran all five DEPS_DEV_V1 Q1 trials following the tie-validator update in #86. Four pass the current extraction and validation rules; the fifth needs your review. Each query’s updated setup was frozen across all five trials, and all 20 results are included. No completed trials were omitted. Updated
DEPS_DEV_V1 Q1 run 4 begins with this answer: The result table repeats those same five correct pairs. A later “Ties (full disclosure)” section mentions Including the BookReview extraction correction you previously confirmed, the projected scores are:
|
[Leaderboard Update] Permute EQ - Claude Opus 5 - 90.58% Pass@1
This PR updates the Permute EQ submission from #88 as a followup from emails with the benchmark team.
Prompt assignment
No query mixes prompt versions. CRM Q3, Q7, and Q9 all use the original Prompt 1.
Replacements
BookReview remains unchanged (the benchmark team confirmed its extraction error separately). The other 255 answer and trace pairs are unchanged from #88.
Revised result: 252/270 raw, 0.9058 stratified Pass@1.
Submission file: leaderboard_submissions/permute_eq.json
Traces: permute_eq_traces.zip