Skip to content

Pull requests: evaleval/every_eval_ever

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Retry a Hub server error instead of ending the adapter's run
#292 opened Sep 10, 2026 by borgr Collaborator Loading…
Read the Terminal-Bench 2.0 rows from the source the page reads
#290 opened Sep 5, 2026 by borgr Collaborator Loading…
Address the hal datastore path by the model id it publishes
#289 opened Sep 5, 2026 by borgr Collaborator Loading…
Declare a metric's scale only where something backs the claim
#287 opened Sep 5, 2026 by borgr Collaborator Loading…
5 of 7 tasks
Stop publishing a pre-v0.4 lm-eval standard error as if it were a score
#285 opened Sep 2, 2026 by borgr Collaborator Loading…
Add the LM-Harmony adapter, keeping its two evaluation protocols unjoinable
#284 opened Sep 2, 2026 by borgr Collaborator Loading…
Add aiXamine adapter
#283 opened Aug 30, 2026 by fatihdeniz Loading…
docs: harden datastore PR review workflow
#241 opened Aug 9, 2026 by nelaturuharsha Collaborator Loading…
19 of 20 tasks
feat: add sayf-eval results-record converter
#220 opened Aug 5, 2026 by Aymen311 Loading…
feat: add generalized Kaggle Community Benchmarks adapter
#191 opened Jun 21, 2026 by mrshu Contributor Loading…
[DRAFT] Transparent Compression stale
#133 opened May 8, 2026 by Erotemic Collaborator Loading…
ProTip! Exclude everything labeled bug with -label:bug.