Now: finished (idle, 0s)
in table (v3/v4): ok 0 (unreliable 0)
· failed 0 · skipped 0
· ignored old formats 7000 · queue 7000
this session: ok 420 · failed 39 · skipped 6541
Score = ARC-Easy, 100 questions, zero-shot, answer picked by next-token letter logits (chance = 0.25; one score is only good to about ±10 points at this size). fmt = share of probability the model puts on the answer letters; grey rows with “?” are below 0.02 (the model doesn't answer with a letter, so the score is unreliable). Rows marked v3 use the previous scoring and are replaced as the v5 run reaches them. Raw results: bucket (top 300 shown, table rebuilt every 120s) · recent logs (1000 lines)
| model | params | score | n | fmt | prompt tok/s | task | file | v |
|---|