English WSD benchmark

SenseBench Leaderboard

SenseBench measures how well language models disambiguate English words: each model sees a word in its sentence context together with its candidate WordNet senses and must answer with the index of the correct sense. Every row is recomputed from verified, fully auditable run artifacts on the lexEN dataset, and anyone can submit a run by pull request.

Verified Runs
225
Models
69
Top Accuracy
95.66%
Dataset
lexen-v1

Which LLM is best at word sense disambiguation?

As of , the best verified result on lexEN v1 is 95.66% (95% CI 95.08–96.22), from GPT 6 Astra at xhigh reasoning effort under registered prompt p001 — 4,650 of 4,861 polysemous English items. GPT-5.5 (95.60%, McNemar p = 0.88), Gemini 3.8 Flash (95.25%, p = 0.16) and Claude Fable 5 (95.21%, p = 0.13) are statistically indistinguishable from it; GPT 6 Sol, Gemini 3.1 Pro and Claude Opus 5 all fall significantly below. WordNet's most-frequent-sense heuristic scores 61.55% on the same items; among supervised systems ConSeC, trained on the original human labels, reaches 84.88%, while Glite's own LENS, retrained on model-relabelled SemCor, reaches 89.69%. Figures use the default labels — lexEN v1 gold at WordNet fine granularity; coarser sense inventories score substantially higher. Every number is recomputed in CI from the stored raw API responses.

225 verified runs · 69 models · latest run 25 September 2026

Re-scores the entire leaderboard — table, chart, Pareto frontier, ranks, and pairwise tests — against the chosen gold labels and sense granularity. Default is lexEN v1 · WordNet fine-grained. What do these mean? → About coarsening →

Reference Baselines

System Accuracy Dataset Provenance
MFS (WordNet first sense)
Computed at build time
61.55%
±1.39%
lexen-v1 Most frequent sense baseline: WordNet 3.0's first (frequency-ranked) sense for the target lemma and part of speech, computed directly on the dataset items.
BEM
Published predictions
79.65%
±1.13%
lexen-v1 Bi-Encoder Model (Blevins & Zettlemoyer 2020); per-item predictions released by Maru et al. 2022, scored on this dataset's items.
Reproduced predictions
81.42%
±1.08%
lexen-v1 ESCHER (Barba et al. 2021; SemCor training); predictions reproduced by Glite, 79.6 F1 on Raganato ALL (-1.1 of the published 80.7 F1), scored on this dataset's items.
Reproduced predictions
84.88%
±0.99%
lexen-v1 ConSeC (Barba et al. 2021); predictions reproduced by Glite (SemCor + WordNet Gloss+Examples training, 82.9 F1 on Raganato ALL, -0.3 of the published 83.2 F1), scored on this dataset's items.
Published predictions
89.69%
±0.84%
lexen-v1 Glite LENS (ModernBERT bi-encoder); shipped seed-42 predictions trained on SemCor-GPT5.5, the GPT-5.5-relabeled corpus rather than original SemCor (83.7 F1 on Raganato ALL; 3-seed mean 83.6). This row demonstrates the relabel-and-retrain result, so the LENS-ESCHER margin is not an architecture-only comparison; because its training labels share a model family with the lexEN triage, its lexEN score is confirmatory under the paper's Section 6.4 rule.

Classic WSD systems scored from per-item predictions on exactly the same dataset items as the model runs, with the same correctness rule. They appear as dashed lines on the chart.

Compare Prompt
1
OpenAI · Proprietary
★ 95.66%
±0.57%
$9,957 p001
2
OpenAI · Proprietary
95.60%
±0.59%
$10,700 p001
3
Google · Proprietary
★ 95.25%
±0.59%
$923 p001
4
OpenAI · Proprietary
95.25%
±0.60%
$6,077 p001
5
Anthropic · Proprietary
95.21%
±0.58%
$14,555 p001
6
OpenAI · Proprietary
95.19%
±0.60%
$7,718 p001
7
OpenAI · Proprietary
95.15%
±0.62%
$6,233 p004
8
OpenAI · Proprietary
95.15%
±0.60%
$11,301 p004
9
OpenAI · Proprietary
95.10%
±0.62%
$6,289 p003
10
Google · Proprietary
95.04%
±0.61%
$4,107 p001
11
OpenAI · Proprietary
95.00%
±0.63%
$1,417 p001
12
OpenAI · Proprietary
95.00%
±0.63%
$2,153 p001
13
OpenAI · Proprietary
95.00%
±0.61%
$3,516 p003
14
OpenAI · Proprietary
95.00%
±0.62%
$5,040 p001
15
OpenAI · Proprietary
94.98%
±0.63%
$4,562 p003
16
OpenAI · Proprietary
94.94%
±0.64%
$4,259 p001
17
Google · Proprietary
94.92%
±0.61%
$7,227 p001
18
Anthropic · Proprietary
94.75%
±0.62%
$8,660 p001
19
Moonshot · Open weights
94.63%
±0.66%
$4,112 p001
20
Google · Proprietary
94.61%
±0.63%
$3,691 p001
21
Anthropic · Proprietary
94.57%
±0.64%
$7,285 p001
22
Google · Proprietary
94.53%
±0.64%
$7,586 p004
23
OpenAI · Proprietary
94.51%
±0.63%
$3,393 p001
24
Google · Proprietary
94.43%
±0.63%
$4,711 p001
25
Google · Proprietary
94.40%
±0.65%
$3,081 p003
26
OpenAI · Proprietary
94.22%
±0.66%
$2,129 p001
27
OpenAI · Proprietary
94.22%
±0.67%
$5,625 p002
28
Google · Proprietary
94.20%
±0.64%
$5,557 p004
29
OpenAI · Proprietary
94.18%
±0.67%
$1,928 p001
30
Anthropic · Proprietary
94.16%
±0.65%
$5,243 p001
31
Google · Proprietary
94.16%
±0.63%
$5,408 p001
32
OpenAI · Proprietary
94.03%
±0.67%
$3,790 p002
33
OpenAI · Proprietary
93.95%
±0.67%
$1,515 p003
34
xAI · Proprietary
93.85%
±0.67%
$3,307 p001
35
Anthropic · Proprietary
93.85%
±0.67%
$4,834 p001
36
Anthropic · Proprietary
93.81%
±0.66%
$7,655 p004
37
Google · Open weights · fp8 · H100 80GB
★ 93.73%
±0.68%
$300 p003
38
OpenAI · Proprietary
93.68%
±0.67%
$1,783 p001
39
Anthropic · Proprietary
93.68%
±0.66%
$4,850 p001
40
Anthropic · Proprietary
93.62%
±0.68%
$4,994 p001
41
OpenAI · Proprietary
93.50%
±0.69%
$1,718 p001
42
Google · Open weights · fp8 · H100 80GB
93.44%
±0.69%
$358 p004
43
Google · Open weights · fp8 · H100 80GB
93.38%
±0.69%
$362 p001
44
Z.ai · Open weights
93.38%
±0.70%
$3,333 p001
45
Moonshot · Open weights
93.31%
±0.69%
$2,565 p001
46
OpenAI · Proprietary
93.23%
±0.72%
$1,168 p003
47
xAI · Proprietary
93.13%
±0.71%
$2,583 p001
48
Z.ai · Open weights
93.11%
±0.72%
$1,125 p001
49
OpenAI · Proprietary
93.09%
±0.69%
$1,300 p003
50
xAI · Proprietary
93.09%
±0.69%
$1,763 p001
51
OpenAI · Proprietary
★ 93.05%
±0.73%
$111 p001
52
Anthropic · Proprietary
93.01%
±0.73%
$4,894 p001
53
OpenAI · Proprietary
92.88%
±0.73%
$2,260 p003
54
xAI · Proprietary
92.88%
±0.71%
$2,816 p001
55
Alibaba · Open weights
92.82%
±0.74%
$2,152 p002
56
Alibaba · Open weights
92.70%
±0.73%
$794 p001
57
Anthropic · Proprietary
92.70%
±0.74%
$3,933 p001
58
Moonshot · Open weights
92.66%
±0.74%
$3,091 p001
59
Google · Proprietary
92.57%
±0.73%
$847 p002
60
OpenAI · Proprietary
92.53%
±0.71%
$1,167 p003
61
DeepSeek · Open weights
92.41%
±0.75%
$1,984 p001
62
Google · Open weights · fp8 · H100 80GB
★ 92.39%
±0.74%
$34.83 p001
63
Moonshot · Open weights
92.39%
±0.76%
$2,615 p002
64
Google · Open weights · fp8 · H100 80GB
★ 92.29%
±0.75%
$21.11 p003
65
Anthropic · Proprietary
92.22%
±0.74%
$1,934 p001
66
Z.ai · Open weights
92.22%
±0.75%
$3,246 p002
67
Anthropic · Proprietary
92.22%
±0.74%
$3,338 p003
68
OpenAI · Proprietary
92.20%
±0.78%
$254 p001
69
Anthropic · Proprietary
91.96%
±0.75%
$1,935 p001
70
Anthropic · Proprietary
91.83%
±0.75%
$2,025 p001
71
Anthropic · Proprietary
91.79%
±0.74%
$2,121 p002
72
Anthropic · Proprietary
91.75%
±0.78%
$1,939 p001
73
OpenAI · Proprietary
91.61%
±0.79%
$224 p001
74
DeepSeek · Open weights
91.50%
±0.77%
$1,338 p002
75
Alibaba · Open weights
91.46%
±0.78%
$659 p002
76
OpenAI · Proprietary
91.26%
±0.83%
$193 p003
77
OpenAI · Proprietary
91.22%
±0.82%
$690 p003
78
Google · Open weights · fp8 · H100 80GB
★ 91.20%
±0.79%
$11.29 p002
79
OpenAI · Proprietary
91.20%
±0.80%
$1,331 p001
80
OpenAI · Proprietary
91.11%
±0.84%
$156 p003
81
OpenAI · Proprietary
91.05%
±0.81%
$904 p003
82
OpenAI · Proprietary
90.99%
±0.82%
$386 p003
83
DeepSeek · Open weights
90.93%
±0.83%
$132 p001
84
OpenAI · Proprietary
90.89%
±0.81%
$512 p003
85
OpenAI · Proprietary
90.76%
±0.81%
$73.76 p001
86
OpenAI · Proprietary
90.72%
±0.81%
$803 p001
87
Anthropic · Proprietary
90.72%
±0.84%
$2,141 p001
88
MiniMax · Open weights
90.62%
±0.81%
$545 p001
89
OpenAI · Proprietary
90.54%
±0.81%
$480 p001
90
OpenAI · Proprietary
90.48%
±0.83%
$217 p003
91
Anthropic · Proprietary
90.48%
±0.81%
$3,348 p001
92
Google · Proprietary
90.41%
±0.83%
$171 p001
93
OpenAI · Proprietary
90.25%
±0.82%
$176 p001
94
Anthropic · Proprietary
90.17%
±0.83%
$1,296 p003
95
OpenAI · Proprietary
90.10%
±0.84%
$685 p001
96
OpenAI · Proprietary
90.04%
±0.83%
$363 p002
97
Anthropic · Proprietary
90.04%
±0.83%
$1,296 p003
98
OpenAI · Proprietary
89.94%
±0.84%
$306 p001
99
OpenAI · Proprietary
89.80%
±0.87%
$127 p003
100
DeepSeek · Open weights
89.69%
±0.85%
$88.64 p002
101
Anthropic · Proprietary
89.67%
±0.86%
$1,325 p003
102
Google · Proprietary
89.61%
±0.82%
$674 p002
103
Alibaba · Open weights · fp8 · H200 141GB
89.51%
±0.85%
$40.91 p001
104
OpenAI · Proprietary
89.49%
±0.88%
$374 p002
105
Google · Open weights · fp8 · A100 80GB
89.43%
±0.88%
$13.04 p001
106
Google · Proprietary
89.41%
±0.86%
$2,529 p001
107
Alibaba · Open weights · fp8 · H100 80GB
89.36%
±0.86%
$25.36 p001
108
Google · Open weights · fp8 · H200 141GB
89.32%
±0.87%
$15.59 p001
109
Alibaba · Open weights · gptq-int4 · B300 288GB
89.24%
±0.87%
$158 p001
110
Google · Open weights · fp8 · H100 80GB
★ 89.20%
±0.87%
$9.13 p001
111
OpenAI · Proprietary
89.14%
±0.88%
$866 p003
112
Google · Open weights · fp8 · H100 80GB
★ 89.12%
±0.88%
$4.75 p003
113
Google · Proprietary
89.08%
±0.88%
$59.42 p002
114
Alibaba · Open weights · bf16 · A100 80GB
89.06%
±0.86%
$48.22 p001
115
Alibaba · Open weights · bf16 · H200 141GB
88.97%
±0.87%
$58.36 p001
116
Alibaba · Open weights · bf16 · H100 80GB
88.89%
±0.87%
$36.33 p001
117
Anthropic · Proprietary
88.89%
±0.88%
$809 p002
118
OpenAI · Proprietary
88.83%
±0.92%
$108 p003
119
Z.ai · Open weights · awq-int4 · B300 288GB
88.71%
±0.89%
$222 p001
120
OpenAI · Proprietary
88.69%
±0.87%
$1,281 p001
121
OpenAI · Proprietary
88.34%
±0.93%
$130 p001
122
OpenAI · Proprietary
87.86%
±0.92%
$465 p002
123
OpenAI · Proprietary
87.80%
±0.92%
$90.38 p003
124
OpenAI · Proprietary
87.72%
±0.92%
$118 p001
125
Anthropic · Proprietary
87.60%
±0.92%
$1,997 p002
126
Meta · Open weights · awq-int4 · B300 288GB
87.55%
±0.94%
$1,062 p001
127
Google · Open weights · fp8 · H100 80GB
★ 87.25%
±0.95%
$2.59 p002
128
Google · Open weights · fp8 · H200 141GB
87.25%
±0.96%
$7.11 p002
129
Alibaba · Open weights · fp8 · H100 80GB
87.20%
±0.96%
$9.40 p002
130
Alibaba · Open weights · fp8 · H200 141GB
87.18%
±0.96%
$15.62 p002
131
Z.ai · Open weights · awq-int4 · B300 288GB
87.12%
±0.95%
$89.31 p002
132
Google · Open weights · fp8 · A100 80GB
87.10%
±0.99%
$4.25 p002
133
Alibaba · Open weights · fp8 · H200 141GB
87.06%
±0.95%
$42.45 p001
134
Alibaba · Open weights · fp8 · B300 288GB
87.02%
±0.96%
$122 p001
135
Alibaba · Open weights · gptq-int4 · B300 288GB
86.69%
±0.95%
$61.21 p002
136
OpenAI · Proprietary
86.63%
±0.92%
$145 p003
137
Alibaba · Open weights · awq-int4 · H200 141GB
86.59%
±0.99%
$88.23 p001
138
Alibaba · Open weights · bf16 · H100 80GB
86.55%
±0.98%
$13.43 p002
139
Alibaba · Open weights · bf16 · H200 141GB
86.55%
±0.98%
$21.98 p002
140
OpenAI · Proprietary
86.48%
±0.95%
$209 p003
141
Alibaba · Open weights · bf16 · A100 80GB
86.46%
±0.97%
$17.73 p002
142
OpenAI · Proprietary
86.26%
±0.96%
$221 p001
143
Google · Open weights · fp8 · H100 80GB
85.58%
±1.01%
$32.90 p003
144
Alibaba · Open weights · fp8 · H100 80GB
85.56%
±1.00%
$8.62 p001
145
Alibaba · Open weights · fp8 · H200 141GB
85.50%
±1.00%
$14.85 p001
146
OpenAI · Proprietary
85.50%
±1.00%
$209 p002
147
Alibaba · Open weights · fp8 · H200 141GB
85.48%
±0.99%
$18.80 p002
148
Alibaba · Open weights · fp8 · B300 288GB
85.29%
±0.97%
$35.77 p002
149
Alibaba · Open weights · fp8 · B300 288GB
85.27%
±1.00%
$56.87 p001
150
Cohere · Open weights · fp8 · H200 141GB
85.23%
±1.01%
$131 p001
151
Google · Open weights · fp8 · H100 80GB
85.19%
±1.00%
$54.16 p001
152
OpenAI · Proprietary
85.00%
±1.01%
$185 p001
153
Alibaba · Open weights · awq-int4 · A100 80GB
84.84%
±1.00%
$53.82 p002
154
Alibaba · Open weights · awq-int4 · H100 80GB
84.82%
±1.02%
$32.60 p002
155
Alibaba · Open weights · awq-int4 · H200 141GB
84.82%
±1.02%
$52.61 p002
156
Alibaba · Open weights · awq-int4 · H200 141GB
84.74%
±1.01%
$34.60 p002
157
Alibaba · Open weights · bf16 · H200 141GB
84.45%
±1.02%
$17.33 p001
158
OpenAI · Proprietary
84.32%
±1.02%
$982 p001
159
Alibaba · Open weights · awq-int4 · H200 141GB
84.08%
±1.03%
$89.36 p001
160
NVIDIA · Open weights · fp8 · H200 141GB
83.93%
±1.02%
$22.03 p003
161
OpenAI · Proprietary
83.93%
±0.99%
$105 p002
162
OpenAI · Proprietary
83.77%
±1.04%
$127 p003
163
OpenAI · Proprietary
83.69%
±1.07%
$256 p001
164
OpenAI · Proprietary
83.56%
±1.02%
$63.56 p003
165
OpenAI · Proprietary
83.56%
±1.06%
$400 p001
166
Google · Open weights · fp8 · H200 141GB
83.44%
±1.02%
$12.38 p001
167
Alibaba · Open weights · awq-int4 · A100 80GB
83.44%
±1.04%
$137 p001
168
Alibaba · Open weights · awq-int4 · H200 141GB
83.42%
±1.03%
$137 p001
169
Alibaba · Open weights · awq-int4 · H100 80GB
83.40%
±1.03%
$83.31 p001
170
Google · Open weights · fp8 · H100 80GB
83.30%
±1.02%
$27.65 p002
171
OpenAI · Proprietary
82.90%
±1.05%
$35.26 p002
172
Alibaba · Open weights · fp8 · H100 80GB
82.72%
±1.04%
$2.92 p002
173
Alibaba · Open weights · fp8 · H200 141GB
82.70%
±1.03%
$8.25 p002
174
OpenAI · Proprietary
82.70%
±1.07%
$96.14 p001
175
NVIDIA · Open weights · fp8 · H200 141GB
82.35%
±1.06%
$15.97 p002
176
Alibaba · Open weights · fp8 · B300 288GB
82.27%
±1.08%
$23.12 p002
177
Alibaba · Open weights · bf16 · H200 141GB
82.10%
±1.09%
$8.91 p002
178
NVIDIA · Open weights · fp8 · H200 141GB
81.86%
±1.12%
$41.77 p001
179
Meta · Open weights · awq-int4 · B300 288GB
81.53%
±1.09%
$81.28 p002
180
Google · Open weights · fp8 · H100 80GB
81.44%
±1.10%
$3.82 p003
181
Alibaba · Open weights · awq-int4 · H200 141GB
81.03%
±1.10%
$35.20 p002
182
Alibaba · Open weights · bf16 · A100 80GB
80.91%
±1.13%
$14.99 p001
183
Alibaba · Open weights · bf16 · H200 141GB
80.89%
±1.15%
$19.26 p001
184
Alibaba · Open weights · bf16 · H100 80GB
80.79%
±1.13%
$11.16 p001
185
OpenAI · Proprietary
80.54%
±1.14%
$93.99 p002
186
Z.ai · Open weights · bf16 · A100 80GB
80.46%
±1.12%
$4.97 p002
187
Z.ai · Open weights · bf16 · A100 80GB
79.78%
±1.11%
$13.51 p001
188
Mistral AI · Open weights · fp8 · H200 141GB
79.39%
±1.11%
$13.80 p002
189
Mistral AI · Open weights · fp8 · H100 80GB
79.28%
±1.13%
$8.47 p002
190
Mistral AI · Open weights · fp8 · A100 80GB
79.00%
±1.13%
$20.41 p002
191
Google · Open weights · fp8 · H200 141GB
78.89%
±1.15%
$5.38 p002
192
Google · Open weights · fp8 · H100 80GB
78.89%
±1.10%
$28.04 p003
193
Google · Open weights · fp8 · H100 80GB
★ 78.87%
±1.15%
$2.21 p002
194
Tencent · Open weights · gptq-int4 · A100 80GB
78.73%
±1.16%
$32.98 p001
195
Z.ai · Open weights · bf16 · A100 80GB
77.60%
±1.17%
$12.16 p001
196
Mistral AI · Open weights · fp8 · A100 80GB
77.17%
±1.19%
$48.92 p001
197
Google · Open weights · fp8 · H100 80GB
77.12%
±1.17%
$21.64 p002
198
OpenAI · Proprietary
76.98%
±1.17%
$125 p001
199
Mistral AI · Open weights · fp8 · H100 80GB
76.75%
±1.22%
$18.36 p001
200
Mistral AI · Open weights · fp8 · H200 141GB
76.61%
±1.25%
$28.17 p001
201
Alibaba · Open weights · bf16 · H100 80GB
76.34%
±1.17%
$4.29 p002
202
Alibaba · Open weights · bf16 · H200 141GB
76.28%
±1.19%
$7.56 p002
203
Z.ai · Open weights · bf16 · A100 80GB
75.70%
±1.19%
$4.42 p002
204
AllenAI · Open weights · fp8 · H100 80GB
73.98%
±1.21%
$24.33 p001
205
AllenAI · Open weights · fp8 · H100 80GB
73.32%
±1.25%
$9.56 p002
206
OpenAI · Proprietary
73.28%
±1.23%
$23.50 p002
207
IBM · Open weights · fp8 · H200 141GB
70.17%
±1.25%
$30.30 p001
208
Google · Open weights · fp8 · H200 141GB
70.13%
±1.27%
$10.45 p001
209
OpenAI · Proprietary
70.07%
±1.18%
$64.03 p001
210
IBM · Open weights · fp8 · H100 80GB
69.92%
±1.27%
$17.50 p001
211
NVIDIA · Open weights · fp8 · H100 80GB
69.22%
±1.36%
$2.89 p002
212
IBM · Open weights · bf16 · A100 80GB
69.02%
±1.26%
$28.04 p001
213
NVIDIA · Open weights · fp8 · H100 80GB
65.34%
±1.33%
$4.70 p003
214
Google · Open weights · fp8 · H100 80GB
65.07%
±1.33%
$2.50 p003
215
Microsoft · Open weights · bf16 · A100 80GB
63.98%
±1.29%
$5.78 p001
216
NVIDIA · Open weights · fp8 · H100 80GB
63.55%
±1.38%
$7.07 p001
217
Google · Open weights · fp8 · H100 80GB
★ 61.65%
±1.37%
$1.72 p002
218
Google · Open weights · fp8 · H200 141GB
61.63%
±1.36%
$5.89 p002
219
OpenAI · Proprietary
59.06%
±1.36%
$37.51 p001
220
Meta · Open weights · bf16 · A100 80GB
58.82%
±1.41%
$11.30 p001
221
Meta · Open weights · fp8 · H200 141GB
58.05%
±1.42%
$12.18 p001
222
Meta · Open weights · fp8 · H100 80GB
57.99%
±1.41%
$6.09 p001
223
Meta · Open weights · fp8 · H100 80GB
44.31%
±1.39%
$2.49 p002
224
Meta · Open weights · fp8 · H200 141GB
44.21%
±1.38%
$5.45 p002
225
Meta · Open weights · bf16 · A100 80GB
39.85%
±1.37%
$4.27 p002

Compare Selected Runs

Select up to six runs from the table.

Method Notes

Public rows are rebuilt from verified run artifacts in the repository. Accuracy is recomputed from predictions, and official runs are checked against the registered dataset hash and prompt ID.

The ± value under each accuracy is half the width of a fixed-seed bootstrap 95% confidence interval. The small range under each rank lists the positions a run could plausibly occupy among the visible rows given overlapping intervals.

The Score against and Sense granularity controls re-score every run and baseline against one of nine scoring schemes: the lexEN v1, Maru 2022 (ALLamended), or original Raganato 2017 gold labels, each at WordNet 3.0 fine-grained sense or at one of two coarse concept levels — Glite or CSI. A prediction is coarse-correct when its WordNet sense maps to the same coarse concept as a gold sense (so coarse accuracy is always at least the fine-grained accuracy). The default and official score is lexEN v1 · WordNet fine-grained. See label schemes and coarsening for what each option means.

The Pareto frontier uses the visible rows and the selected chart metric (cost per million items or machine-hours per 1M items): higher accuracy is better, and a lower metric value is better. Starred rows are on the frontier.

Machine-hours per 1M items is recorded for self-hosted runs only. It measures the per-item evaluation loop on the recorded machine, excluding model and asset loading; cloud API runs are excluded from that chart. Because it measures time rather than money, it is the one hardware metric no pricing assumption touches: multiply it by whatever you would pay per hour to cost a model on your own terms.

Cost per million items prices self-hosted runs at a fixed reference rate for their GPU class, not at the rate each machine happened to be rented at, so a lucky or unlucky spot price neither flatters nor penalizes a model. Each run still records the rate it actually paid, and its run page shows both. As of 2026-07-14 the reference rates are $1.09/h for A100 80GB, $2.26/h for H100 80GB, $3.66/h for H200 141GB, $6.72/h for B300 288GB. A cost marked * is the exception: that run is on a GPU class we have no reference rate for, so it is priced at the rate its machine was rented at and is not on the same basis as the rows around it. Such a run can still reach the Pareto frontier, and its star is marked ★* to show the frontier position rests on that non-comparable cost. See methodology for how the rates are set.

Reference baselines are supervised WSD systems and MFS scored from per-item predictions on the same dataset items; dashed chart lines show them for context.

Comparing selected runs on the same dataset uses McNemar's test on paired per-item correctness, which detects differences that overlapping confidence intervals can miss.