English WSD benchmark

SenseBench Leaderboard

SenseBench measures how well language models disambiguate English words: each model sees a word in its sentence context together with its candidate WordNet senses and must answer with the index of the correct sense. Every row is recomputed from verified, fully auditable run artifacts on the lexEN dataset, and anyone can submit a run by pull request.

Verified Runs
218
Models
65
Top Accuracy
95.60%
Dataset
lexen-v1

Which LLM is best at word sense disambiguation?

As of , the best verified result on lexEN v1 is 95.60% (95% CI 95.00–96.17), from GPT-5.5 at xhigh reasoning effort under registered prompt p001 — 4,647 of 4,861 polysemous English items. Only Claude Fable 5 is statistically indistinguishable from it (95.21%, McNemar p = 0.20); Gemini 3.1 Pro, GPT-5.6 Sol and Claude Opus 5 all fall significantly below. WordNet's most-frequent-sense heuristic scores 61.55% on the same items; among supervised systems ConSeC, trained on the original human labels, reaches 84.88%, while Glite's own LENS, retrained on model-relabelled SemCor, reaches 89.69%. Figures use the default labels — lexEN v1 gold at WordNet fine granularity; coarser sense inventories score substantially higher. Every number is recomputed in CI from the stored raw API responses.

218 verified runs · 65 models · latest run 31 July 2026

Re-scores the entire leaderboard — table, chart, Pareto frontier, ranks, and pairwise tests — against the chosen gold labels and sense granularity. Default is lexEN v1 · WordNet fine-grained. What do these mean? → About coarsening →

Reference Baselines

System Accuracy Dataset Provenance
MFS (WordNet first sense)
Computed at build time
61.55%
±1.39%
lexen-v1 Most frequent sense baseline: WordNet 3.0's first (frequency-ranked) sense for the target lemma and part of speech, computed directly on the dataset items.
BEM
Published predictions
79.65%
±1.13%
lexen-v1 Bi-Encoder Model (Blevins & Zettlemoyer 2020); per-item predictions released by Maru et al. 2022, scored on this dataset's items.
Reproduced predictions
81.42%
±1.08%
lexen-v1 ESCHER (Barba et al. 2021; SemCor training); predictions reproduced by Glite, 79.6 F1 on Raganato ALL (-1.1 of the published 80.7 F1), scored on this dataset's items.
Reproduced predictions
84.88%
±0.99%
lexen-v1 ConSeC (Barba et al. 2021); predictions reproduced by Glite (SemCor + WordNet Gloss+Examples training, 82.9 F1 on Raganato ALL, -0.3 of the published 83.2 F1), scored on this dataset's items.
Published predictions
89.69%
±0.84%
lexen-v1 Glite LENS (ModernBERT bi-encoder); shipped seed-42 predictions trained on SemCor-GPT5.5, the GPT-5.5-relabeled corpus rather than original SemCor (83.7 F1 on Raganato ALL; 3-seed mean 83.6). This row demonstrates the relabel-and-retrain result, so the LENS-ESCHER margin is not an architecture-only comparison; because its training labels share a model family with the lexEN triage, its lexEN score is confirmatory under the paper's Section 6.4 rule.

Classic WSD systems scored from per-item predictions on exactly the same dataset items as the model runs, with the same correctness rule. They appear as dashed lines on the chart.

Compare Prompt
1
OpenAI · Proprietary
95.60%
±0.59%
$10,700 p001
2
OpenAI · Proprietary
95.25%
±0.60%
$6,077 p001
3
Anthropic · Proprietary
95.21%
±0.58%
$14,555 p001
4
OpenAI · Proprietary
95.19%
±0.60%
$7,718 p001
5
OpenAI · Proprietary
95.15%
±0.62%
$6,233 p004
6
OpenAI · Proprietary
95.15%
±0.60%
$11,301 p004
7
OpenAI · Proprietary
95.10%
±0.62%
$6,289 p003
8
OpenAI · Proprietary
95.00%
±0.61%
$3,516 p003
9
OpenAI · Proprietary
95.00%
±0.62%
$5,040 p001
10
OpenAI · Proprietary
94.98%
±0.63%
$4,562 p003
11
OpenAI · Proprietary
94.94%
±0.64%
$4,259 p001
12
Google · Proprietary
94.92%
±0.61%
$7,227 p001
13
Anthropic · Proprietary
94.75%
±0.62%
$8,660 p001
14
Moonshot · Open weights
94.63%
±0.66%
$4,112 p001
15
Google · Proprietary
94.61%
±0.63%
$3,691 p001
16
Anthropic · Proprietary
94.57%
±0.64%
$7,285 p001
17
Google · Proprietary
94.53%
±0.64%
$7,586 p004
18
OpenAI · Proprietary
94.51%
±0.63%
$3,393 p001
19
Google · Proprietary
94.43%
±0.63%
$4,711 p001
20
Google · Proprietary
94.40%
±0.65%
$3,081 p003
21
OpenAI · Proprietary
94.22%
±0.66%
$2,129 p001
22
OpenAI · Proprietary
94.22%
±0.67%
$5,625 p002
23
Google · Proprietary
94.20%
±0.64%
$5,557 p004
24
OpenAI · Proprietary
94.18%
±0.67%
$1,928 p001
25
Anthropic · Proprietary
94.16%
±0.65%
$5,243 p001
26
Google · Proprietary
94.16%
±0.63%
$5,408 p001
27
OpenAI · Proprietary
94.03%
±0.67%
$3,790 p002
28
OpenAI · Proprietary
93.95%
±0.67%
$1,515 p003
29
xAI · Proprietary
93.85%
±0.67%
$3,307 p001
30
Anthropic · Proprietary
93.85%
±0.67%
$4,834 p001
31
Anthropic · Proprietary
93.81%
±0.66%
$7,655 p004
32
Google · Open weights · fp8 · H100 80GB
93.73%
±0.68%
$300 p003
33
OpenAI · Proprietary
93.68%
±0.67%
$1,783 p001
34
Anthropic · Proprietary
93.68%
±0.66%
$4,850 p001
35
Anthropic · Proprietary
93.62%
±0.68%
$4,994 p001
36
OpenAI · Proprietary
93.50%
±0.69%
$1,718 p001
37
Google · Open weights · fp8 · H100 80GB
93.44%
±0.69%
$358 p004
38
Google · Open weights · fp8 · H100 80GB
93.38%
±0.69%
$362 p001
39
Z.ai · Open weights
93.38%
±0.70%
$3,333 p001
40
Moonshot · Open weights
93.31%
±0.69%
$2,565 p001
41
OpenAI · Proprietary
93.23%
±0.72%
$1,168 p003
42
xAI · Proprietary
93.13%
±0.71%
$2,583 p001
43
Z.ai · Open weights
93.11%
±0.72%
$1,125 p001
44
OpenAI · Proprietary
93.09%
±0.69%
$1,300 p003
45
xAI · Proprietary
93.09%
±0.69%
$1,763 p001
46
Anthropic · Proprietary
93.01%
±0.73%
$4,894 p001
47
OpenAI · Proprietary
92.88%
±0.73%
$2,260 p003
48
xAI · Proprietary
92.88%
±0.71%
$2,816 p001
49
Alibaba · Open weights
92.82%
±0.74%
$2,152 p002
50
Alibaba · Open weights
92.70%
±0.73%
$794 p001
51
Anthropic · Proprietary
92.70%
±0.74%
$3,933 p001
52
Moonshot · Open weights
92.66%
±0.74%
$3,091 p001
53
Google · Proprietary
92.57%
±0.73%
$847 p002
54
OpenAI · Proprietary
92.53%
±0.71%
$1,167 p003
55
DeepSeek · Open weights
92.41%
±0.75%
$1,984 p001
56
Google · Open weights · fp8 · H100 80GB
92.39%
±0.74%
$34.83 p001
57
Moonshot · Open weights
92.39%
±0.76%
$2,615 p002
58
Google · Open weights · fp8 · H100 80GB
92.29%
±0.75%
$21.11 p003
59
Anthropic · Proprietary
92.22%
±0.74%
$1,934 p001
60
Z.ai · Open weights
92.22%
±0.75%
$3,246 p002
61
Anthropic · Proprietary
92.22%
±0.74%
$3,338 p003
62
OpenAI · Proprietary
92.20%
±0.78%
$254 p001
63
Anthropic · Proprietary
91.96%
±0.75%
$1,935 p001
64
Anthropic · Proprietary
91.83%
±0.75%
$2,025 p001
65
Anthropic · Proprietary
91.79%
±0.74%
$2,121 p002
66
Anthropic · Proprietary
91.75%
±0.78%
$1,939 p001
67
OpenAI · Proprietary
91.61%
±0.79%
$224 p001
68
DeepSeek · Open weights
91.50%
±0.77%
$1,338 p002
69
Alibaba · Open weights
91.46%
±0.78%
$659 p002
70
OpenAI · Proprietary
91.26%
±0.83%
$193 p003
71
OpenAI · Proprietary
91.22%
±0.82%
$690 p003
72
Google · Open weights · fp8 · H100 80GB
91.20%
±0.79%
$11.29 p002
73
OpenAI · Proprietary
91.20%
±0.80%
$1,331 p001
74
OpenAI · Proprietary
91.11%
±0.84%
$156 p003
75
OpenAI · Proprietary
91.05%
±0.81%
$904 p003
76
OpenAI · Proprietary
90.99%
±0.82%
$386 p003
77
DeepSeek · Open weights
90.93%
±0.83%
$132 p001
78
OpenAI · Proprietary
90.89%
±0.81%
$512 p003
79
OpenAI · Proprietary
90.72%
±0.81%
$803 p001
80
Anthropic · Proprietary
90.72%
±0.84%
$2,141 p001
81
MiniMax · Open weights
90.62%
±0.81%
$545 p001
82
OpenAI · Proprietary
90.54%
±0.81%
$480 p001
83
OpenAI · Proprietary
90.48%
±0.83%
$217 p003
84
Anthropic · Proprietary
90.48%
±0.81%
$3,348 p001
85
Google · Proprietary
90.41%
±0.83%
$171 p001
86
OpenAI · Proprietary
90.25%
±0.82%
$176 p001
87
Anthropic · Proprietary
90.17%
±0.83%
$1,296 p003
88
OpenAI · Proprietary
90.10%
±0.84%
$685 p001
89
OpenAI · Proprietary
90.04%
±0.83%
$363 p002
90
Anthropic · Proprietary
90.04%
±0.83%
$1,296 p003
91
OpenAI · Proprietary
89.94%
±0.84%
$306 p001
92
OpenAI · Proprietary
89.80%
±0.87%
$127 p003
93
DeepSeek · Open weights
89.69%
±0.85%
$88.64 p002
94
Anthropic · Proprietary
89.67%
±0.86%
$1,325 p003
95
Google · Proprietary
89.61%
±0.82%
$674 p002
96
Alibaba · Open weights · fp8 · H200 141GB
89.51%
±0.85%
$40.91 p001
97
OpenAI · Proprietary
89.49%
±0.88%
$374 p002
98
Google · Open weights · fp8 · A100 80GB
89.43%
±0.88%
$13.04 p001
99
Google · Proprietary
89.41%
±0.86%
$2,529 p001
100
Alibaba · Open weights · fp8 · H100 80GB
89.36%
±0.86%
$25.36 p001
101
Google · Open weights · fp8 · H200 141GB
89.32%
±0.87%
$15.59 p001
102
Alibaba · Open weights · gptq-int4 · B300 288GB
89.24%
±0.87%
$158 p001
103
Google · Open weights · fp8 · H100 80GB
89.20%
±0.87%
$9.13 p001
104
OpenAI · Proprietary
89.14%
±0.88%
$866 p003
105
Google · Open weights · fp8 · H100 80GB
89.12%
±0.88%
$4.75 p003
106
Google · Proprietary
89.08%
±0.88%
$59.42 p002
107
Alibaba · Open weights · bf16 · A100 80GB
89.06%
±0.86%
$48.22 p001
108
Alibaba · Open weights · bf16 · H200 141GB
88.97%
±0.87%
$58.36 p001
109
Alibaba · Open weights · bf16 · H100 80GB
88.89%
±0.87%
$36.33 p001
110
Anthropic · Proprietary
88.89%
±0.88%
$809 p002
111
OpenAI · Proprietary
88.83%
±0.92%
$108 p003
112
Z.ai · Open weights · awq-int4 · B300 288GB
88.71%
±0.89%
$222 p001
113
OpenAI · Proprietary
88.69%
±0.87%
$1,281 p001
114
OpenAI · Proprietary
88.34%
±0.93%
$130 p001
115
OpenAI · Proprietary
87.86%
±0.92%
$465 p002
116
OpenAI · Proprietary
87.80%
±0.92%
$90.38 p003
117
OpenAI · Proprietary
87.72%
±0.92%
$118 p001
118
Anthropic · Proprietary
87.60%
±0.92%
$1,997 p002
119
Meta · Open weights · awq-int4 · B300 288GB
87.55%
±0.94%
$1,062 p001
120
Google · Open weights · fp8 · H100 80GB
87.25%
±0.95%
$2.59 p002
121
Google · Open weights · fp8 · H200 141GB
87.25%
±0.96%
$7.11 p002
122
Alibaba · Open weights · fp8 · H100 80GB
87.20%
±0.96%
$9.40 p002
123
Alibaba · Open weights · fp8 · H200 141GB
87.18%
±0.96%
$15.62 p002
124
Z.ai · Open weights · awq-int4 · B300 288GB
87.12%
±0.95%
$89.31 p002
125
Google · Open weights · fp8 · A100 80GB
87.10%
±0.99%
$4.25 p002
126
Alibaba · Open weights · fp8 · H200 141GB
87.06%
±0.95%
$42.45 p001
127
Alibaba · Open weights · fp8 · B300 288GB
87.02%
±0.96%
$122 p001
128
Alibaba · Open weights · gptq-int4 · B300 288GB
86.69%
±0.95%
$61.21 p002
129
OpenAI · Proprietary
86.63%
±0.92%
$145 p003
130
Alibaba · Open weights · awq-int4 · H200 141GB
86.59%
±0.99%
$88.23 p001
131
Alibaba · Open weights · bf16 · H100 80GB
86.55%
±0.98%
$13.43 p002
132
Alibaba · Open weights · bf16 · H200 141GB
86.55%
±0.98%
$21.98 p002
133
OpenAI · Proprietary
86.48%
±0.95%
$209 p003
134
Alibaba · Open weights · bf16 · A100 80GB
86.46%
±0.97%
$17.73 p002
135
OpenAI · Proprietary
86.26%
±0.96%
$221 p001
136
Google · Open weights · fp8 · H100 80GB
85.58%
±1.01%
$32.90 p003
137
Alibaba · Open weights · fp8 · H100 80GB
85.56%
±1.00%
$8.62 p001
138
Alibaba · Open weights · fp8 · H200 141GB
85.50%
±1.00%
$14.85 p001
139
OpenAI · Proprietary
85.50%
±1.00%
$209 p002
140
Alibaba · Open weights · fp8 · H200 141GB
85.48%
±0.99%
$18.80 p002
141
Alibaba · Open weights · fp8 · B300 288GB
85.29%
±0.97%
$35.77 p002
142
Alibaba · Open weights · fp8 · B300 288GB
85.27%
±1.00%
$56.87 p001
143
Cohere · Open weights · fp8 · H200 141GB
85.23%
±1.01%
$131 p001
144
Google · Open weights · fp8 · H100 80GB
85.19%
±1.00%
$54.16 p001
145
OpenAI · Proprietary
85.00%
±1.01%
$185 p001
146
Alibaba · Open weights · awq-int4 · A100 80GB
84.84%
±1.00%
$53.82 p002
147
Alibaba · Open weights · awq-int4 · H100 80GB
84.82%
±1.02%
$32.60 p002
148
Alibaba · Open weights · awq-int4 · H200 141GB
84.82%
±1.02%
$52.61 p002
149
Alibaba · Open weights · awq-int4 · H200 141GB
84.74%
±1.01%
$34.60 p002
150
Alibaba · Open weights · bf16 · H200 141GB
84.45%
±1.02%
$17.33 p001
151
OpenAI · Proprietary
84.32%
±1.02%
$982 p001
152
Alibaba · Open weights · awq-int4 · H200 141GB
84.08%
±1.03%
$89.36 p001
153
NVIDIA · Open weights · fp8 · H200 141GB
83.93%
±1.02%
$22.03 p003
154
OpenAI · Proprietary
83.93%
±0.99%
$105 p002
155
OpenAI · Proprietary
83.77%
±1.04%
$127 p003
156
OpenAI · Proprietary
83.69%
±1.07%
$256 p001
157
OpenAI · Proprietary
83.56%
±1.02%
$63.56 p003
158
OpenAI · Proprietary
83.56%
±1.06%
$400 p001
159
Google · Open weights · fp8 · H200 141GB
83.44%
±1.02%
$12.38 p001
160
Alibaba · Open weights · awq-int4 · A100 80GB
83.44%
±1.04%
$137 p001
161
Alibaba · Open weights · awq-int4 · H200 141GB
83.42%
±1.03%
$137 p001
162
Alibaba · Open weights · awq-int4 · H100 80GB
83.40%
±1.03%
$83.31 p001
163
Google · Open weights · fp8 · H100 80GB
83.30%
±1.02%
$27.65 p002
164
OpenAI · Proprietary
82.90%
±1.05%
$35.26 p002
165
Alibaba · Open weights · fp8 · H100 80GB
82.72%
±1.04%
$2.92 p002
166
Alibaba · Open weights · fp8 · H200 141GB
82.70%
±1.03%
$8.25 p002
167
OpenAI · Proprietary
82.70%
±1.07%
$96.14 p001
168
NVIDIA · Open weights · fp8 · H200 141GB
82.35%
±1.06%
$15.97 p002
169
Alibaba · Open weights · fp8 · B300 288GB
82.27%
±1.08%
$23.12 p002
170
Alibaba · Open weights · bf16 · H200 141GB
82.10%
±1.09%
$8.91 p002
171
NVIDIA · Open weights · fp8 · H200 141GB
81.86%
±1.12%
$41.77 p001
172
Meta · Open weights · awq-int4 · B300 288GB
81.53%
±1.09%
$81.28 p002
173
Google · Open weights · fp8 · H100 80GB
81.44%
±1.10%
$3.82 p003
174
Alibaba · Open weights · awq-int4 · H200 141GB
81.03%
±1.10%
$35.20 p002
175
Alibaba · Open weights · bf16 · A100 80GB
80.91%
±1.13%
$14.99 p001
176
Alibaba · Open weights · bf16 · H200 141GB
80.89%
±1.15%
$19.26 p001
177
Alibaba · Open weights · bf16 · H100 80GB
80.79%
±1.13%
$11.16 p001
178
OpenAI · Proprietary
80.54%
±1.14%
$93.99 p002
179
Z.ai · Open weights · bf16 · A100 80GB
80.46%
±1.12%
$4.97 p002
180
Z.ai · Open weights · bf16 · A100 80GB
79.78%
±1.11%
$13.51 p001
181
Mistral AI · Open weights · fp8 · H200 141GB
79.39%
±1.11%
$13.80 p002
182
Mistral AI · Open weights · fp8 · H100 80GB
79.28%
±1.13%
$8.47 p002
183
Mistral AI · Open weights · fp8 · A100 80GB
79.00%
±1.13%
$20.41 p002
184
Google · Open weights · fp8 · H200 141GB
78.89%
±1.15%
$5.38 p002
185
Google · Open weights · fp8 · H100 80GB
78.89%
±1.10%
$28.04 p003
186
Google · Open weights · fp8 · H100 80GB
78.87%
±1.15%
$2.21 p002
187
Tencent · Open weights · gptq-int4 · A100 80GB
78.73%
±1.16%
$32.98 p001
188
Z.ai · Open weights · bf16 · A100 80GB
77.60%
±1.17%
$12.16 p001
189
Mistral AI · Open weights · fp8 · A100 80GB
77.17%
±1.19%
$48.92 p001
190
Google · Open weights · fp8 · H100 80GB
77.12%
±1.17%
$21.64 p002
191
OpenAI · Proprietary
76.98%
±1.17%
$125 p001
192
Mistral AI · Open weights · fp8 · H100 80GB
76.75%
±1.22%
$18.36 p001
193
Mistral AI · Open weights · fp8 · H200 141GB
76.61%
±1.25%
$28.17 p001
194
Alibaba · Open weights · bf16 · H100 80GB
76.34%
±1.17%
$4.29 p002
195
Alibaba · Open weights · bf16 · H200 141GB
76.28%
±1.19%
$7.56 p002
196
Z.ai · Open weights · bf16 · A100 80GB
75.70%
±1.19%
$4.42 p002
197
AllenAI · Open weights · fp8 · H100 80GB
73.98%
±1.21%
$24.33 p001
198
AllenAI · Open weights · fp8 · H100 80GB
73.32%
±1.25%
$9.56 p002
199
OpenAI · Proprietary
73.28%
±1.23%
$23.50 p002
200
IBM · Open weights · fp8 · H200 141GB
70.17%
±1.25%
$30.30 p001
201
Google · Open weights · fp8 · H200 141GB
70.13%
±1.27%
$10.45 p001
202
OpenAI · Proprietary
70.07%
±1.18%
$64.03 p001
203
IBM · Open weights · fp8 · H100 80GB
69.92%
±1.27%
$17.50 p001
204
NVIDIA · Open weights · fp8 · H100 80GB
69.22%
±1.36%
$2.89 p002
205
IBM · Open weights · bf16 · A100 80GB
69.02%
±1.26%
$28.04 p001
206
NVIDIA · Open weights · fp8 · H100 80GB
65.34%
±1.33%
$4.70 p003
207
Google · Open weights · fp8 · H100 80GB
65.07%
±1.33%
$2.50 p003
208
Microsoft · Open weights · bf16 · A100 80GB
63.98%
±1.29%
$5.78 p001
209
NVIDIA · Open weights · fp8 · H100 80GB
63.55%
±1.38%
$7.07 p001
210
Google · Open weights · fp8 · H100 80GB
61.65%
±1.37%
$1.72 p002
211
Google · Open weights · fp8 · H200 141GB
61.63%
±1.36%
$5.89 p002
212
OpenAI · Proprietary
59.06%
±1.36%
$37.51 p001
213
Meta · Open weights · bf16 · A100 80GB
58.82%
±1.41%
$11.30 p001
214
Meta · Open weights · fp8 · H200 141GB
58.05%
±1.42%
$12.18 p001
215
Meta · Open weights · fp8 · H100 80GB
57.99%
±1.41%
$6.09 p001
216
Meta · Open weights · fp8 · H100 80GB
44.31%
±1.39%
$2.49 p002
217
Meta · Open weights · fp8 · H200 141GB
44.21%
±1.38%
$5.45 p002
218
Meta · Open weights · bf16 · A100 80GB
39.85%
±1.37%
$4.27 p002

Compare Selected Runs

Select up to six runs from the table.

Method Notes

Public rows are rebuilt from verified run artifacts in the repository. Accuracy is recomputed from predictions, and official runs are checked against the registered dataset hash and prompt ID.

The ± value under each accuracy is half the width of a fixed-seed bootstrap 95% confidence interval. The small range under each rank lists the positions a run could plausibly occupy among the visible rows given overlapping intervals.

The Score against and Sense granularity controls re-score every run and baseline against one of nine scoring schemes: the lexEN v1, Maru 2022 (ALLamended), or original Raganato 2017 gold labels, each at WordNet 3.0 fine-grained sense or at one of two coarse concept levels — Glite or CSI. A prediction is coarse-correct when its WordNet sense maps to the same coarse concept as a gold sense (so coarse accuracy is always at least the fine-grained accuracy). The default and official score is lexEN v1 · WordNet fine-grained. See label schemes and coarsening for what each option means.

The Pareto frontier uses the visible rows and the selected chart metric (cost per million items or machine-hours per 1M items): higher accuracy is better, and a lower metric value is better. Starred rows are on the frontier.

Machine-hours per 1M items is recorded for self-hosted runs only. It measures the per-item evaluation loop on the recorded machine, excluding model and asset loading; cloud API runs are excluded from that chart. Because it measures time rather than money, it is the one hardware metric no pricing assumption touches: multiply it by whatever you would pay per hour to cost a model on your own terms.

Cost per million items prices self-hosted runs at a fixed reference rate for their GPU class, not at the rate each machine happened to be rented at, so a lucky or unlucky spot price neither flatters nor penalizes a model. Each run still records the rate it actually paid, and its run page shows both. As of 2026-07-14 the reference rates are $1.09/h for A100 80GB, $2.26/h for H100 80GB, $3.66/h for H200 141GB, $6.72/h for B300 288GB. A cost marked * is the exception: that run is on a GPU class we have no reference rate for, so it is priced at the rate its machine was rented at and is not on the same basis as the rows around it. Such a run can still reach the Pareto frontier, and its star is marked ★* to show the frontier position rests on that non-comparable cost. See methodology for how the rates are set.

Reference baselines are supervised WSD systems and MFS scored from per-item predictions on the same dataset items; dashed chart lines show them for context.

Comparing selected runs on the same dataset uses McNemar's test on paired per-item correctness, which detects differences that overlapping confidence intervals can miss.