Why do only 4 out of 40 models beat a coin flip on hard questions?
https://future-wiki.win/index.php/Onboarding_Documentation_from_AI_Sessions:_Transforming_Ephemeral_Conversations_into_Enterprise_Assets
In the field of Applied NLP, we have developed a dangerous habit: we treat LLM leaderboards like the final scoreboard of a professional sports league. If a model tops the chart, we assume it is "smarter