A consolidated ranking based on the evaluations and latest reports in our dataset, scored with Bradley-Terry and Elo pairwise methods. How the ranking works →
Filters open-weight models with fewer than 40 billion parameters from the Overall ranking, preserves its raw Elo ratings and relative order, then normalizes within this board.