All articles
Analysis & Method11 min6 August 2026

The Maths That Decided the U16 Worlds

A ranking formula, not match points, decided who could compete for medals at the U16 World Championships in Zagreb, and the same ranking had already shaped the draw, months before a ball was played. What the formula did, why, and what I rebuilt to test it.

A ranking formula, not match points, decided who could compete for medals at the World Men's U16 Championships in Zagreb this month. The same ranking had already shaped the draw, months before the tournament began.

This piece looks at how the Tournament Performance Index worked, what it produced in practice, how I reconstructed it, and why the real issue is the assumptions built into it.


What happened

Russia beat Montenegro in a penalty shootout in Zagreb this month and earned 99.2 Tournament Performance Index points, one of the highest single-match scores reported at the tournament. Montenegro, meanwhile, pushed the eventual group leaders all the way to penalties, leading at half-time and drawing level twice in the fourth quarter. For that performance they received 38 points and dropped from ninth to eighteenth in the standings overnight.

That gap is the product of the Tournament Performance Index. Because Russia returned after four years away from international competition, they entered the tournament without a world ranking, and under the regulations that meant they were treated as the lowest-ranked team in the field for scoring purposes, even as they went on to top their group.

Much of the discussion around the tournament has focused on the mathematics behind the formula. The more interesting question is what the model was actually trying to measure. The calculations did exactly what they were designed to do. The real issue is whether the assumptions behind them were appropriate for deciding qualification at a world championship. To understand why, it's worth looking at how the system worked.


What the formula is

Group positions in Zagreb weren't decided by match points. Instead, each team received a Tournament Performance Index, calculated as the average of its four group-stage performances.

World Aquatics has never published a dedicated set of U16 competition regulations explaining the calculation. Every published tournament value, however, matches the Rating Points calculation from version 2.2 of the World Aquatics Water Polo World Ranking methodology, released in January 2026. That calculation has three components.

Basis Points

The first component rewards the match result using broad bands rather than a continuous scale. A win by one to four goals scores 65 points, by five to ten goals scores 70, and by eleven goals or more scores 75, the maximum available regardless of how much larger the margin gets. A shootout win scores 62 and a shootout loss scores 38. A defeat by one to four goals scores 35, by five to ten scores 30, and by eleven or more scores 25. The shootout figures are worth flagging: World Aquatics' own general ranking document states 60 and 40 for shootout results, not 62 and 38. The values that actually fit Zagreb's published scores are different from the ones World Aquatics has made public, which means the U16 event ran on parameters that aren't in any document I could find.

Home, Away and Neutral Points

The second component is straightforward. Hosts receive a three-point deduction, and every team that plays the host receives a three-point bonus. Every other match is treated as neutral.

Opposition Ranking Points

The final component adjusts the score according to the strength of the opposition, and it's where the model becomes more complicated. Each opponent's value is based on its World Aquatics ranking position, measured relative to the lowest-ranked team in the tournament. Version 2.2 also introduced a clause stating that any nation without an official world ranking is treated as occupying the last position in the rankings, plus one. In practice, that meant an unranked team was automatically treated as the weakest opponent in the field before a single ball had been thrown.

I rebuilt the calculation using the per-match values published throughout the tournament, and every score reproduced exactly. That gives me confidence the structure has been reconstructed correctly, even without an official U16 regulation to check it against.


What the formula produced

Once the index is broken down, some of its consequences become easier to understand. None of them are mathematical errors. They're the result of the assumptions built into the model.

Beating the best team was worth less than beating the worst

The clearest example came from Russia's return to international competition. Because Russia had no official world ranking, the index treated them as the weakest team in the tournament regardless of how they actually performed, so every side that faced them received a lower opposition value than they would have against almost anyone else in the field. Montenegro felt this directly: taking the eventual group winners to a penalty shootout earned them 38 points, the same kind of value the system awards for a narrow defeat to one of the weakest teams in the competition. Russia received 99.2 for the win, because Montenegro carried a strong world ranking. By the end of the group stage, the team that had topped it had been treated throughout as the least valuable opponent anyone could face. The formula was working exactly as built, relying on a pre-tournament ranking that no longer reflected the reality in the pool.

A narrow defeat could be worth almost as much as a comfortable win

The same logic showed up elsewhere. Italy lost to Montenegro by two goals and received 72.2 points. The United States beat Zimbabwe by twenty-five goals and received 76.2. Those scores aren't identical, but they're remarkably close given the gap between the two performances, because the model weighted the perceived strength of the opposition heavily enough that a narrow defeat to a highly ranked team could score close to a dominant win over a much weaker one. Whether that's desirable depends on what the index was meant to reward.

Once the margin reached eleven, extra goals stopped mattering

The Basis Points bands introduce a second modelling choice: any win by eleven goals or more receives the maximum available, and nothing beyond that point changes the score. Montenegro's 34:2 win over Zimbabwe and the United States' 26:1 win over the same opponent both scored exactly 76.2, one match won by thirty-two goals, the other by twenty-five, with no difference between them as far as the index was concerned. That's a deliberate design decision: the model assumes that, past a certain point, an increasingly large margin adds little further information about team strength. Whether that assumption is the right one is worth debating.

The standings reflected the model

Going into the final round of group matches, both the United States and Serbia had won all three of their games, yet sat tenth and eleventh, because other teams had accumulated higher index scores.


The same ranking shaped the draw

The index wasn't the only place the World Aquatics ranking mattered. Months before the first whistle in Zagreb, the same ranking was used to seed the tournament. The 36 teams were split into two pots of 18 based on the World Ranking, and each team then played two matches against opponents from its own pot and two from the other. Russia and Egypt were both placed in the lower pot; China, ranked sixth in the world, was seeded into the stronger one.

On its own, that's unremarkable. Seeding by world ranking is standard practice across most sports. The difficulty is that the same ranking was then used a second time, to decide how much each result was worth. Any inaccuracy in the underlying ranking therefore shaped both who you played and how valuable beating them turned out to be.

China illustrates this well. Despite being seeded sixth, they lost all three of their group matches played so far. Egypt beat them and received 96.2 points for the win, the second-highest single-match score reported at the tournament, because China's pre-tournament ranking still classified them as one of the strongest teams in the field. The ranking had overestimated China's strength before the tournament began, and when Egypt beat them, that overestimation was carried straight into the index. The same piece of information shaped the competition twice: once building the draw, and again scoring the result.


Why the problem sits upstream

It's tempting to conclude that the Tournament Performance Index itself is flawed. I don't think that's quite right. The index only uses the information it's given, and if the underlying ranking doesn't estimate team strength accurately, no amount of arithmetic downstream can correct for that.

The World Aquatics ranking wasn't built as a pure measure of playing strength. It's a cumulative system running on an eight-year cycle, and one of its stated purposes is to encourage participation in World Aquatics competitions. It records activity as much as ability, which is sensible if the goal is to reward federations for showing up regularly. It becomes a problem when that same ranking is treated as a current estimate of team strength. A federation playing dozens of age-group matches over several seasons steadily accumulates ranking points; another may play far fewer matches while still fielding a stronger team at a given championship. Those are different questions, and one ranking isn't built to answer both.

Seen that way, most of what happened in Zagreb stops looking strange. Russia were undervalued because the ranking held no meaningful information about their current level. China were overvalued because the model assumed their pre-tournament ranking still reflected their strength. The index simply carried those assumptions forward, faithfully. The issue sits upstream of the maths: the model behaved exactly as designed, and the real question is whether the data feeding it was fit to decide a world championship.


What I rebuilt

Rather than asking whether the index was right or wrong, I wanted to test a different question: could a ranking built entirely from the matches played in Zagreb produce a table that better reflected what actually happened in the pool?

I rebuilt the first three rounds using only the 54 matches played, with no external ranking involved. Instead of fixed bands, results are scored on a continuous curve, so winning by twelve goals is worth slightly more than winning by eleven, with diminishing returns as the margin grows, and no threshold where extra goals suddenly stop mattering. Opponent strength isn't looked up from a pre-tournament list either. Every team's rating is solved simultaneously from the tournament's own results: a team's value depends on who they beat, how convincingly, and how those opponents performed against everyone else. An unranked entrant needs no special rule under this approach. Their strength simply emerges from what they did that week.

What changed

Ranked first by wins, with the rebuilt rating used only to separate teams on the same record, every team that had won its opening three matches occupied the top eight places, regardless of how heavily opponent strength was weighted. China illustrates the difference well. The original system treated them as an elite opponent because of their pre-tournament ranking, despite three group defeats. The rebuilt model has no knowledge of that ranking, sees three losses, and rates China accordingly, so Egypt's win over them is credited for what it proved on the day rather than what China had been expected to be beforehand.

The rebuild doesn't suggest the original seeding pots were completely wrong, either. On average, teams from the stronger pot still rate higher than teams from the weaker one, so the broad picture was correct. Where the model differs is in recognising that specific teams, Russia and Egypt among them, performed well above the level their seeding implied.

The limits of any model

No ranking system removes uncertainty from a four-game tournament. New Zealand also won all three of their opening matches, but the rebuilt model rates them close to the average strength of the stronger pot rather than among the tournament's favourites. Their schedule hadn't yet produced enough evidence to separate genuine title-winning strength from a strong record against moderate opposition. That's an unavoidable limit of short tournaments. The goal isn't to remove uncertainty entirely. It's to avoid adding more of it through assumptions the competition itself doesn't support.


What could change

The most complete fix addresses both the draw and the scoring. Rather than fixing the schedule months in advance from an external ranking, teams could be paired after each round against opponents with the same record, in the Swiss-system style chess has used for decades wherever a large field plays too few rounds for everyone to face everyone else. Ties could then be broken using Buchholz, which scores a team by the strength of the opposition it actually faced during the tournament, not on a list compiled beforehand. This wouldn't require more matches or a different format. Zagreb already runs as four sequential match days rather than fixed round-robin groups, so the only change is how opponents are assigned after the opening round.

A simpler option keeps the existing index but shrinks its role. Rank teams by wins first, and use the Tournament Performance Index only to separate teams tied on the same record. That keeps the extra information the model captures without letting it outweigh the results it's meant to be measuring.


Final thoughts

No statistical model is perfect. Every ranking system reflects a set of assumptions about what should matter and how different pieces of information should be combined, and the Tournament Performance Index is no exception. The calculations are internally consistent, and I've been able to reproduce the published values exactly. The more important question is whether the assumptions behind them were the right ones for deciding who competed for medals at a world championship.

That's the real lesson from Zagreb. The model didn't fail to do what it was designed to do. It relied on an external ranking that shaped the tournament twice, once by building the draw, and again by deciding how much each result was worth. Statistical models are most useful when they measure the thing they're meant to measure. Once they become part of the rules of competition, transparency, validation and careful design matter just as much as the mathematics itself.