Reference

Method & assumptions

What drives the numbers, the judgement calls behind them, and how much trust each division's clusters deserve. This applies to all nine briefs; anything division-specific is stated on that division's own page.

How this was made — and how to read it. The data pipeline, the clustering, and the written analysis in these briefs and the player dashboard were produced with substantial help from an AI assistant, working from data collected from public GeoGuessr profiles and augmented manually where the public API only exposes a limited recent-games window. Every effort was made to keep both the data and the analysis accurate — cross-checking figures, flagging thin samples, and stating each division's data gaps openly. But two kinds of error are possible and worth naming plainly: data error (a manually captured win-rate or duel count that's slightly off, a profile that had changed since capture) and analytical error (a clustering that reads structure into noise, a narrative line that over-claims from a small sample). Read these as a careful enthusiast's scouting notes, not as an official record. Where a specific number matters to you, the source is your own GeoGuessr profile.
01

What is PCA? (in plain terms)

The briefs lean on two techniques with intimidating names — principal component analysis and clustering. Neither is as complicated as it sounds.

Imagine describing every player by a fistful of numbers at once — win-rate, accuracy, how much they play, how they lean across game modes. That's six numbers per player, and you can't plot six dimensions on a flat screen. Principal component analysis (PCA) is a way of squashing those six numbers down to two, while throwing away as little information as possible — like choosing the single best camera angle to photograph a crowd so the fewest people are hidden behind others.

The two axes it produces (PC1 and PC2) aren't "win-rate" or "accuracy" directly. Each is a blend of the original numbers — whichever blend spreads the players out the most. On each brief, the axis labels tell you what that division's blend actually turned out to be, because the blend is different in every division. Two players sitting close together on the map are similar across all six measures; two far apart are different. That's the whole idea: a 2-D picture where distance means dissimilarity.

Clustering is the companion step: it groups the players who sit near one another into a handful of named types (the "archetypes" on the dashboard, the coloured groups on each brief). The caveat is that clustering will always hand back groups, even when the players don't really fall into neat types — so each division reports a "silhouette" score measuring how real the grouping is, and the weaker ones are flagged as such. A tidy-looking set of clusters is not proof the division actually divides that cleanly.

Why the axes change between divisions. Because PCA finds whichever blend spreads out that particular division's players, the same axis name means different things in different briefs — in one division PC1 might be mostly about raw skill, in another mostly about mode preference. Don't compare one division's map to another's; read each on its own terms, using the axis labels on that page.

02

The two models

Every division is segmented twice, independently. Where the two agree, a finding is robust; where they diverge, the disagreement is itself the signal.

Win-rate model (primary)

  • current_overall — present skill level (mean of real current per-mode ELO)
  • lt_overall_pct — lifetime overall win-rate
  • wr_lean — moving-vs-restrictive lean, from win-rates
  • wr_spread — win-rate consistency across modes
  • avg_distance_km — raw accuracy
  • log_duels — lifetime ranked experience / volume

Lifetime per-mode win-rate and duel counts were captured manually — the API exposes only a 40-game window. Samples are large (hundreds to tens of thousands of duels), so no imputation. A per-mode win-rate is used only where a player has ≥30 games in that mode; thinner modes — most often NMPZ — fall back to the player's overall rate.

ELO model (secondary)

  • Current per-mode ratings, not peak — present form, not ceiling
  • Missing current rating → the player's seeded (enrollment) value stands in, with no decay or inflation: there are no last-played dates to justify any
  • Mode-lean from real modes only — never imputed; undefined where a player lacks the needed current ratings
  • Delisted players are therefore fully seeded-derived. Which players those are is stated in each division's own data-quality note

Tested & dropped

  • Team-duel signal — from a player's best team; a strong or impulsive partner confounds it
  • modes_played — after imputation it measured data-completeness, not playstyle
  • geo_us — near-independent of skill; added a noise axis. Kept as a separate matchup lens instead (§01–03 of each brief)
  • imputed restrictive_lean — mixed peak with current on different scales, flipping signs
  • volume-only — games and mode-mix without win-rate gave unstable clusters; the two algorithms disagreed
  • Each removal that raised separation confirmed the feature was noise
03

How much to trust each division

Silhouette scores measure cluster separation (roughly: 0 = no structure, 1 = perfectly distinct). At these sample sizes every value below is a soft structure — real, but not sharp enough for a placement near a boundary to be treated as fixed.

DivisionnWin-rate modelELO model Data flags
Div 1A80.390.362 no live rating
Div 1B80.340.32
Div 2120.390.253 no live rating, 1 thin NMPZ
Div 3120.330.35
Div 4120.360.272 no live rating
Div 5120.290.33
Div 6110.270.302 no live rating, 5 thin NMPZ
Div 7120.340.351 no live rating, 1 thin NMPZ
Div 8100.430.256 no live rating, 7 thin NMPZ

Reading the table. A high silhouette is not automatically a better analysis. Division 8 scores the highest of any division on the win-rate model — and that is precisely the trap: with six of ten players lacking a live rating and several carrying double-digit lifetime samples, the model is cleanly separating players by how much data exists about them, which looks like structure and isn't skill. Conversely Division 5's low score reflects a genuinely homogeneous field, which is itself a finding. Read the score together with the data flags, and with each brief's own data-quality note.

04

Standing caveats

  • Seeded ELO is a frozen snapshot — each player's peak rating at the moment of enrollment. It is not a live peak, and for some players their true peak has since moved. It is labelled "seeded" throughout rather than "peak" for that reason.
  • Current ELO is a snapshot from when the tournament began, and is deliberately not refreshed. Since the tournament started, GeoGuessr introduced a soft rating reset (to curb rating inflation) and rating decay for inactive players. Both move a player's live rating for reasons that have nothing to do with their form in this tournament — so pulling fresh ELOs now would quietly change the meaning of every "current" figure and break comparison with the seeded values. The numbers here are frozen at the tournament's starting line on purpose; treat them as "where everyone stood at kickoff", not "where they are today".
  • Geographic data is ordinal — top-3 and bottom-3 country lists, not continuous scores. It is deliberately excluded from segmentation and kept as a separate matchup lens.
  • Lifetime records are not current form. A player whose ranked history predates a change in how they play will be mis-read by both models — the record is old, the player isn't. Round-by-round results correct this faster than lifetime data ever will.
  • Percentages on tiny samples are anecdotes. A 100% win-rate over three games is not a rating. Where a number rests on a thin sample, the brief says so beside it.
  • Win-rates mix modes and eras. A lifetime win-rate blends Moving, No-Move and NMPZ games played across a player's whole history, against opponents of every strength. Two players with the same headline win-rate can have arrived there completely differently — the per-mode breakdowns exist because the overall figure hides as much as it shows.
  • The clustering is unsupervised and pre-tournament. Nobody told the models what a "good" grouping looks like; they found whatever structure the numbers held, before a single tournament game was played. They know nothing about head-to-head history, current motivation, or who has been quietly practising. They are a starting hypothesis, not a prediction.
  • This is one analyst's lens. Different reasonable choices — which features to include, how many groups to look for, where to draw the line on a thin sample — would produce somewhat different briefs. The specific reads here are defensible, not uniquely correct.