Reference
Method & assumptions
What drives the numbers, the judgement calls behind them, and how much trust
each division's clusters deserve. This applies to all nine briefs; anything division-specific
is stated on that division's own page.
How this was made — and how to read it. The data pipeline, the clustering, and the
written analysis in these briefs and the player dashboard were produced with substantial help
from an AI assistant, working from data collected from public GeoGuessr profiles and augmented
manually where the public API only exposes a limited recent-games window. Every effort was
made to keep both the data and the analysis accurate — cross-checking figures, flagging thin
samples, and stating each division's data gaps openly. But two kinds of error are possible and
worth naming plainly: data error (a manually captured win-rate or duel count that's
slightly off, a profile that had changed since capture) and analytical error (a
clustering that reads structure into noise, a narrative line that over-claims from a small
sample). Read these as a careful enthusiast's scouting notes, not as an official record.
Where a specific number matters to you, the source is your own GeoGuessr profile.
01
What is PCA? (in plain terms)
The briefs lean on two techniques with intimidating names — principal
component analysis and clustering. Neither is as complicated as it sounds.
Imagine describing every player
by a fistful of numbers at once — win-rate, accuracy, how much they play, how they lean
across game modes. That's six numbers per player, and you can't plot six dimensions on a
flat screen. Principal component analysis (PCA) is a way of
squashing those six numbers down to two, while throwing away as little information as
possible — like choosing the single best camera angle to photograph a crowd so the fewest
people are hidden behind others.
The two axes it produces
(PC1 and PC2) aren't "win-rate" or "accuracy" directly. Each is a blend of the
original numbers — whichever blend spreads the players out the most. On each brief, the
axis labels tell you what that division's blend actually turned out to be, because
the blend is different in every division. Two players sitting
close together on the map are similar across all six measures; two far apart are different.
That's the whole idea: a 2-D picture where distance means dissimilarity.
Clustering
is the companion step: it groups the players who sit near one another into a handful of
named types (the "archetypes" on the dashboard, the coloured groups on each brief). The
caveat is that clustering will always hand back groups, even when the players don't
really fall into neat types — so each division reports a "silhouette" score measuring how
real the grouping is, and the weaker ones are flagged as such. A tidy-looking set of clusters
is not proof the division actually divides that cleanly.
Why the axes change between divisions. Because PCA finds whichever blend
spreads out that particular division's players, the same axis name means different
things in different briefs — in one division PC1 might be mostly about raw skill, in another
mostly about mode preference. Don't compare one division's map to another's; read each on its
own terms, using the axis labels on that page.
02
The two models
Every division is segmented twice, independently. Where the two agree, a
finding is robust; where they diverge, the disagreement is itself the signal.
Win-rate model (primary)
- current_overall — present skill level (mean of real current per-mode ELO)
- lt_overall_pct — lifetime overall win-rate
- wr_lean — moving-vs-restrictive lean, from win-rates
- wr_spread — win-rate consistency across modes
- avg_distance_km — raw accuracy
- log_duels — lifetime ranked experience / volume
Lifetime per-mode win-rate
and duel counts were captured manually — the API exposes only a 40-game window. Samples
are large (hundreds to tens of thousands of duels), so no imputation. A per-mode win-rate
is used only where a player has ≥30 games in that mode; thinner modes — most often NMPZ —
fall back to the player's overall rate.
ELO model (secondary)
- Current per-mode ratings, not peak — present form, not ceiling
- Missing current rating → the player's seeded (enrollment) value stands in, with
no decay or inflation: there are no last-played dates to justify any
- Mode-lean from real modes only — never imputed; undefined where a player lacks
the needed current ratings
- Delisted players are therefore fully seeded-derived. Which players those are is stated
in each division's own data-quality note
Tested & dropped
- Team-duel signal — from a player's best team; a strong or impulsive partner
confounds it
- modes_played — after imputation it measured data-completeness, not playstyle
- geo_us — near-independent of skill; added a noise axis. Kept as a separate
matchup lens instead (§01–03 of each brief)
- imputed restrictive_lean — mixed peak with current on different scales,
flipping signs
- volume-only — games and mode-mix without win-rate gave unstable clusters; the
two algorithms disagreed
- Each removal that raised separation confirmed the feature was noise
03
How much to trust each division
Silhouette scores measure cluster separation (roughly: 0 = no structure,
1 = perfectly distinct). At these sample sizes every value below is a soft structure — real,
but not sharp enough for a placement near a boundary to be treated as fixed.
| Division | n | Win-rate model | ELO model |
Data flags |
| Div 1A | 8 | 0.39 | 0.36 | 2 no live rating |
| Div 1B | 8 | 0.34 | 0.32 | — |
| Div 2 | 12 | 0.39 | 0.25 | 3 no live rating, 1 thin NMPZ |
| Div 3 | 12 | 0.33 | 0.35 | — |
| Div 4 | 12 | 0.36 | 0.27 | 2 no live rating |
| Div 5 | 12 | 0.29 | 0.33 | — |
| Div 6 | 11 | 0.27 | 0.30 | 2 no live rating, 5 thin NMPZ |
| Div 7 | 12 | 0.34 | 0.35 | 1 no live rating, 1 thin NMPZ |
| Div 8 | 10 | 0.43 | 0.25 | 6 no live rating, 7 thin NMPZ |
Reading the table. A high silhouette is not automatically a better
analysis. Division 8 scores the highest of any division on the win-rate model — and that is
precisely the trap: with six of ten players lacking a live rating and several carrying
double-digit lifetime samples, the model is cleanly separating players by how much data
exists about them, which looks like structure and isn't skill. Conversely Division 5's
low score reflects a genuinely homogeneous field, which is itself a finding. Read the score
together with the data flags, and with each brief's own data-quality note.
04
Standing caveats
- Seeded ELO is a frozen
snapshot — each player's peak rating at the moment of enrollment. It is not a live
peak, and for some players their true peak has since moved. It is labelled "seeded"
throughout rather than "peak" for that reason.
- Current ELO is a snapshot from
when the tournament began, and is deliberately not refreshed. Since the
tournament started, GeoGuessr introduced a soft rating reset (to curb rating inflation) and
rating decay for inactive players. Both move a player's live rating for reasons that have
nothing to do with their form in this tournament — so pulling fresh ELOs now would quietly
change the meaning of every "current" figure and break comparison with the seeded values.
The numbers here are frozen at the tournament's starting line on purpose; treat them as
"where everyone stood at kickoff", not "where they are today".
- Geographic data is ordinal —
top-3 and bottom-3 country lists, not continuous scores. It is deliberately excluded from
segmentation and kept as a separate matchup lens.
- Lifetime records are not
current form. A player whose ranked history predates a change in how they play will
be mis-read by both models — the record is old, the player isn't. Round-by-round results
correct this faster than lifetime data ever will.
- Percentages on tiny samples are anecdotes. A 100%
win-rate over three games is not a rating. Where a number rests on a thin sample, the
brief says so beside it.
- Win-rates mix modes and eras. A
lifetime win-rate blends Moving, No-Move and NMPZ games played across a player's whole
history, against opponents of every strength. Two players with the same headline win-rate
can have arrived there completely differently — the per-mode breakdowns exist because the
overall figure hides as much as it shows.
- The clustering is unsupervised
and pre-tournament. Nobody told the models what a "good" grouping looks like; they
found whatever structure the numbers held, before a single tournament game was played. They
know nothing about head-to-head history, current motivation, or who has been quietly
practising. They are a starting hypothesis, not a prediction.
- This is one analyst's lens. Different reasonable choices
— which features to include, how many groups to look for, where to draw the line on a thin
sample — would produce somewhat different briefs. The specific reads here are defensible,
not uniquely correct.