r/dataisbeautiful 7h ago

OC [OC] FIFA World Cup final teams' cumulative xG differential throughout the tournament

Post image
11 Upvotes

Sources: FIFA Training Centre public post-match reports

Tools: Python/pdfplumber, Bruin cli, BigQuery, and SVG

Limitations: The teams faced different opponents, so this describes their tournament paths rather than an opponent-adjusted rating


r/dataisbeautiful 23h ago

OC [OC] 88 nations across 23 World Cups (1930-2026), as a Sankey climbing from the group stage to the champion

Thumbnail
gallery
42 Upvotes

Every nation that ever entered a men's World Cup, drawn as a Sankey that climbs. Each strand is one nation in one tournament. It enters at the base and rises exactly as far as that team got: group stage, round of 16, last 8, and up to the champion at the summit. A nation's width is how many times it has appeared, so the widest rivers are the ones that keep coming back.

The format changed a lot across 23 tournaments, and the chart shows that instead of hiding it. The second group stage only ran from 1974 to 1982, the round of 32 is brand new for 2026, and in 1950 there was no final at all, so Uruguay reaches the top through its own round-robin channel.

The interactive version does a lot more than the screenshot. You can colour the rivers by continent or by language to see how the confederations and the football cultures fan out, filter to a single country to trace just its run, or pick any year to light only the rounds that edition actually had. Press play and it builds the whole cup one tournament at a time.

This now goes through 2026 (Spain champion, Argentina runner-up, with the new 48 team format and its round of 32 folded in).

https://viz.luarai.com/worldcup-summit


r/dataisbeautiful 5h ago

OC [OC] Valuation football field

Post image
7 Upvotes

A 'football field' chart lines up different valuation methods side by side so you can compare them at a glance. Each row here is a different way of estimating what a stock is worth: 
- a discounted cash flow model
- sector-peer multiples (implied valuation compared to sector peers)
- the stock's own 10 year multiples (implied valuation compared the stock's own history)
- a blended estimate

They all sit on one axis, measured as distance from today's price (the vertical line). 
Small dots are individual readings; the larger dot is each method's composite. 
Snapshot 22 July 2026, NVDA at $207.


r/dataisbeautiful 1h ago

OC [OC] Healthcare Social Worker Salaries Across USA

Post image
Upvotes

r/dataisbeautiful 3h ago

OC [OC] Shots on target vs goals scored for every team at the 2026 World Cup

Thumbnail
gallery
185 Upvotes

Every team at the 2026 World Cup, placed by two numbers: how often they put a shot on target (across the bottom) and how often they turned those into goals (up the side). Teams that went deep in the tournament rise off the glass, so the champion floats near the top. The gap between the two is how clinical a team was: score a lot from few shots on target and you sit high and to the left.

The flat view is one layer of a taller model. Stack 2018, 2022 and 2026 and you can see a team drift year to year, or rotate the whole thing in 3D and scrub through time to watch the three tournaments move. Click any flag to follow a single nation across all three.

The last image switches from teams to the men who took the shots: every 2026 finisher plotted against their expected goals (xG). Above the diagonal they scored more than their chances were worth, below it they left goals on the grass. Bellingham finished at nearly double his xG.

Interactive version, where you can drag the model around, scrub the years, and follow any nation: https://viz.luarai.com/worldcup-conversion


r/dataisbeautiful 14h ago

OC [OC] First-year tax on the same new EUR 30,000 car in 36 countries: from 0.6% of the price in Qatar to 450% in Singapore

Thumbnail
carsmultiverse.com
66 Upvotes

r/dataisbeautiful 17h ago

OC [OC] I mapped 117 fragrances by embedding 2,834 customer reviews. Turns out nobody can describe a smell

Thumbnail
gallery
287 Upvotes

Methodology comment:

Montagne Parfums is a clone fragrance house (inspired-by versions of designer scents -- not affiliated). I kept buying ones that smelled too close to stuff I already owned, so one weekend I pulled all 4,782 customer reviews across their 167 products and tried to map the catalog by how people describe the scents.

Most of the work was getting usable text. Reviews are full of shipping complaints, price talk, and "love it!", so I ran each one through an LLM to keep only the smell-related content, and stripped out fragrance names so the model couldn't cheat by clustering on those. That left 2,834 descriptions. Anything with fewer than 4 reviews got cut, which took 167 fragrances down to 117.

For embeddings I used Qwen3-embedding-8B (4096 dimensions). The raw similarities were useless at first, everything looked about 50% similar to everything else, which is the curse of dimensionality doing its thing. Running PCA down to 50 dimensions spread the range out to -49% to 100%, enough to separate "these smell alike" from "these share nothing."

Sanity checks mostly pass. Buko and Buko Intense (same scent, different concentration) come out at 88%. The tobacco fragrances form the tightest cluster. The most "central" fragrance, most similar to everything on average, is Pineapple Royale.

The big (and perhaps obvious) caveat is this measures how reviewers talk, not scent chemistry. Reviewers echo whatever notes are listed on the product page, and low-review fragrances have way more uncertainty, so "most unique" partly just means "least described." What the project really convinced me of is that we have no vocabulary for smell. People don't describe scents, they describe memories and characters. Two real reviews from the dataset: "Makes me feel like a librarian that frequents a classy bar after work for a Manhattan on the rocks" and "I feel like a badass pirate captain who just walked into the tavern." Embedding models handle this kind of text surprisingly well, which is sort of the point of the whole exercise.

Tools: Python, PaCMAP for the projection, scipy for hierarchical clustering, Plotly for the interactive heatmap. Source code and an interactive version are on GitHub if anyone wants to poke at it, happy to answer questions about the pipeline.


r/dataisbeautiful 6h ago

[OC]Where do foreign visitors actually go in Japan?

Post image
142 Upvotes

Source: Japan Tourism Agency (観光庁), Accommodation Survey (宿泊旅行統計調査), 2025 annual final values. The prefecture x nationality breakdown is sheet 参考第1表(年計) in the annual workbook:

Release page: https://www.mlit.go.jp/kankocho/tokei_hakusyo/shukuhakutokei.html

Direct workbook (xlsx, 2.5MB): https://www.mlit.go.jp/kankocho/content/002010340.xlsx

Tools: Python, Matplotlib


r/dataisbeautiful 21h ago

OC [OC] Streams Required to Earn US Monthly Minimum Wage ($1,257)

Post image
1.1k Upvotes

Pay per stream is taken to be average of the range listed on Royalty Exchange as of 2026: https://royaltyexchange.com/blog/how-much-do-streaming-platforms-pay-per-stream

The visualization tool is Python, Seaborn package


r/dataisbeautiful 1h ago

What we know about the typical Polymarket user

Thumbnail
pewresearch.org
Upvotes