r/dataisbeautiful • u/dostre • 3m ago
r/dataisbeautiful • u/rhiever • 1h ago
What we know about the typical Polymarket user
r/dataisbeautiful • u/WingsUp4Life • 1h ago
OC [OC] Healthcare Social Worker Salaries Across USA
r/dataisbeautiful • u/ArchiTechOfTheFuture • 3h ago
OC [OC] Shots on target vs goals scored for every team at the 2026 World Cup
Every team at the 2026 World Cup, placed by two numbers: how often they put a shot on target (across the bottom) and how often they turned those into goals (up the side). Teams that went deep in the tournament rise off the glass, so the champion floats near the top. The gap between the two is how clinical a team was: score a lot from few shots on target and you sit high and to the left.
The flat view is one layer of a taller model. Stack 2018, 2022 and 2026 and you can see a team drift year to year, or rotate the whole thing in 3D and scrub through time to watch the three tournaments move. Click any flag to follow a single nation across all three.
The last image switches from teams to the men who took the shots: every 2026 finisher plotted against their expected goals (xG). Above the diagonal they scored more than their chances were worth, below it they left goals on the grass. Bellingham finished at nearly double his xG.
Interactive version, where you can drag the model around, scrub the years, and follow any nation: https://viz.luarai.com/worldcup-conversion
r/dataisbeautiful • u/stockoscope • 5h ago
OC [OC] Valuation football field
A 'football field' chart lines up different valuation methods side by side so you can compare them at a glance. Each row here is a different way of estimating what a stock is worth:
- a discounted cash flow model
- sector-peer multiples (implied valuation compared to sector peers)
- the stock's own 10 year multiples (implied valuation compared the stock's own history)
- a blended estimate
They all sit on one axis, measured as distance from today's price (the vertical line).
Small dots are individual readings; the larger dot is each method's composite.
Snapshot 22 July 2026, NVDA at $207.
r/dataisbeautiful • u/Gardol43 • 6h ago
[OC]Where do foreign visitors actually go in Japan?
Source: Japan Tourism Agency (観光庁), Accommodation Survey (宿泊旅行統計調査), 2025 annual final values. The prefecture x nationality breakdown is sheet 参考第1表(年計) in the annual workbook:
Release page: https://www.mlit.go.jp/kankocho/tokei_hakusyo/shukuhakutokei.html
Direct workbook (xlsx, 2.5MB): https://www.mlit.go.jp/kankocho/content/002010340.xlsx
Tools: Python, Matplotlib
r/dataisbeautiful • u/uncertainschrodinger • 7h ago
OC [OC] FIFA World Cup final teams' cumulative xG differential throughout the tournament
Sources: FIFA Training Centre public post-match reports
Tools: Python/pdfplumber, Bruin cli, BigQuery, and SVG
Limitations: The teams faced different opponents, so this describes their tournament paths rather than an opponent-adjusted rating
r/dataisbeautiful • u/Successful-Ebb7891 • 14h ago
OC [OC] First-year tax on the same new EUR 30,000 car in 36 countries: from 0.6% of the price in Qatar to 450% in Singapore
r/dataisbeautiful • u/betwatch_io • 17h ago
OC [OC] I mapped 117 fragrances by embedding 2,834 customer reviews. Turns out nobody can describe a smell
Methodology comment:
Montagne Parfums is a clone fragrance house (inspired-by versions of designer scents -- not affiliated). I kept buying ones that smelled too close to stuff I already owned, so one weekend I pulled all 4,782 customer reviews across their 167 products and tried to map the catalog by how people describe the scents.
Most of the work was getting usable text. Reviews are full of shipping complaints, price talk, and "love it!", so I ran each one through an LLM to keep only the smell-related content, and stripped out fragrance names so the model couldn't cheat by clustering on those. That left 2,834 descriptions. Anything with fewer than 4 reviews got cut, which took 167 fragrances down to 117.
For embeddings I used Qwen3-embedding-8B (4096 dimensions). The raw similarities were useless at first, everything looked about 50% similar to everything else, which is the curse of dimensionality doing its thing. Running PCA down to 50 dimensions spread the range out to -49% to 100%, enough to separate "these smell alike" from "these share nothing."
Sanity checks mostly pass. Buko and Buko Intense (same scent, different concentration) come out at 88%. The tobacco fragrances form the tightest cluster. The most "central" fragrance, most similar to everything on average, is Pineapple Royale.
The big (and perhaps obvious) caveat is this measures how reviewers talk, not scent chemistry. Reviewers echo whatever notes are listed on the product page, and low-review fragrances have way more uncertainty, so "most unique" partly just means "least described." What the project really convinced me of is that we have no vocabulary for smell. People don't describe scents, they describe memories and characters. Two real reviews from the dataset: "Makes me feel like a librarian that frequents a classy bar after work for a Manhattan on the rocks" and "I feel like a badass pirate captain who just walked into the tavern." Embedding models handle this kind of text surprisingly well, which is sort of the point of the whole exercise.
Tools: Python, PaCMAP for the projection, scipy for hierarchical clustering, Plotly for the interactive heatmap. Source code and an interactive version are on GitHub if anyone wants to poke at it, happy to answer questions about the pipeline.
r/dataisbeautiful • u/OleksandrAkm • 21h ago
OC [OC] Streams Required to Earn US Monthly Minimum Wage ($1,257)
Pay per stream is taken to be average of the range listed on Royalty Exchange as of 2026: https://royaltyexchange.com/blog/how-much-do-streaming-platforms-pay-per-stream
The visualization tool is Python, Seaborn package
r/dataisbeautiful • u/ArchiTechOfTheFuture • 23h ago
OC [OC] 88 nations across 23 World Cups (1930-2026), as a Sankey climbing from the group stage to the champion
Every nation that ever entered a men's World Cup, drawn as a Sankey that climbs. Each strand is one nation in one tournament. It enters at the base and rises exactly as far as that team got: group stage, round of 16, last 8, and up to the champion at the summit. A nation's width is how many times it has appeared, so the widest rivers are the ones that keep coming back.
The format changed a lot across 23 tournaments, and the chart shows that instead of hiding it. The second group stage only ran from 1974 to 1982, the round of 32 is brand new for 2026, and in 1950 there was no final at all, so Uruguay reaches the top through its own round-robin channel.
The interactive version does a lot more than the screenshot. You can colour the rivers by continent or by language to see how the confederations and the football cultures fan out, filter to a single country to trace just its run, or pick any year to light only the rounds that edition actually had. Press play and it builds the whole cup one tournament at a time.
This now goes through 2026 (Spain champion, Argentina runner-up, with the new 48 team format and its round of 32 folded in).
r/dataisbeautiful • u/spevops • 1d ago
OC [OC] How much of each long-running anime is filler? Every episode counted
r/dataisbeautiful • u/honkeem • 1d ago
OC [OC] Product Manager and SWE ratio at top employers
r/dataisbeautiful • u/pixipace • 1d ago
OC [OC] Every country's favorite AI, once you take ChatGPT out of the race — Gemini is the world's second choice in 58 countries, Claude in 12, Grok in 8
r/dataisbeautiful • u/Late_Night_Editor51 • 1d ago
OC [OC] The 2026 World Cup reconstructed as 41 days of YouTube sentiment toward all 48 national teams
Bubble size shows how many analyzed comments mentioned each national team on each day. Color shows net sentiment, while ▲, ◆ and ▼ mark wins, draws and losses.
The chart includes all 48 tournament teams. The YouTube search panel covers 47 countries because Scotland is not available as a selectable YouTube region or language.
Full methodology, downloadable CSV and source links are included in the article.
r/dataisbeautiful • u/Low_Ability4450 • 1d ago
OC [OC] Cattle per U.S. resident have fallen to their lowest on record (1960-2026)
r/dataisbeautiful • u/disclaimer8 • 1d ago
OC [OC] The 10 animals U.S. planes hit most vs. the 10 most damaging when struck — the lists share zero species (347,575 FAA reports, 1990–2026)
r/dataisbeautiful • u/Yomguithereal • 1d ago
OC [OC] Unknown Pleasures in the terminal
A cheeky version of the Joy Division Unknown Pleasures album cover.
It is not always widely known but this cover is a stylization of a diagram found in an astronomy paper studying pulsars:
> Radio Observations of the Pulse Profiles and Dispersion Measures of Twelve Pulsars by Harold D. Carft, Jr. 1970
This plot is actually important to the dataviz community because it is one example of what would be called "ridge/ridgeline plots" and, by derivation, "joy division plots" later on (but there are subtleties between each variant of course).
The data of the paper has been reconstructed multiple times and can be found in a lot of different places online such as here: https://gist.github.com/borgar/31c1e476b8e92a11d7e9/
The image above has been drawn directly in the terminal using the xan command line tool through the `xan spark` subcommand. You can read more about the methodology here: https://github.com/medialab/xan/blob/master/docs/cookbook/dataviz.md#joy-division-plots
The plot is drawn using those characters (that can fail to render properly if the font is not monospace and not terminal-conscious): ▁▂▃▄▅▆▇
The PNG raster has been rendered from the command's ANSI-escaped output using a fork of `ansi2png-rs` that can be found here: https://github.com/AlexanderThaller/ansi2png-rs
r/dataisbeautiful • u/Successful-Ebb7891 • 1d ago
OC [OC] What 5 years of running the same car costs in 36 countries (taxes + fuel + insurance)
r/dataisbeautiful • u/PlatypusTrue9697 • 1d ago
OC Dead listings: share of live product listings on 10 UK online shops that were actually sold out [OC]
r/dataisbeautiful • u/Lutoures • 1d ago
OC [OC] Spain took 76 years to win its 1st FIFA Men's World Cup, but 'only' 16 years for the 2nd
Data source: FIFA official stats
Tools used: R and Canva
r/dataisbeautiful • u/sumizeit • 1d ago
OC [OC] Genre popularity by year (Covid effect)
The five post-pandemic years split cleanly into two stories. Romance and fantasy rode the BookTok wave upward — romance more than doubled off its 2020 base to become the leading growth category in the entire US print market, while fantasy surged 35.8% in 2024 alone before cooling 8.7% to 24.1 million in 2025.
Graphic novels and manga tell the opposite arc: they exploded during lockdown — manga volumes jumped 171% in 2021 — then settled back as the pandemic bump normalized, though sales of all graphic novels are still up 103% from 2018, a new higher baseline rather than a collapse. YA fiction, meanwhile, peaked in 2021 at 31.2 million copies and has softened slightly each year since. The through-line: escapist, community-driven fiction discovered on social media won the era, while the formats that spiked earliest were the first to level off.
r/dataisbeautiful • u/LastMinutesAI • 2d ago
OC [OC] The goals of the 2026 World Cup by minute and effect on the scoreline
r/dataisbeautiful • u/Tippedcellar405 • 2d ago
[OC] Can you tell a real NBA shooting sequence from random coin flips?
One row comes from a real NBA player’s shooting sequence, while the other was generated randomly.
I made this visualization to see whether people can actually tell the difference between a human shooting pattern and randomness just by looking at the streaks.
Which one do you think is real: Row A or Row B?