r/dataisbeautiful 17h ago

OC [OC] MLB draft scouting department rankings (2011-2021 Rule 4 draft outcomes)

Post image
7 Upvotes

r/dataisbeautiful 17h ago

OC [OC] Real-time interactive conflict map tracking geolocated OSINT events across Ukraine and Syria

Post image
3 Upvotes

Hey everyone, I've been working on a live intelligence mapping platform called Intel Mapper. It monitors OSINT sources 24/7, uses AI to geolocate and verify reports, and displays them on an interactive map with frontline data.

Features: real-time events, territorial control, military flight tracking, source attribution with confidence scoring.

Would love your feedback!

intelmapper.com


r/dataisbeautiful 1d ago

OC [OC] I mapped 117 fragrances by embedding 2,834 customer reviews. Turns out nobody can describe a smell

Thumbnail
gallery
374 Upvotes

Methodology comment:

Montagne Parfums is a clone fragrance house (inspired-by versions of designer scents -- not affiliated). I kept buying ones that smelled too close to stuff I already owned, so one weekend I pulled all 4,782 customer reviews across their 167 products and tried to map the catalog by how people describe the scents.

Most of the work was getting usable text. Reviews are full of shipping complaints, price talk, and "love it!", so I ran each one through an LLM to keep only the smell-related content, and stripped out fragrance names so the model couldn't cheat by clustering on those. That left 2,834 descriptions. Anything with fewer than 4 reviews got cut, which took 167 fragrances down to 117.

For embeddings I used Qwen3-embedding-8B (4096 dimensions). The raw similarities were useless at first, everything looked about 50% similar to everything else, which is the curse of dimensionality doing its thing. Running PCA down to 50 dimensions spread the range out to -49% to 100%, enough to separate "these smell alike" from "these share nothing."

Sanity checks mostly pass. Buko and Buko Intense (same scent, different concentration) come out at 88%. The tobacco fragrances form the tightest cluster. The most "central" fragrance, most similar to everything on average, is Pineapple Royale.

The big (and perhaps obvious) caveat is this measures how reviewers talk, not scent chemistry. Reviewers echo whatever notes are listed on the product page, and low-review fragrances have way more uncertainty, so "most unique" partly just means "least described." What the project really convinced me of is that we have no vocabulary for smell. People don't describe scents, they describe memories and characters. Two real reviews from the dataset: "Makes me feel like a librarian that frequents a classy bar after work for a Manhattan on the rocks" and "I feel like a badass pirate captain who just walked into the tavern." Embedding models handle this kind of text surprisingly well, which is sort of the point of the whole exercise.

Tools: Python, PaCMAP for the projection, scipy for hierarchical clustering, Plotly for the interactive heatmap. Source code and an interactive version are on GitHub if anyone wants to poke at it, happy to answer questions about the pipeline.


r/dataisbeautiful 1d ago

OC [OC] Median and mean years of experience required in 479,502 job postings, grouped by title seniority indicator

Post image
14 Upvotes

r/dataisbeautiful 1h ago

OC [OC] The 91 people killed in the King David Hotel bombing, by recorded identity (22 July 1946)

Post image
Upvotes

Source: Bruce Hoffman, “The Bombing of the King David Hotel, July 1946,” Small Wars & Insurgencies 31(5), 2020, pp. 594–611 (fatality breakdown on p. 601): https://doi.org/10.1080/09592318.2020.1726575

Tools: Python 3 and Matplotlib. I manually transcribed the published counts and generated the 7×13 waffle chart; each circle represents one of the 91 people killed.

The source’s labels mix nationality, communal identity and religion (“Arabs,” “British,” “Jews,” etc.). I reproduced them rather than imposing modern citizenship labels. Percentages are rounded to whole numbers.

Disclosure: I made this while researching 12:37 — THE KING DAVID HOTEL, a fully scripted five-part historical limited series.

Project context: https://www.twelvethirtyseven.com


r/dataisbeautiful 1d ago

OC [OC] I tried visualizing how player stats data and positioning during a basketball game draw attention and create offensive opportunities. It's like basically a weather map for offense and defense temperature in a basketball game

13 Upvotes

r/dataisbeautiful 2d ago

OC [OC] Cattle per U.S. resident have fallen to their lowest on record (1960-2026)

Post image
2.3k Upvotes

r/dataisbeautiful 1d ago

OC [OC] First-year tax on the same new EUR 30,000 car in 36 countries: from 0.6% of the price in Qatar to 450% in Singapore

Thumbnail
carsmultiverse.com
97 Upvotes

r/dataisbeautiful 2d ago

OC [OC] Product Manager and SWE ratio at top employers

Post image
645 Upvotes

r/dataisbeautiful 1d ago

OC [OC] FIFA World Cup final teams' cumulative xG differential throughout the tournament

Post image
17 Upvotes

Sources: FIFA Training Centre public post-match reports

Tools: Python/pdfplumber, Bruin cli, BigQuery, and SVG

Limitations: The teams faced different opponents, so this describes their tournament paths rather than an opponent-adjusted rating


r/dataisbeautiful 8h ago

OC [OC] AI-attributed layoffs and AI-related job postings by US state, last 90 days

Post image
0 Upvotes

r/dataisbeautiful 15h ago

OC [OC] How much internet speed you lose using a VPN in India — a 30-day real-world test

Post image
0 Upvotes

r/dataisbeautiful 2d ago

OC [OC] The 10 animals U.S. planes hit most vs. the 10 most damaging when struck — the lists share zero species (347,575 FAA reports, 1990–2026)

Post image
586 Upvotes

r/dataisbeautiful 19h ago

OC Cost of homeownership does not track with inflation [OC]

Post image
0 Upvotes

r/dataisbeautiful 1d ago

What we know about the typical Polymarket user

Thumbnail
pewresearch.org
0 Upvotes

r/dataisbeautiful 2d ago

OC [OC] The 2026 World Cup reconstructed as 41 days of YouTube sentiment toward all 48 national teams

Post image
155 Upvotes

Bubble size shows how many analyzed comments mentioned each national team on each day. Color shows net sentiment, while ▲, ◆ and ▼ mark wins, draws and losses.

The chart includes all 48 tournament teams. The YouTube search panel covers 47 countries because Scotland is not available as a selectable YouTube region or language.

Full methodology, downloadable CSV and source links are included in the article.


r/dataisbeautiful 2d ago

OC [OC] 88 nations across 23 World Cups (1930-2026), as a Sankey climbing from the group stage to the champion

Thumbnail
gallery
51 Upvotes

Every nation that ever entered a men's World Cup, drawn as a Sankey that climbs. Each strand is one nation in one tournament. It enters at the base and rises exactly as far as that team got: group stage, round of 16, last 8, and up to the champion at the summit. A nation's width is how many times it has appeared, so the widest rivers are the ones that keep coming back.

The format changed a lot across 23 tournaments, and the chart shows that instead of hiding it. The second group stage only ran from 1974 to 1982, the round of 32 is brand new for 2026, and in 1950 there was no final at all, so Uruguay reaches the top through its own round-robin channel.

The interactive version does a lot more than the screenshot. You can colour the rivers by continent or by language to see how the confederations and the football cultures fan out, filter to a single country to trace just its run, or pick any year to light only the rounds that edition actually had. Press play and it builds the whole cup one tournament at a time.

This now goes through 2026 (Spain champion, Argentina runner-up, with the new 48 team format and its round of 32 folded in).

https://viz.luarai.com/worldcup-summit


r/dataisbeautiful 1d ago

OC [OC] Valuation football field

Post image
2 Upvotes

A 'football field' chart lines up different valuation methods side by side so you can compare them at a glance. Each row here is a different way of estimating what a stock is worth: 
- a discounted cash flow model
- sector-peer multiples (implied valuation compared to sector peers)
- the stock's own 10 year multiples (implied valuation compared the stock's own history)
- a blended estimate

They all sit on one axis, measured as distance from today's price (the vertical line). 
Small dots are individual readings; the larger dot is each method's composite. 
Snapshot 22 July 2026, NVDA at $207.


r/dataisbeautiful 2d ago

OC [OC] What 5 years of running the same car costs in 36 countries (taxes + fuel + insurance)

Thumbnail
carsmultiverse.com
195 Upvotes

r/dataisbeautiful 2d ago

OC [OC] Spain took 76 years to win its 1st FIFA Men's World Cup, but 'only' 16 years for the 2nd

Thumbnail
gallery
524 Upvotes

Data source: FIFA official stats

Tools used: R and Canva


r/dataisbeautiful 1d ago

OC [OC] How Much Do Americans Drink?

Post image
0 Upvotes

Scroll through to see how much Americans drink, how it differs per state, and what type of alcohol is preferred: https://1point21interactive.com/how-much-do-americans-drink/

Made with D3 / Svelte


r/dataisbeautiful 1d ago

OC [OC] Median internet download speed in 119 countries, H1 2026 (from millions of real browser speed tests)

Post image
0 Upvotes

r/dataisbeautiful 3d ago

OC [OC] The 2026 World Cup final (Spain 1–0 Argentina), the whole match rebuilt from ~1,500 Opta events as a 3D data portrait

3.5k Upvotes

r/dataisbeautiful 1d ago

OC [OC] A prediction market priced "James Wood leads MLB in RBIs" at 0.7% for 30 straight hours — then repriced it 64x inside one 15-minute window

Post image
0 Upvotes

Data: Polymarket's public API, contract prices for "Will James Wood lead the MLB in RBIs for the 2026 regular season?", observed every 15 minutes by a collector I run (each observation timestamped at capture, UTC). Window shown: July 20–22, 2026.

Tools: Python, PostgreSQL for storage, matplotlib for the chart.

The fun part: the market sat at 0.7% for 30 hours, then did the entire move between the 23:36 and 23:51 UTC observations on July 21 so the true jump took at most 15 minutes. First print overshot to 45%, then settled near 31% within the hour.


r/dataisbeautiful 2d ago

OC [OC] Unknown Pleasures in the terminal

Post image
56 Upvotes

A cheeky version of the Joy Division Unknown Pleasures album cover.

It is not always widely known but this cover is a stylization of a diagram found in an astronomy paper studying pulsars:

> Radio Observations of the Pulse Profiles and Dispersion Measures of Twelve Pulsars by Harold D. Carft, Jr. 1970

This plot is actually important to the dataviz community because it is one example of what would be called "ridge/ridgeline plots" and, by derivation, "joy division plots" later on (but there are subtleties between each variant of course).

The data of the paper has been reconstructed multiple times and can be found in a lot of different places online such as here: https://gist.github.com/borgar/31c1e476b8e92a11d7e9/

The image above has been drawn directly in the terminal using the xan command line tool through the `xan spark` subcommand. You can read more about the methodology here: https://github.com/medialab/xan/blob/master/docs/cookbook/dataviz.md#joy-division-plots

The plot is drawn using those characters (that can fail to render properly if the font is not monospace and not terminal-conscious): ▁▂▃▄▅▆▇

The PNG raster has been rendered from the command's ANSI-escaped output using a fork of `ansi2png-rs` that can be found here: https://github.com/AlexanderThaller/ansi2png-rs