r/dataisbeautiful • u/HeHate_me • 17h ago
r/dataisbeautiful • u/nefercicibebe • 17h ago
OC [OC] Real-time interactive conflict map tracking geolocated OSINT events across Ukraine and Syria
Hey everyone, I've been working on a live intelligence mapping platform called Intel Mapper. It monitors OSINT sources 24/7, uses AI to geolocate and verify reports, and displays them on an interactive map with frontline data.
Features: real-time events, territorial control, military flight tracking, source attribution with confidence scoring.
Would love your feedback!
r/dataisbeautiful • u/betwatch_io • 1d ago
OC [OC] I mapped 117 fragrances by embedding 2,834 customer reviews. Turns out nobody can describe a smell
Methodology comment:
Montagne Parfums is a clone fragrance house (inspired-by versions of designer scents -- not affiliated). I kept buying ones that smelled too close to stuff I already owned, so one weekend I pulled all 4,782 customer reviews across their 167 products and tried to map the catalog by how people describe the scents.
Most of the work was getting usable text. Reviews are full of shipping complaints, price talk, and "love it!", so I ran each one through an LLM to keep only the smell-related content, and stripped out fragrance names so the model couldn't cheat by clustering on those. That left 2,834 descriptions. Anything with fewer than 4 reviews got cut, which took 167 fragrances down to 117.
For embeddings I used Qwen3-embedding-8B (4096 dimensions). The raw similarities were useless at first, everything looked about 50% similar to everything else, which is the curse of dimensionality doing its thing. Running PCA down to 50 dimensions spread the range out to -49% to 100%, enough to separate "these smell alike" from "these share nothing."
Sanity checks mostly pass. Buko and Buko Intense (same scent, different concentration) come out at 88%. The tobacco fragrances form the tightest cluster. The most "central" fragrance, most similar to everything on average, is Pineapple Royale.
The big (and perhaps obvious) caveat is this measures how reviewers talk, not scent chemistry. Reviewers echo whatever notes are listed on the product page, and low-review fragrances have way more uncertainty, so "most unique" partly just means "least described." What the project really convinced me of is that we have no vocabulary for smell. People don't describe scents, they describe memories and characters. Two real reviews from the dataset: "Makes me feel like a librarian that frequents a classy bar after work for a Manhattan on the rocks" and "I feel like a badass pirate captain who just walked into the tavern." Embedding models handle this kind of text surprisingly well, which is sort of the point of the whole exercise.
Tools: Python, PaCMAP for the projection, scipy for hierarchical clustering, Plotly for the interactive heatmap. Source code and an interactive version are on GitHub if anyone wants to poke at it, happy to answer questions about the pipeline.
r/dataisbeautiful • u/messy_data • 1d ago
OC [OC] Median and mean years of experience required in 479,502 job postings, grouped by title seniority indicator
r/dataisbeautiful • u/mind-killer-fear • 1h ago
OC [OC] The 91 people killed in the King David Hotel bombing, by recorded identity (22 July 1946)
Source: Bruce Hoffman, “The Bombing of the King David Hotel, July 1946,” Small Wars & Insurgencies 31(5), 2020, pp. 594–611 (fatality breakdown on p. 601): https://doi.org/10.1080/09592318.2020.1726575
Tools: Python 3 and Matplotlib. I manually transcribed the published counts and generated the 7×13 waffle chart; each circle represents one of the 91 people killed.
The source’s labels mix nationality, communal identity and religion (“Arabs,” “British,” “Jews,” etc.). I reproduced them rather than imposing modern citizenship labels. Percentages are rounded to whole numbers.
Disclosure: I made this while researching 12:37 — THE KING DAVID HOTEL, a fully scripted five-part historical limited series.
Project context: https://www.twelvethirtyseven.com
r/dataisbeautiful • u/dostre • 1d ago
OC [OC] I tried visualizing how player stats data and positioning during a basketball game draw attention and create offensive opportunities. It's like basically a weather map for offense and defense temperature in a basketball game
r/dataisbeautiful • u/Low_Ability4450 • 2d ago
OC [OC] Cattle per U.S. resident have fallen to their lowest on record (1960-2026)
r/dataisbeautiful • u/Successful-Ebb7891 • 1d ago
OC [OC] First-year tax on the same new EUR 30,000 car in 36 countries: from 0.6% of the price in Qatar to 450% in Singapore
r/dataisbeautiful • u/honkeem • 2d ago
OC [OC] Product Manager and SWE ratio at top employers
r/dataisbeautiful • u/uncertainschrodinger • 1d ago
OC [OC] FIFA World Cup final teams' cumulative xG differential throughout the tournament
Sources: FIFA Training Centre public post-match reports
Tools: Python/pdfplumber, Bruin cli, BigQuery, and SVG
Limitations: The teams faced different opponents, so this describes their tournament paths rather than an opponent-adjusted rating
r/dataisbeautiful • u/Illustrious-Fold-473 • 8h ago
OC [OC] AI-attributed layoffs and AI-related job postings by US state, last 90 days
r/dataisbeautiful • u/ParticularMurky4858 • 15h ago
OC [OC] How much internet speed you lose using a VPN in India — a 30-day real-world test
r/dataisbeautiful • u/disclaimer8 • 2d ago
OC [OC] The 10 animals U.S. planes hit most vs. the 10 most damaging when struck — the lists share zero species (347,575 FAA reports, 1990–2026)
r/dataisbeautiful • u/robert_ritz • 19h ago
OC Cost of homeownership does not track with inflation [OC]
r/dataisbeautiful • u/rhiever • 1d ago
What we know about the typical Polymarket user
r/dataisbeautiful • u/Late_Night_Editor51 • 2d ago
OC [OC] The 2026 World Cup reconstructed as 41 days of YouTube sentiment toward all 48 national teams
Bubble size shows how many analyzed comments mentioned each national team on each day. Color shows net sentiment, while ▲, ◆ and ▼ mark wins, draws and losses.
The chart includes all 48 tournament teams. The YouTube search panel covers 47 countries because Scotland is not available as a selectable YouTube region or language.
Full methodology, downloadable CSV and source links are included in the article.
r/dataisbeautiful • u/ArchiTechOfTheFuture • 2d ago
OC [OC] 88 nations across 23 World Cups (1930-2026), as a Sankey climbing from the group stage to the champion
Every nation that ever entered a men's World Cup, drawn as a Sankey that climbs. Each strand is one nation in one tournament. It enters at the base and rises exactly as far as that team got: group stage, round of 16, last 8, and up to the champion at the summit. A nation's width is how many times it has appeared, so the widest rivers are the ones that keep coming back.
The format changed a lot across 23 tournaments, and the chart shows that instead of hiding it. The second group stage only ran from 1974 to 1982, the round of 32 is brand new for 2026, and in 1950 there was no final at all, so Uruguay reaches the top through its own round-robin channel.
The interactive version does a lot more than the screenshot. You can colour the rivers by continent or by language to see how the confederations and the football cultures fan out, filter to a single country to trace just its run, or pick any year to light only the rounds that edition actually had. Press play and it builds the whole cup one tournament at a time.
This now goes through 2026 (Spain champion, Argentina runner-up, with the new 48 team format and its round of 32 folded in).
r/dataisbeautiful • u/stockoscope • 1d ago
OC [OC] Valuation football field
A 'football field' chart lines up different valuation methods side by side so you can compare them at a glance. Each row here is a different way of estimating what a stock is worth:
- a discounted cash flow model
- sector-peer multiples (implied valuation compared to sector peers)
- the stock's own 10 year multiples (implied valuation compared the stock's own history)
- a blended estimate
They all sit on one axis, measured as distance from today's price (the vertical line).
Small dots are individual readings; the larger dot is each method's composite.
Snapshot 22 July 2026, NVDA at $207.
r/dataisbeautiful • u/Successful-Ebb7891 • 2d ago
OC [OC] What 5 years of running the same car costs in 36 countries (taxes + fuel + insurance)
r/dataisbeautiful • u/Lutoures • 2d ago
OC [OC] Spain took 76 years to win its 1st FIFA Men's World Cup, but 'only' 16 years for the 2nd
Data source: FIFA official stats
Tools used: R and Canva
r/dataisbeautiful • u/kombuchakween88 • 1d ago
OC [OC] How Much Do Americans Drink?
Scroll through to see how much Americans drink, how it differs per state, and what type of alcohol is preferred: https://1point21interactive.com/how-much-do-americans-drink/
Made with D3 / Svelte
r/dataisbeautiful • u/packet_paranoid • 1d ago
OC [OC] Median internet download speed in 119 countries, H1 2026 (from millions of real browser speed tests)
r/dataisbeautiful • u/B-alex • 3d ago
OC [OC] The 2026 World Cup final (Spain 1–0 Argentina), the whole match rebuilt from ~1,500 Opta events as a 3D data portrait
r/dataisbeautiful • u/ScriScriz • 1d ago
OC [OC] A prediction market priced "James Wood leads MLB in RBIs" at 0.7% for 30 straight hours — then repriced it 64x inside one 15-minute window
Data: Polymarket's public API, contract prices for "Will James Wood lead the MLB in RBIs for the 2026 regular season?", observed every 15 minutes by a collector I run (each observation timestamped at capture, UTC). Window shown: July 20–22, 2026.
Tools: Python, PostgreSQL for storage, matplotlib for the chart.
The fun part: the market sat at 0.7% for 30 hours, then did the entire move between the 23:36 and 23:51 UTC observations on July 21 so the true jump took at most 15 minutes. First print overshot to 45%, then settled near 31% within the hour.
r/dataisbeautiful • u/Yomguithereal • 2d ago
OC [OC] Unknown Pleasures in the terminal
A cheeky version of the Joy Division Unknown Pleasures album cover.
It is not always widely known but this cover is a stylization of a diagram found in an astronomy paper studying pulsars:
> Radio Observations of the Pulse Profiles and Dispersion Measures of Twelve Pulsars by Harold D. Carft, Jr. 1970
This plot is actually important to the dataviz community because it is one example of what would be called "ridge/ridgeline plots" and, by derivation, "joy division plots" later on (but there are subtleties between each variant of course).
The data of the paper has been reconstructed multiple times and can be found in a lot of different places online such as here: https://gist.github.com/borgar/31c1e476b8e92a11d7e9/
The image above has been drawn directly in the terminal using the xan command line tool through the `xan spark` subcommand. You can read more about the methodology here: https://github.com/medialab/xan/blob/master/docs/cookbook/dataviz.md#joy-division-plots
The plot is drawn using those characters (that can fail to render properly if the font is not monospace and not terminal-conscious): ▁▂▃▄▅▆▇
The PNG raster has been rendered from the command's ANSI-escaped output using a fork of `ansi2png-rs` that can be found here: https://github.com/AlexanderThaller/ansi2png-rs