r/dataisbeautiful 8h ago

OC [OC] 40% of Americans read zero books last year. The top 10% account for half of all the books read in the country.

Post image
2.8k Upvotes

Last year, 4 out of every 10 grown ups in America didn't finish a single book. But a small group, about 1 in 10 people, read 25 or more books. That small group read so much that they probably account for about half of all the books read in the whole country. So basically, a few people are doing almost all the reading, while most people are barely reading at all.

EDIT - This includes audio books. The survey question was "Have you read or listened to ..."?

Also a graph for another day - most people who did read, read Mystery/Crime.


r/dataisbeautiful 19h ago

OC [oc] how many times i cried in a month before/after getting on lexapro

Thumbnail
gallery
2.1k Upvotes

posted months ago my crying stats from the first quarter of the year and i mentioned in that that i was tracking the data because i was planning on getting medicated over the summer and someone mentioned they wanted a before and after so here that is lol. no need to worry any longer, i am fixed!

source: my life

tools for visualization: canva (slides 1-2) + i tracked the data in my calendar app (slides 3-whatever, since people seemed to enjoy seeing my reasons last time)


r/dataisbeautiful 7h ago

[OC] The price curve for items sold at government auction in the US.

Post image
642 Upvotes

Some recent fun data from my government auction platform: the distribution of winning bids across >100k listings from the last 6 months or so. X axis is a log scale of bidding price, and Y axis is the number of listings sold at that price.

People love to bid (and win!) things for $10 - you can get a lot of stuff for that much, like a classroom full of desks and furniture. The expensive part is transporting it all.

The median listing is a 1997 Ford Ranger that went for $210 - vehicles are by far the most popular thing that people bid on, at a wide range of prices and conditions.

There's a link below to the interactive chart and a bit more context on the data - you can filter by category to see how the curve changes for things like vehicles or Real Estate.

Source: The Govauctions.app auction database, including listings from GSA, GovDeals, PublicSurplus and more.

Tools: SVG and Javascript

Interactive version of the chart, where you can sort by category and see listings at a given bid range: https://govauctions.app/research/what-the-government-sells-for


r/dataisbeautiful 6h ago

OC [OC] Route 66 turns 100 this year. I mapped every business still open on it, then made all 2,199 miles drivable in one scroll

225 Upvotes

r/dataisbeautiful 12h ago

OC Spread of simulated League of Legends draft win probability by rank, from 400,000 ten-champion lobbies per rank [OC]

Post image
128 Upvotes

Public champion tier lists, all 11 rank filters x 5 roles, patch 16.17.1, snapshotted 2026-08-29. 2,765 champion/role/rank entries.

Tools: I built a lobby advisor ( https://shouldidodge.lol ) and this came out of calibrating it. C# to collect the daily snapshot, Python + NumPy for the simulation, matplotlib for the chart.


r/dataisbeautiful 7h ago

OC [OC] With almost a month until the first round of Brazilian Presidential Election, Lula's nationwide approval rating is of 45,4%

Post image
79 Upvotes

r/dataisbeautiful 8h ago

OC [OC] September and October are the best times to look for work. Everyone waits for January to job hunt, but January had the worst interview rate of 2025. I analyzed 1.99 million job applications and found that interview rates and total interviews were highest in October.

Post image
79 Upvotes

r/dataisbeautiful 7h ago

Money spent on building data centers in the US has grown 5-fold since late 2022

Thumbnail
ourworldindata.org
71 Upvotes

r/dataisbeautiful 3h ago

[OC] Ranges of european, amerindian and subsaharian african ancestry in Latin America

Thumbnail
gallery
54 Upvotes

Own Work. Sources at the Bottom.

Own Work. Sources: *Norris et al. (2018), PubMed Central. (https://pubmed.ncbi.nlm.nih.gov/30537949/) for Colombia/Mexico/Peru* Borda et al. (2024), ScienceDirect. (https://www.sciencedirect.com/science/article/pii/S2666979X24003215) for Chile/Peru/Brazil/Uruguay/Mexico/Guatemala/Costa Rica/Dom Rep *Ruiz-Linares et al. (2014), PubMed Central. (https://pmc.ncbi.nlm.nih.gov/articles/PMC4177621/) for Colombia/Mexico/Peru/Brazil *Horimoto et al. (2021), KI Reports (https://www.kireports.org/article/S2468-0249(21)00593-3/fulltext) for central america *Homburger et at. (2015), PubMed Central (https://pubmed.ncbi.nlm.nih.gov/26636962/) for South America *Moreno-Estrada et al. (2013), PubMed. (https://pubmed.ncbi.nlm.nih.gov/24244192/) for The Caribbean. *Johnson et al. (2011), PubMed (https://pmc.ncbi.nlm.nih.gov/articles/PMC3240599/) for Mexico *Suarez-Curtz. (2010), ResearchGate (https://www.researchgate.net/publication/51563335_Pharmacogenetics_in_the_Brazilian_Population) for Brazil *Eyheramendy et at. (2015), Nature Communications (https://www.nature.com/articles/ncomms7472) for Chile *Avena et al. (2012) ResearchGate (https://www.nature.com/articles/ncomms7472) for Argentina *Asgari et al. (2020), ResearchGate. (https://www.nature.com/articles/ncomms7472) for Peru *Ibarra et al. (2014), Plos Journals (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0087202) for Colombia *Söchtig et al. (2015), ResearchGate (https://www.researchgate.net/publication/273703459_Genomic_insights_on_the_ethno-history_of_the_Maya_and_the_'Ladinos'_from_Guatemala) for Guatemala *Castro-Pérez et al. (2016), ISPUB (https://ispub.com/IJBA/9/1/44045) for Panama


r/dataisbeautiful 6h ago

OC [oc] Best NFL Front Offices Drafting Talent (2016 - 2025) *Updated*

Post image
14 Upvotes

Rank Position by Score, 1-32

Score Headline index. Draft z-score and undrafted z-score mixed 79 / 21, re-standardized, displayed as 50 + 16z. 50 = league average, higher is better

Draft The draft-only score, same scale

Surplus/Pick Career wAV above slot expectation per pick, standardized within draft class, restated in AV-equivalent units. +6.7 ≈ each pick returned 6.7 AV more than its slot should yield

Star% Share of picks finishing in the top 10% of their own draft class by career wAV

Bust% Share of rounds 1-3 picks finishing ≥0.5 class SD below slot expectation. Lower is better (shading inverted)

Best pick Highest career-wAV draft pick of the window (year and overall pick)

Top-Third% Share of picks finishing in the top 33% of their own draft class

Day 3 Surplus Same surplus measure, rounds 4-7 only the late-round scouting signal

All-Pro First-team All-Pro selections by the team's picks (PFR's field; includes specialists). Displayed, not scored

Pro Bowl Pro Bowl selections by the team's picks. Displayed, not scored

Kept% dr_av ÷ w_av share of the career value a team drafted that was produced for that team. Retention, not selection. Displayed, not scored

2026 picks Picks made in the April 2026 draft. Capital only no games played, so ungraded


r/dataisbeautiful 11h ago

OC Work, School, or Both? What Americans Ages 18–24 Are Doing [OC]

Post image
17 Upvotes

r/dataisbeautiful 10h ago

OC [OC] Interactive Dashboard of Official Spanish Car Registrations & EV Market Share.

Thumbnail
cardatasales.com
9 Upvotes

Interactive live dashboard.

Key insights from August 2026 (69,365 new passenger cars):

• Pure Electric (BEV) share: 13.6% (9,450 units)

• Electrified total (BEV + PHEV): 27.6%

• Hybrid Gasoline (HEV): 43.4%

• Pure Gasoline: 18.2% | Diesel: 3.8% | Gas (LPG/CNG): 4.6%

Top EV models in August:

  1. Kia EV3 (433 un.)

  2. BYD Atto 2 (373 un.)

  3. BYD Dolphin Surf (343 un.)

  4. Leapmotor B10 (332 un.)

  5. Citroën ë-C3 (302 un.)


r/dataisbeautiful 9h ago

OC [oc] Cross Country National Championships

6 Upvotes

The 2025 Nationals course modeled medians 47 seconds faster for women and 73 seconds faster for men after standardization.


r/dataisbeautiful 8h ago

OC [OC] All 180 named figures in the Odyssey, mapped from the text with every (288) relationship between them sourced to a quote

Post image
6 Upvotes

r/dataisbeautiful 5h ago

OC [OC] Daily cloud cover over India, analyzed from ~1,200 weather station meteograms, Feb-Sep 2026

Post image
5 Upvotes

r/dataisbeautiful 3h ago

OC [OC] What languages are programming languages built with? A graph of compiler and runtime lineage

Thumbnail
languagelineage.org
4 Upvotes

[OC] What programming languages are programming languages actually built with? An interactive lineage graph

I’ve been building an interactive graph that maps implementation relationships between programming languages, compilers, and runtimes.

For example:

• Go’s compiler moved from C to Go around Go 1.5
• modern Rust uses staged self-hosting
• CPython is primarily implemented in C
• HotSpot is largely C++
• javac is implemented in Java
• V8 is primarily C++

I wanted to visualize something a little different from the usual “which language influenced which language?” trees.

The graph focuses on implementation lineage:

compiler_written_in
runtime_written_in
bootstrap_written_in

Each relationship has a time range, confidence score, and evidence source rather than treating every historical relationship as equally certain.

Interactive version:
https://www.languagelineage.org/explore

I’m expanding the dataset now. I’d especially like feedback on questionable relationships, missing historical transitions, or better primary sources.


r/dataisbeautiful 10h ago

Discussion [Topic][Open] Open Discussion Thread — Anybody can post a general visualization question or start a fresh discussion!

3 Upvotes

Anybody can post a question related to data visualization or discussion in the monthly topical threads. Meta questions are fine too, but if you want a more direct line to the mods, click here

If you have a general question you need answered, or a discussion you'd like to start, feel free to make a top-level comment.

Beginners are encouraged to ask basic questions, so please be patient responding to people who might not know as much as yourself.


To view all Open Discussion threads, click here.

To view all topical threads, click here.

Want to suggest a topic? Click here.


r/dataisbeautiful 5h ago

OC [OC] mappd - was looking for decent places to move to in US filled in a lot of data on places to NOT move to and places that are turning it around

Thumbnail
rabmach.github.io
0 Upvotes

As the subject reads - looking for a place within US to maybe move to. Came up with a ton of data points on why not to move to a certain spot and some data detailing how spots (regions, states) are turning it around. If you highlight a button to see data and your IP is within 100 miles of that data point opportunity exists to send off a quick note to a rep.

Plenty of tags to filter and drill down into. mappd


r/dataisbeautiful 4h ago

OC [OC] How exposed is the European job market to AI? Visualizing 436 standard occupations across all 27 EU countries with real Eurostat census data and Gemini AI scoring

Post image
0 Upvotes

Live Interactive Tool: https://eu-jobs.alexandrucruceanu.com
Source Code (GitHub): https://github.com/alexandrucruceanu/EU-jobs

Data & Methodology:

  • Employment & Wage Data: Official Eurostat census database (2023 Eurostat Structural Earnings Survey & Labour Force Survey).
  • Taxonomy: 436 ISCO-08 4-digit occupations mapped to the European Skills, Competences, Qualifications and Occupations (ESCO) framework.
  • AI Exposure Scores (0–10): Generated using Gemini by evaluating routine digital knowledge processing vs physical/interpersonal dexterity tasks from ESCO occupational profiles.
  • Tools Used: Python for ETL pipelines, Vanilla JavaScript, HTML5 Canvas 2D for squarified treemaps and scatter matrix, Docker + Nginx.

Key Observations:

  1. High-earning roles (finance, law, software) exhibit the highest digital AI exposure (7–9/10), but benefit significantly from productivity augmentation.
  2. Hands-on skilled trades (plumbing, electricians, healthcare nursing) remain heavily shielded against near-term AI disruption (0–2/10).

r/dataisbeautiful 5h ago

OC [OC] I watched a language model think. Here's what its mind looks like as it processes a question, layer by layer

0 Upvotes

Last week I posted a topographic map of Qwen 2.5-7B's vocabulary, 10,000+ English words projected from 3,584-dimensional space into a terrain where similar words form mountains. That map was static, a snapshot of the model before it thinks.

This time I made it think.

I fed the model a question: "What material would be hardest for a craftsman to combine with gold using only fire: quartz, silver, or copper? Answer in one word." Then captured its internal state at every one of its 28 transformer layers. Then I projected each layer's hidden state back onto the vocabulary terrain and visualized which regions light up.

Models think across layers, which means that each layer transforms the model's internal representation a little further. Think of it like a 28-step thought process. The input enters as raw words, no understanding. Each layer adds context, combines meanings, and resolves ambiguities. By the final layer, the representation has been refined into a specific answer.

The challenge is that we can't normally see what's happening in the middle steps. The model's thinking at layer 14 exists as a 3,584-dimensional vector that isn't directly readable as words. It's like watching someone solve a math problem but only being able to see the final answer, the scratch work is in a language you can't read. The J-lens translates that scratch work back into vocabulary at every step.

What you're seeing in the GIF:

The grey mountains are the model's vocabulary landscape, fixed and unchanging. The amber glow shows which words the model's thinking is connected to at each layer depth. Bright means "this region is active." Dark means "the model isn't thinking about this right now."

White labels mark the strongest activations. Amber labels (⬡) mark words that are uniquely active at that specific layer compared to all others.

The story across layers:

Layers 0–5: The model is parsing basic language structure. Top words are "one", "how", "were", they are syntax, not content.

Layers 6–12: Task recognition. "Strategist", "answer", "puzzle" appear. The model has figured out it's being asked to solve something, but hasn't started thinking about materials yet.

Layers 13–22: "Answer" dominates for ten straight layers. This is the model's workspace, it's holding "I need to produce an answer" while working through the decision. The distinctive words evolve alongside it: "correct" and "decision" appear at layer 19, "single" at layer 20. It knows it needs one correct answer.

Layer 23: The pivot. "Copper" appears for the first time. The model transitions from processing the task format to processing the content.

Layer 28: Crystallization. The glow concentrates on a single cluster: copper, silver, quartz, silica, gold. The distinctive words at this layer are "hard", "solid", "specific", the actual reasoning about material properties. The model answers: copper (55% confidence), silver (14%), quartz (7%).

The model's final reply:

Copper. Copper would be the hardest material for a craftsman to combine with gold using only fire due to its lower melting point and tendency to form an oxide layer when heated, which can make it difficult to fuse with gold. In contrast, silver and quartz have higher melting points and are more compatible with gold when melted together. However, among the given options, copper is the most challenging to work with in this context. 

Qwen got it wrong:

The model's inner deliberation showed "answer, correct, decision" in the mid-layers, it weighed all three options (copper 55%, silver 14%, quartz 7%), committed to copper with confidence, and got it wrong. The visualization of thinking doesn't guarantee the thinking is correct. The map shows you how the model reaches its answer, not whether the answer is right. I asked for one word, and Qwen couldn't help itself and answered a whole paragraph.

What we're still missing:

Most of the visible action happens in the last 4 layers, where the terrain suddenly erupts with activation before collapsing onto the final answer. There's clearly a lot happening in that transition. The model goes from "I know what kind of question this is" to "the answer is copper" in just a few layers, and right now we're seeing the explosion but not the detail of how it resolves. Zooming into that window, maybe tracking individual word trajectories frame by frame, is the next step.

Also worth noting: the J-lens covers layers 0–26, and the final layers fall back to a cruder method. So the most dramatic moment in the model's thinking is the one we're least equipped to read. Working on it.

Next up: the hunt for the em-dash.

Certain models have strong stylistic preferences — they love em-dashes. Somewhere in those 3,584 dimensions there's a direction that means "dramatic pause energy." I want to find it.

The method:

This uses the J-lens (from "Verbalizable Representations Form a Global Workspace in Language Models," 2026). At each transformer layer, the model's hidden state exists in a coordinate system that's been rotated by all previous layers. A naive projection back to vocabulary (called the logit lens) produces noise in the middle layers because it's reading a rotated map. The J-lens corrects for that rotation using a learned per-layer transformation (the averaged Jacobian of the network from that layer to the output), making the mid-layer thinking readable.

I ran both lenses on the same prompt. I don't show the logit lens here, but I have a similar visual for it. The logit lens showed the word "libertine" as the top activation for 22 out of 29 layers, pure geometric noise. The J-lens showed "answer", "puzzle", "hint", "true" , the actual deliberation.

Tools: Qwen2.5-7B-Instruct, PyTorch (hidden state capture), pre-fitted J-lens from HuggingFace, UMAP + KDE for the terrain, Plotly for the 3D visualization.

Visualization and research: me. Write-up polished with Claude's help.

Previous post: https://www.reddit.com/r/dataisbeautiful/comments/1vyq0ug/oc_i_turned_qwen_257bs_embedding_space_into_a/


r/dataisbeautiful 2h ago

We built a live dashboard tracking 14+ global health indicators — CO₂, GDP, debt, poverty, and more — with AI‑powered interpretation.

Thumbnail
global-vitals-dashboard.traffictorch.workers.dev
0 Upvotes

Global Vital Signs Dashboard

14 interactive charts covering:

· Climate: CO₂, temperature anomaly
· Economy: GDP, growth, fiscal balance, inflation, unemployment, tax revenue
· Social: poverty, infant mortality, population
· Productivity: GDP per capita, labour productivity

All data from World Bank, NASA, and Global Warming API. Updates automatically.

AI Feature: Click "Interpret Current Data" and get a concise plain‑English summary of the latest numbers — climate, economy, and social conditions in one flowing paragraph.

Built with Cloudflare Workers + D1 + Workers AI. Open source, free, no ads.

Would love feedback on the visualisation and the AI output.


r/dataisbeautiful 2h ago

Data

Thumbnail
gallery
0 Upvotes

#data


r/dataisbeautiful 14h ago

OC [OC] The whole world is shifting right: how 106 countries' wind + solar electricity share moved, 2010–2025

Post image
0 Upvotes

Each ridge is one year. The curve shows the distribution of 106 countries by the share of their

electricity that comes from wind + solar; the dot marks the median. In 2010 almost every country

was near 0% (median 0.1%). By 2025 the entire distribution has slid right — median 13.4%, with a

long tail out to Denmark at 72%. What struck me is that it's not a handful of leaders pulling the

average; the whole mass moves.

Data: Ember (yearly electricity data) + Energy Institute. Tool: R — ggplot2, kernel density per year.

Cross-checked against Ember's country pages (e.g. China 1,347 TWh in 2000 → 10,573 TWh in 2025).


r/dataisbeautiful 5h ago

Data

Thumbnail
gallery
0 Upvotes

Read


r/dataisbeautiful 6h ago

OC [OC] Time to reach each rank in nine martial arts

Post image
0 Upvotes

Hung nine martial arts on one wall as belts, to scale in years. The wooden rail at the top is the door you walk in through, and every rank hangs below it at the earliest year each federation's own published minimums allow. Where several ranks share a color, the dashed lines mark the divisions.

Two surprises. First, every art on the wall places its black rank within seven years of the door. Kendo's shodan can arrive in six months, and BJJ is the slowest at 6.5 years. What federations actually regulate in detail is everything below black.

Second, the black belt is nowhere near the end of the ladder. Judo's fastest legal route to 10th dan takes 58 years. The JKA reaches 10th around year 47 and marks nothing on the belt the entire way: fifty years of rank, zero stripes.

One honest gap: where a belt fades into washed cloth, the federation publishes no time at all. Colored ranks in judo, Shotokan and taekwondo are set by each school, so those stretches show no divisions instead of a guess. That gap turned out to be the finding.

And a reality check from the mats: BJJ's rulebook floor is 4.5 years to black, but a survey of 1,948 practitioners puts the real one around year 16.

Interactive version: https://viz.luarai.com/belt-ladder

EDIT: A few people are reading the red lines as typical times, so to be clear: they mark the earliest year the written rules allow, not the average journey. For taekwondo and judo the colored years aren't even regulated (that's the washed cloth), so those placements are drawn, and judo rides the elite competitor route. On "the data isn't hard to come by": I looked hard. No federation publishes time-to-rank statistics (the Kodokan publishes none at all), and the one academic study (n=285) sampled champions only. BJJ is the sole art with real survey data, which is why it's the only belt wearing a reality line at year 16. If anyone knows a dataset for the others, genuinely, send it and I'll add the reality lines.