Estimating global data creation
Memory Loss sums annual estimates of the global datasphere from January 1, 1999, then projects them forward second by second. It is a model of data created, captured, copied and consumed—not a live measurement, a count of unique internet content, or a count of what still exists.
01 The headline number
All visitors with the same clock see the same total. The counter is evaluated from a UTC timestamp and the configuration; elapsed browser frames never accumulate into it.
Each year contributes its configured volume. Within a year, the rate grows exponentially, normalised so that the integral equals that year’s total.
rY(τ) = AY × ln(k) × kτ/T / (T × (k − 1))
A is the annual total, k is the following year’s total divided by this year’s, T is the number of seconds in the year, and τ is the time elapsed. When k = 1, the formula becomes linear. The headline adds the completed years to the current year’s integral. Leap days are included; leap seconds are ignored.
Missing years are interpolated geometrically between known totals. After 2029, the final known annual ratio is extended forward. The cumulative total is continuous at year boundaries; the rate can change there because adjacent forecasts imply different growth curves.
| year | ZB | basis |
|---|---|---|
| 1999 | 0.008 | approximate |
| 2000 | 0.012 | approximate |
| 2001 | 0.02 | approximate |
| 2002 | 0.03 | approximate |
| 2003 | 0.05 | approximate |
| 2004 | 0.08 | approximate |
| 2005 | 0.13 | approximate |
| 2006 | 0.16 | approximate |
| 2007 | 0.28 | approximate |
| 2008 | 0.49 | approximate |
| 2009 | 0.8 | approximate |
| 2010 | 2 | approximate |
| 2011 | 5 | approximate |
| 2012 | 6.5 | approximate |
| 2013 | 9 | approximate |
| 2014 | 12.5 | approximate |
| 2015 | 15.5 | estimate |
| 2016 | 18 | estimate |
| 2017 | 26 | estimate |
| 2018 | 33 | estimate |
| 2019 | 41 | estimate |
| 2020 | 64.2 | estimate |
| 2021 | 79 | estimate |
| 2022 | 97 | estimate |
| 2023 | 120 | estimate |
| 2024 | 149 | estimate |
| 2025 | 181 | forecast |
| 2026 | 221 | forecast |
| 2027 | 295.083 | interpolated |
| 2028 | 394 | forecast |
| 2029 | 527.5 | forecast |
The 1999–2014 values are approximate and contribute 3.66% of the launch total. The earliest years are back-extrapolations, not independent measurements.
Statista series, reproduced by DemandSage ↗
Background for the 2015–2026 series. The current 2024 figure is 147 ZB; this model version uses 149 ZB.
IDC: The Digitization of the World ↗
Primary background for the global datasphere definition. This older forecast is not a primary source for the complete 2026 config.
DesignRush: daily data generation ↗
Secondary compilation for later forecasts. Forecast editions differ.
Reading the launch values
At 2026-09-19T00:00:00Z, this configuration gives 1.012156 YB (1.012156e+24 bytes) and 7.43 PB per second. One yottabyte is 1024 bytes. The unit label uses this decimal definition throughout.
The full integer is evaluated with 36-place fixed-point arithmetic and BigInt. Those digits describe the projection; their length does not imply byte-level observational accuracy.
02 The categories
Category counters begin at 2026-09-19 with a zero offset unless an offset is configured. Their byte totals use assumed average file sizes. Six newer categories use dated published activity figures held flat; their AI attribution is unknown, so they are excluded from the origin calculation. They overlap, are illustrative, and do not add up to the headline.
| category / source | items / s | annual growth | AI weight |
|---|---|---|---|
| hours of video uploaded to youtubeSocialRails upload statistics ↗ | 8.333 | +0% | 0% |
| photos & videos posted to instagramEarthWeb data statistics ↗ | 1100 | +5% | 0% |
| instagram storiesBusinesstats: content per minute ↗ | 11583 | +5% | 0% |
| snaps sentSnap figures via SocialRails ↗ | 63657 | +5% | 0% |
| whatsapp messagesBusinesstats: content per minute ↗ | 683333 | +5% | 0% |
| emails sentThe Radicati Group ↗ | 4351852 | +3% | 0% |
| messages sent to ai assistantsNBER working paper 34255 ↗ | 46300 | +100% | 100% |
| ai images generatedWearView: AI image statistics ↗ | 1389 | +100% | 100% |
| ai videos generatedAdwave: AI video generation statistics ↗ | 58 | +150% | 100% |
| songs uploaded to streamingMusic Ally / Luminate ↗ | 1.74 | +20% | 50% |
| ai songs generatedBillboard: streaming upload report ↗ | 81 | +100% | 100% |
| posts uploaded to tiktokSteel et al.: Just Another Hour on TikTok (v5, 2026) ↗ | 3116.898 | +0% | unknown |
| podcast episodes publishedListen Notes: podcast statistics dataset ↗ | 0.863 | +0% | unknown |
| wordpress blog postsWordPress.com: publishing activity ↗ | 26.618 | +0% | unknown |
| reddit posts + commentsReddit: July–December 2025 transparency report ↗ | 139.523 | +0% | unknown |
| code commits pushed to githubGitHub: 986 million commits in its 2025 report ↗ | 31.266 | +0% | unknown |
| hours broadcast on twitchStreamlabs / Stream Hatchet: Q4 2025 report ↗ | 26.092 | +0% | unknown |
hours of video uploaded to youtube · source note
500 uploaded hours per minute; rounded to 8.333 per second. Secondary compilation; the flat rate is a model assumption.
photos & videos posted to instagram · source note
95 million posts per day; rounded to 1,100 per second. Historical secondary estimate. Includes video posts, so the picture comparison is illustrative.
instagram stories · source note
695,000 Stories per minute. Secondary estimate; not a live platform count.
snaps sent · source note
5.5 billion Snaps per day. Snap newsroom figure cited by a secondary compilation.
whatsapp messages · source note
41 million messages per minute. Secondary estimate; messages use a 2 KB average.
emails sent · source note
Approximately 376 billion emails per day. Approximate vendor projection. The exact report edition has not been independently verified.
messages sent to ai assistants · source note
ChatGPT 2.5 billion messages per day, multiplied by 1.6 for other assistants; approximately 4 billion per day. The cross-platform multiplier and 100% annual growth are model assumptions. The underlying paper has not been independently verified for this model version.
ai images generated · source note
120 million images per day; an assumed midpoint within an 80–200 million range. Model estimate informed by platform disclosures and secondary compilations; not a measured cross-platform total.
ai videos generated · source note
Approximately 5 million generated video clips per day. Model estimate informed by Veo/Flow and Kling figures; clips vary substantially in duration and size.
songs uploaded to streaming · source note
150,000 tracks uploaded per day, rounded to 1.74 per second. September 2026 report has not been independently verified for this model version. A 50% AI weighting is an assumption, not a verified industry-wide share.
ai songs generated · source note
Approximately 7 million generated songs per day, based on a reported Suno figure. Generated songs and streaming uploads may overlap. This is an estimate, not a count of unique released songs.
posts uploaded to tiktok · source note
269.3 million posts over the sampled day, April 10, 2024; divided by 86,400 seconds. Historical one-day research estimate, carried forward flat. Version 5 revises earlier estimates. Includes posts later unavailable. The current platform rate and AI share are unknown.
podcast episodes published · source note
27,214,183 new episodes in 2025 in the dataset checked September 19, 2026; divided by 31,536,000 seconds. Public RSS podcasts indexed by Listen Notes, not every audio or video podcast. The dataset is revised over time and filters some AI/spam content. Held flat; AI share is unknown.
wordpress blog posts · source note
70 million new posts per month; multiplied by 12 and divided by an average 365.25-day year. WordPress.com and connected Jetpack network only, not all blogs or web pages. The page does not date this long-running estimate. Held flat; AI share is unknown.
reddit posts + comments · source note
4.4 billion posts and comments during 2025; divided by 31,536,000 seconds. Posts and comments combined, excluding private chats. Held flat. Text-byte estimate excludes attached media; AI share is unknown.
code commits pushed to github · source note
986 million commits reported for the 2025 reporting year, normalised to 365 days. A commit is a change, not a unique file or software release. GitHub only; may include forks, automation and AI-assisted work. Held flat; AI share is unknown.
hours broadcast on twitch · source note
207.4 million creator-hours broadcast on Twitch in Q4 2025; divided by 92 days. Counts hours broadcast, not audience-hours watched. Held flat. Byte estimate represents one encoded stream, excluding viewer copies and multiple renditions; AI share is unknown.
Average sizes
| item | bytes / item | basis |
|---|---|---|
| hours of video uploaded to youtube | 2 250 000 000 | One hour of 1080p video at 5 Mbps. |
| photos & videos posted to instagram | 3 000 000 | Typical 12 MP HEIC/JPEG; applied to a mixed post category. |
| instagram stories | 4 000 000 | Mixed photos and short videos. |
| snaps sent | 4 000 000 | Mixed photos and short videos. |
| whatsapp messages | 2 000 | A text message; attachments averaged into a simple 2 KB assumption. |
| emails sent | 75 000 | Body and headers, with attachments amortised. |
| messages sent to ai assistants | 4 000 | A text-only prompt and reply pair. |
| ai images generated | 1 500 000 | A 1024 × 1024 PNG/WebP image. |
| ai videos generated | 15 000 000 | An 8-second clip at 1080p. |
| songs uploaded to streaming | 9 000 000 | A 4-minute track at 320 kbps, rounded to 9 MB. |
| ai songs generated | 9 000 000 | A 4-minute track at 320 kbps, rounded to 9 MB. |
| posts uploaded to tiktok | 10 210 000 | Assumed 4 Mbps × 20.42 seconds ÷ 8 = 10.21 MB. Duration comes from the sampled posts; bitrate is an illustrative assumption. |
| podcast episodes published | 43 200 000 | Assumed 45-minute audio episode at 128 kbps = 43.2 MB. Duration and bitrate are illustrative, not measured averages. |
| wordpress blog posts | 100 000 | Assumed 100 KB of text and markup per post, excluding embedded media and page assets. |
| reddit posts + comments | 2 000 | Assumed 2 KB per post/comment for text and metadata, excluding attached images and video. |
| code commits pushed to github | 50 000 | Assumed 50 KB per code change, not an entire repository or a measured average diff size. |
| hours broadcast on twitch | 2 250 000 000 | Assumed 5 Mbps × 3,600 seconds ÷ 8 = 2.25 GB per creator-hour; one encoded stream. |
03 AI-generated and other data
÷ Σ(rate × bytes/item)
AI images, videos, songs and assistant exchanges have a machine weight of 1. Streaming uploads have a weight of 0.5; the original other-data categories have a weight of 0. New categories with unknown attribution have no weight and are excluded from both the numerator and denominator. “Other” is the remaining model bucket, not verified human authorship: ordinary messages and posts may also be automated.
A reported Deezer example informed the 50/50 streaming assumption, but one service’s share is not proof of an industry-wide split. AI assistant exchanges include a human prompt. Generated songs can later be uploaded, so the categories are not mutually exclusive.
At launch, this model yields 0.5920% AI-generated by bytes across the selected categories. It does not estimate the machine share of the entire global datasphere.
The picture readout divides the combined rate of Instagram posts, Stories and Snaps by the AI-image rate: 54.96 platform items per AI image at launch. Because some platform items are videos and can overlap, this is not a measured fraction of unique new pictures.
The song readout divides AI-song generation by all streaming uploads: 46.55 at launch. It says “streaming upload” because the denominator includes the configured AI share; describing it as a human-only upload would be incorrect.
04 Growth assumptions
Annual growth rates are modelling choices, not sourced predictions. The baseline is 2026-09-19. Most human categories grow slowly; YouTube upload hours remain flat. The initial doubling of several AI rates represents a period of rapid adoption, with greater uncertainty than the headline series.
For any category with a growth multiplier above 1.5, the instantaneous annual multiplier decays linearly toward 1.2 over 5 years. It then stays at that floor. This keeps totals increasing while avoiding indefinite doubling.
rate(t) = rate₀ × exp(∫₀ᵗ ln(g(u)) du / Y)
items(t) = offset + ∫₀ᵗ rate(u) du
Y is 31,557,600 seconds. The logarithmic growth integral is evaluated analytically. Item totals use 64-slice Simpson integration within the decay period, then an exact exponential tail. Categories without decay use a closed-form integral. Times before the baseline clamp to the baseline.
Values and origin ratios in a frame use the same timestamp. A bounded cache reduces repeated work; it does not accumulate totals. Hidden tabs stop rendering and recompute from the wall clock when visible again.
05 Units
Explore the visual Units of Measure guide ↗
Decimal is the default: 1 TB = 1012 bytes, 1 YB = 1024 bytes. Binary uses powers of 1,024: 1 TiB = 240 bytes. This preference changes compact labels; the full byte integer and the model stay the same.
Saved on this device when browser storage is available.
06 Known gaps
The following forms of media are visible on the counter page without invented totals. “No reliable estimate” means this model has no verified, comparable creation rate; it does not mean zero activity.
Facebook, X, Threads + other social posts
No current, comparable publishing rate has been verified for this model. Platform audiences and impressions do not measure posts created.
documents, PDFs, spreadsheets + slides
Creation is spread across private devices, workplaces and cloud services. No defensible global creation rate is included.
voice messages + video calls
These create audio and video, but participant-minutes, recordings and transmitted copies are different quantities. No comparable creation-rate estimate is included.
photos + videos kept on devices
Social uploads cover only part of photography and video creation. A global rate for files that remain on personal devices has not been verified.
ebooks + audiobooks
ISBNs, editions, formats and retailer listings overlap. No global count of newly created digital files is included.
games, animation + 3D assets
Releases, updates, textures, models and user-created worlds have very different sizes and overlap. No reliable combined creation rate is included.
- The global datasphere includes offline, enterprise and device data, as well as copies and consumption. It is broader than the internet.
- TikTok uses a historical one-day research sample from April 2024, not a current platform disclosure. WordPress covers only its publishing network. Neither represents all social media or web pages.
- Email is a vendor projection. File sizes vary enormously. Category inputs mix dates, reporting scopes and secondary estimates.
- This version uses 149 ZB for 2024; the linked secondary table currently says 147 ZB. That discrepancy is retained and disclosed so the version remains reproducible.
- Some linked platform and 2026 media claims have not been independently verified for this model version. Source notes identify these limitations; a working link does not validate a claim.
- Forecasts are uncertain, especially far beyond the last refresh. The counter depends on the device’s clock. A clock adjustment can move the displayed total backward.
07 Change log
2026-09-19.1
Added TikTok posts, podcast episodes, WordPress posts, Reddit discussions, GitHub commits and Twitch broadcasts, with dated sources and explicit size assumptions. Unmeasured media are listed as gaps. Unknown origins are excluded from the AI-share denominator. Headline and original category baselines are unchanged.
2026-09-19
Initial model: annual datasphere estimates, eleven category counters, origin weights and growth decay. Source limitations recorded.
08 Colophon
Memory Loss is a digital artwork at chorus.computer, inspired by the National Debt Clock.