Before using an archive, count what is actually in it. That sounds obvious, but trend datasets invite a particular mistake: a large row count can be mistaken for a large audience. Our Q1 2026 extract contains 15,780 saved ranking states from eight feeds. It does not contain 15,780 people, searches, posts, or independent events.
We ran this audit because later comparisons depend on this distinction. Each row in TrendGoing's trends table records a source, a timestamp, and an ordered list of internal item identifiers. When the same source is captured again, another row is added. A fast collection schedule therefore produces more rows even when the visible list barely moves.
Captured snapshots by source, Q1 2026
What the row count tells us
The chart is useful for operational questions. It shows which sources were observed often enough to support hour-by-hour work, and it exposes feeds that were collected at a different cadence. It is not a popularity ranking. Weibo having more saved states than another source means our collector saw Weibo more often; it does not mean Weibo generated more public interest.
| Source | Snapshots | Active days | Unique item IDs | Median captures/day |
|---|---|---|---|---|
| 2,109 | 90 | 32,208 | 23 | |
| Tieba | 2,031 | 90 | 1,948 | 23 |
| Toutiao | 2,113 | 90 | 23,072 | 23 |
| Baidu | 2,115 | 90 | 14,754 | 23.5 |
| Hacker News Top | 1,907 | 90 | 9,554 | 22 |
| Zhihu | 1,959 | 90 | 2,714 | 22 |
| Hacker News Best | 1,600 | 89 | 2,941 | 18 |
| V2EX | 1,946 | 90 | 5,468 | 22 |
The unique-ID column is closer to a measure of archive breadth, but it also needs care. An ID identifies an item inside TrendGoing's historical collection. It does not prove that two differently worded items describe different real-world events, and it does not merge translations or near-duplicates. We keep that limitation visible rather than presenting a false count of “unique trends.”
Why eight feeds, not every row in the database?
The database includes legacy and experimental resource IDs. For this note we used eight named feeds with interpretable metadata and observations throughout the selected quarter: Weibo, Tieba, Toutiao, Baidu, Hacker News Top, Zhihu, Hacker News Best, and V2EX. We excluded orphaned resource IDs and feeds without adequate coverage. The decision reduces the headline number, but makes the result reproducible.
The upstream surfaces also measure different things. Baidu says its hot index combines search, information consumption, and field-specific signals; Weibo describes a composite of search and engagement signals; Zhihu describes a list built around activity on questions and answers. Hacker News exposes separate topstories and beststories lists through its official API. Treating these feeds as interchangeable would erase the reason they are interesting.
How to use this archive responsibly
Start with a question that matches the data. “How quickly did the saved list change?” is answerable. “How many people cared?” is not. “Did the same internal item remain visible across two captures?” is answerable. “Was the public persuaded?” is not. A ranking archive is strong evidence about recorded list movement and weak evidence about motives.
For cross-platform work, normalize the observation schedule first. Our later notes use one sample near noon per day or one sample per observed hour so that collector frequency does not dominate the comparison. We also report missing days and avoid filling gaps with invented values.