Trend archives look smooth in a chart because lines connect points. Databases are less polite: collectors stop, responses arrive empty, endpoints change, and a day can vanish. Before publishing any Q1 comparison, we measured those holes.
Seven of the eight selected feeds had observations on most or all of the 90 calendar days. The least complete source in this sample was Hacker News Best, with 89 covered days (98.9%). That does not automatically invalidate the source. It changes which questions we can ask of it.
Calendar-day coverage in the 90-day window
Three kinds of completeness
Day coverage asks whether any usable snapshot exists on a date. Capture interval asks how far apart stored observations usually are. Longest gap looks for outages hidden by the median. A feed can score well on one and poorly on another: one capture every day yields perfect day coverage but cannot support hour-level claims.
| Source | Day coverage | Median interval | Longest observed gap |
|---|---|---|---|
| 100% | 62 min | 1.1 h | |
| Tieba | 100% | 65 min | 1.2 h |
| Toutiao | 100% | 61.9 min | 1.1 h |
| Baidu | 100% | 61.8 min | 1.1 h |
| Hacker News Top | 100% | 69.7 min | 23.2 h |
| Zhihu | 100% | 69.7 min | 1.4 h |
| V2EX | 100% | 69.7 min | 4.7 h |
| Hacker News Best | 98.9% | 89.5 min | 45 h |
The median interval describes our collection system, not an official refresh promise. For example, Baidu publishes its own update rule, but the table reports when TrendGoing successfully stored observations. Network failures, response parsing, and scheduler behavior sit between those two facts.
How gaps change an article
If a missing period overlaps the moment a topic entered or left a list, duration becomes an interval rather than an exact number. A topic last seen at 10:00 and next absent at 14:00 did not necessarily disappear at 14:00; the defensible statement is that the transition happened somewhere between the two observations.
Cross-platform comparisons need an even stricter rule. When one source is missing on a day, we exclude that paired day rather than carrying its previous list forward. Forward-filling would create artificial persistence and make the less frequently observed source appear more stable.
Our publication threshold
For the six-hour retention note, we required a non-empty hourly sample and an exact partner six hours later. For hourly leader changes, we accepted only hour pairs one or two hours apart. For Hacker News overlap, both Top and Best had to have a sample on the same date. Each article states its denominator so a reader can reproduce or reject the choice.
We are publishing this gap audit because methodological honesty is part of the product. A trend dashboard is most useful when it preserves uncertainty instead of converting missing data into a confident story. Future research notes will identify their query window, rank cutoff, sampling rule, and exclusions in the same way.