Research question
Can observable news volume, publisher activity, community attention and topic growth form a useful descriptive index of a selected AI-news corpus? What changes when the sources, weights or classification method change?
Current approach
A deterministic baseline compares the most recent seven days with the previous 28 days. Four signals are transformed and combined. The score describes relative activity in those sources, not technological capability, social impact, personal relevance or sentiment.
The descriptive band names now refer to activity. The numerical baseline remains unchanged, and legacy band IDs remain in the JSON for compatibility. Read the full methodology for assumptions and exact arithmetic.
What the evidence supports
The repository contains a collected corpus, a historical reconstruction and reproducible calculations. The current snapshot includes equal-weight, signal-exclusion and source-exclusion checks. Those checks measure dependence on specified choices; they do not establish construct validity.
There is no independent human benchmark yet. Retrospective agreement with memorable announcements is descriptive evidence, not a held-out validation result. Backfilled headlines and later Hacker News scores are not what an observer necessarily knew at the historical date.
The semantic experiment
Jev / System One is a candidate for assessing headline relevance, reported changes and specificity. Its structured outputs make these judgments inspectable, but correct output types do not guarantee correct judgments. A separate pilot records exact inputs, rubric, model version and returned distributions. It contributes no weight to the baseline.
The first task is a human-annotated comparison, including unclear and irrelevant headlines. Measuring novelty or impact would require richer evidence and a separate rubric. Read the protocol →
An invitation to investigate
Contributions can audit source coverage, label a sample, test event deduplication or compare another specification. Negative results and documented disagreements are useful contributions. Proposed changes should preserve a frozen baseline and report what changes across the full evaluation set.
Reuse & citation
Code is MIT; data is CC BY 4.0. Cite the project, method version, observation date and repository commit. Archive the JSON snapshots with the commit when using results in another study. This is an exploratory working study; no peer-review or DOI status is implied.