← About the index

Method, assumptions & limits

The arithmetic, its limitations and the next experiment. The reference implementation is in fomo-index.ts.

Status: exploratory working study. Method v1.0.0. The index describes relative activity in a selected AI-news corpus. It does not measure technical progress, sentiment, human anxiety or the importance of keeping up. The name “FOMO” is motivational, not the measured construct.

The current baseline preserves the original arithmetic. Windows, weights, keyword rules, point thresholds, pseudocount, curve exponent and band boundaries are modeling choices, not learned or independently validated constants.

1. Two windows

Every reading is taken at a moment TT, the most recent first-seen timestamp in data/seen.json (falling back to the newest publication date if unavailable). Around it we draw two windows:

(T−35d,  T−7d ]⏟baseline: 28 days(T−7d,  T ]⏟window: this week\underbrace{(T - 35\text{d},\; T - 7\text{d}\,]}_{\text{baseline: 28 days}} \qquad \underbrace{(T - 7\text{d},\; T\,]}_{\text{window: this week}}

For any set of stories SS, let CC be how many fell in this week's window and NN how many fell in the baseline. The baseline covers four weeks, so an average week in it holds

B=728 N=N4.B = \frac{7}{28}\,N = \frac{N}{4}.

CC and BB are now on the same scale: stories per week.

2. A ratio, softened

The pace of a signal is this week against an average week:

r=C+kB+k,k=2.r = \frac{C + k}{B + k}, \qquad k = 2.

The pseudocount kk keeps small signals from swinging wildly. Without it, a signal that goes from one story to three reads as 3×3\times the usual pace. With it, the same jump is

r=3+21+2≈1.67,r = \frac{3 + 2}{1 + 2} \approx 1.67,

still a clear rise, but not a panic. For large counts kk barely matters: 262/202262/202 and 260/200260/200 are both 1.301.30.

Ratios compare a source mix against its own recent history. Adding a source can still change the result if its activity pattern differs. A source only counts in a reading when its declared coverage starts before that reading's baseline begins, which prevents immediate entry from being treated as a spike. This does not prove that source coverage is complete or that later source admission has no effect.

3. From ratio to score

A ratio lives on (0,∞)(0, \infty). The index lives on [0,100][0, 100]. One curve maps between them:

s(r)=100⋅r41+r4s(r) = 100 \cdot \frac{r^4}{1 + r^4}

rounded to a whole number. Some properties worth knowing:

The exponent 4 is the steepness. Weekly AI news volume is steady, so a gentle curve (r2r^2) left most weeks bunched between 45 and 63. With r4r^4, ordinary weeks spread across roughly 40–70, and higher values become more sensitive to count changes. This choice was informed by the same historical data shown here, not a held-out validation set.

Inverting the curve gives the smoothed ratio for one signal, not the composite index or an absolute rate of AI progress:

r(s)=(s100−s)1/4r(s) = \left(\frac{s}{100 - s}\right)^{1/4}
Individual signal score Smoothed ratio
19 0.70×0.70\times
35 0.86×0.86\times
50 1.00×1.00\times
55 1.05×1.05\times
75 1.32×1.32\times
79 1.40×1.40\times
90 1.73×1.73\times

4. Three counting signals

Three signals are the formula above applied to a different set of stories.

Signal Stories counted Weight ww
News velocity Every story from every source 0.300.30
Lab activity Lab-source headlines matching model, product or agent keywords 0.250.25
Community attention Stories with 200+200+ points on Hacker News 0.250.25

The lab-source group is OpenAI, Anthropic, Google AI, DeepMind, Mistral and Hugging Face. This grouping includes community posts on Hugging Face; it is not a guarantee of first-party releases. Keywords are applied to titles at build time. A lab post can match the proxy without announcing a launch. Older HN rows missing points can fall back to the legacy importance field (≥3), whose model/heuristic provenance was not recorded per row. For each signal, si=s ⁣(Ci+2Bi+2)s_i = s\!\left(\dfrac{C_i + 2}{B_i + 2}\right).

5. The topic spike

Every story carries topics such as models, agents, hardware or policy. The fourth signal asks whether any single topic is suddenly louder than usual.

Call T\mathcal{T} the topics with at least 5 stories this week. Compute each one's own ratio, and take the largest:

R(T)=max⁡{Ct+2Bt+2  :  t∈T}R(T) = \max \left\{ \frac{C_t + 2}{B_t + 2} \;:\; t \in \mathcal{T} \right\}

(If no topic reaches 5 stories, R(T)=1R(T) = 1.)

There is a catch. Some topic always grows fastest, so R(T)R(T) sits well above 1 even in a dull week. Scoring it directly would pin this signal near the top every week. Instead, it is compared with the biggest spike of each of the four previous weeks:

Rˉ=14∑j=14R(T−7j days),sspike=s ⁣(R(T)Rˉ).\bar{R} = \frac{1}{4}\sum_{j=1}^{4} R(T - 7j\ \text{days}), \qquad s_{\text{spike}} = s\!\left(\frac{R(T)}{\bar{R}}\right).

A score of 50 now means "a typical week's biggest spike", and only an unusually sharp one scores high.

6. The index

The FOMO Index is the weighted mean of the four signal scores:

FOMO(T)=round⁡ ⁣(∑iwi si∑iwi)\text{FOMO}(T) = \operatorname{round}\!\left(\frac{\sum_i w_i\, s_i}{\sum_i w_i}\right)

and since the weights add up to 1, that is simply

FOMO(T)=round⁡( 0.30 svelocity+0.25 slaunches+0.25 sheat+0.20 sspike).\begin{aligned} \text{FOMO}(T) = \operatorname{round}\big(\,&0.30\,s_{\text{velocity}} + 0.25\,s_{\text{launches}} \\ {}+{} &0.25\,s_{\text{heat}} + 0.20\,s_{\text{spike}}\big). \end{aligned}

A composite score cannot be inverted to a unique news-volume multiplier: signals have different counts, overlap and undergo a nonlinear transformation before averaging. The 0–100 scale is neither a percentile nor a probability.

It falls into descriptive bands with chosen boundaries:

Score Band
0≤x<350 \leq x < 35 Low activity
35≤x<5535 \leq x < 55 Near baseline
55≤x<7555 \leq x < 75 Elevated
75≤x≤10075 \leq x \leq 100 High activity

The daily history is the same reading repeated at T, T−1d, T−2d, …T,\ T - 1\text{d},\ T - 2\text{d},\ \dots back to November 2022. The "vs. last week" arrow compares FOMO(T)\text{FOMO}(T) with FOMO(T−7d)\text{FOMO}(T - 7\text{d}).

Long-run level

The index above is relative to its own last month, so over years it returns to 50 by construction and cannot show a trend. The long-run level is a separate, companion measure for that. For each counting signal (velocity, launches, heat) it compares the trailing 13 weeks with the first tracked year (the 52 weeks from historyFrom), using only sources covered since historyFrom, so adding a source never looks like growth:

ρi=Qi+21352Yi+2,Level(T)=round⁡ ⁣(100⋅exp⁡∑iwiln⁡ρi∑iwi),\rho_i = \frac{Q_i + 2}{\tfrac{13}{52} Y_i + 2}, \qquad \text{Level}(T) = \operatorname{round}\!\left(100 \cdot \exp \frac{\sum_i w_i \ln \rho_i}{\sum_i w_i}\right),

where QiQ_i is the signal's count in the 13 weeks ending at TT and YiY_i its count in the reference year. The geometric mean keeps one fast-growing, low-volume signal from dominating. The topic spike is excluded because it is relative by design. 100100 means "as active as the reference year" (336→3.4×336 \to 3.4\times).

The Overall FOMO Index puts that level on the same 0–100 scale and bands as the weekly index, with a gentler curve than section 3 because multi-year growth spans much larger ratios than week-to-week change:

r=Level/100,Overall(T)=round⁡ ⁣(100 r1+r),r = \text{Level}/100,\qquad \text{Overall}(T) = \operatorname{round}\!\left(\frac{100\,r}{1 + r}\right),

so 1×→501\times \to 50, 2×→672\times \to 67, 3×→753\times \to 75, 4×→804\times \to 80 and 9×→909\times \to 90. History is weekly. The level also rises when sources simply publish more or when Hacker News grows; it is not a measure of AI progress.

7. A worked example

Suppose this is a week's raw data:

Signal CC NN B=N/4B = N/4 rr s(r)s(r)
Velocity 260 800 200 262/202=1.297262/202 = 1.297 74
Lab activity 9 20 5 11/7=1.57111/7 = 1.571 86
Attention 14 64 16 16/18=0.88916/18 = 0.889 38
Spike 3.2/2.5=1.283.2/2.5 = 1.28 73

Take velocity: 1.2974≈2.831.297^4 \approx 2.83, so s=100⋅2.83/3.83≈74s = 100 \cdot 2.83 / 3.83 \approx 74.

The index is

0.30⋅74+0.25⋅86+0.25⋅38+0.20⋅73=22.2+21.5+9.5+14.6=67.8  →  68,0.30 \cdot 74 + 0.25 \cdot 86 + 0.25 \cdot 38 + 0.20 \cdot 73 = 22.2 + 21.5 + 9.5 + 14.6 = 67.8 \;\to\; 68,

which lands in Elevated: busier than usual, mostly because the lab-activity proxy rose, while the count of high-attention stories fell. Neither signal verifies releases or measures human emotion.

8. Employment mentions

The jobs signal counts headlines about jobs, layoffs, hiring and automation of work, and is scored exactly like the counting signals:

sjobs=s ⁣(Cjobs+2Bjobs+2).s_{\text{jobs}} = s\!\left(\frac{C_{\text{jobs}} + 2}{B_{\text{jobs}} + 2}\right).

Its weight is 00: it is exported alongside the index and included in the report, but never moves the composite. Counts can be very small, so it stays experimental. It measures employment-related headlines, not labor-market conditions.

Limitations

Sensitivity analysis

The overview and research.json recompute the current snapshot under: equal signal weights; omission of each signal with remaining weights renormalized; and omission of each eligible source from all windows, including the topic-spike reference windows. The published specification is included in the reported minimum and maximum.

This range is not a confidence interval, does not quantify sampling error and is not exhaustive. In particular, it does not vary windows, the exponent, the pseudocount, keyword lists or HN thresholds. Removing a source changes the observed population; a small score change does not prove representativeness. Full-history sensitivity and external validation remain research tasks.

Reproducibility

research.json records the observation timestamp, method version, source counts, specified alternatives and SHA-256 fingerprints of the dataset, coverage, first-seen ledger and calculation code. fomo.json exposes the index and daily reconstruction. Reproduction requires the corresponding repository commit and data, not only a hash. Run npm ci, npm run backtest and npm run build from that checkout.

The timestamp is the newest first-seen record, not evidence that a crawler ran successfully at that time. Rebuilding identical files preserves the score; changing point totals or coverage can revise historical scores. Archive outputs and commits when citing a result. The method is not preregistered or independently validated.

Semantic experiment

Protocol: jev-headlines-v1. Model annotations are not human-validated; zero weight in the index. Jev / System One uses three narrow questions: AI relevance (including an insufficient-information option), whether a concrete change is reported (including an insufficient-information option), and headline specificity on three descriptive levels. Specificity is an information-detail score, not an impact score.

Inputs are the headline and publication date. URLs, source labels, HN points and existing importance are omitted from model state to reduce direct popularity and publisher cues, although names in headlines can still reveal a publisher. Headline-only evidence cannot substantiate real-world importance, novelty, factual truth or a claim of a breakthrough.

The pilot pins a model version and stores the complete request, input and rubric hashes, full answer distributions, confidence, token usage and timestamps. Missing or malformed responses stop the experiment; they are not replaced by neutral scores. npm run semantic:pilot only prepares an offline plan and human-label template. An explicit --execute and a locally configured provider key (TypeSafe or OpenRouter) are required for API calls. Pilot outputs never alter the index automatically. A separate resumable corpus runner appends annotations to data/semantic/annotations.jsonl; the scheduled GitHub pipeline processes only missing or changed inputs after each crawl. Coverage is reported in data/semantic/summary.json and the research snapshot. Historical annotations are retrospective, not real-time predictions.

Before proposing an additional index dimension:

  1. Freeze a source- and time-stratified sample, the rubric and evaluation split before reading model outputs. The provided source-balanced pilot is an audit sample, not a prevalence estimate.
  2. Obtain independent labels from at least two annotators, blinded to model answers. Record disagreement and adjudication instead of forcing unclear items into a class.
  3. Use a development split to refine the rubric; evaluate once on a held-out, preferably later-period split. Keep reports of the same event within one split.
  4. Compare to keyword rules. Report per-class precision/recall, macro-F1, abstention coverage and error at that coverage, with source-level breakdowns. Report Brier score and reliability plots for probabilistic classes. For specificity, report ordinal agreement and absolute error. Evaluate calibration against human labels, not the model's own confidence.
  5. Test repeated runs, headline paraphrases and another pinned model. Keep raw disagreements and failures. Label historical model judgments as retrospective because model training may include later events.
  6. Only then propose a versioned semantic series with a sufficiently covered reference window. Publish it alongside the baseline before considering a combined score. Missing semantic coverage is missing data, never zero activity.

The runnable implementation and fuller protocol are in scripts/semantic-pilot.mjs, scripts/semantic-corpus.mjs and docs/semantic-pilot.md in the repository.

References