RSI monthly review protocol

Only demonstrated progress on the general-intelligence problem can change the score. Narrow-intelligence progress has zero score impact, even when commercially useful or superhuman. A general-purpose foundation model inside a specialized workflow does not make the result general.

What qualifies as general progress?

The same underlying mechanism must show the relevant capability in at least three qualitatively different domains, including unfamiliar task structures introduced after the method was fixed. At least one domain must require open-ended task formulation, adaptation and independent outcome verification, rather than only answering a fixed benchmark. No domain-specific redesign or bespoke retraining may explain the transfer. Ordinary task instructions, tool access and the method’s own learning are allowed and must be disclosed. Physical claims require physical outcomes.

Three domains alone is insufficient: multiple coding datasets, several robot tasks, or several subject labels within a single fixed task format do not establish generality. Conversely, a result is not disqualified merely because it uses frozen model weights or an agent harness. Evaluate the demonstrated behavior and causal improvement, not the implementation label. Full AGI is not required for score 2.

A dated primary source must describe the setup, comparator, measurable improvement, uncertainty or run-level outcomes, human assistance, compute/resources, failures and transfer conditions. Inspect the experiments and methods, not just the title or abstract, before awarding points. A live product claim needs a version/date plus reproducible tests or inspectable traces. Marketing announcements, model cards without relevant demonstrations, surveys, theoretical possibility and benchmark counts alone are insufficient. If a required fact is absent, record “not established” and withhold the increase.

These are this tracker’s operational conventions, not a scientific consensus definition of general intelligence. Freeze rubric version 2.0 before future scoring. Amendments require a dated explanation and a separately labeled restated history, never a silent change.

How scores and trends work

Score each of the 18 components independently on the common ordinal scale. A headline area is the minimum of its component scores, because one prerequisite cannot compensate for another. Do not average the seven areas or label a score as a percentage of AGI. Advancement in one component remains visible in the detailed ledger even if the headline minimum does not move.

Every increase must reference a new qualifying source, its first verifiable public date, relevant experimental section, the prior and new score, the satisfied gate, and a written decision. Narrow results never change scores. Negative evidence may trigger a decrease if it invalidates the evidence supporting the current score; the floor is 1. A completed review with no qualifying result carries the score forward with an explicit reason. An incomplete review is NA and a chart gap, never fabricated stasis.

The initial January–September history is a retrospective, limited-coverage public-evidence reconstruction performed on October 1, 2026. It was not measured monthly at the time. January 1 uses pre-2026 evidence; subsequent snapshots only admit evidence available by that month-end. Flat lines mean no qualifying milestone was established in this reviewed corpus, not that all research stopped or that undisclosed progress is impossible. The complete 2026 literature has not been exhaustively assessed. Initial scores are provisional for review.

Repeatable monthly research procedure

  1. Use the previous calendar month in America/Chicago as the reporting period. Load the prior ledger, rubric version and source watchlist. Revisit a 45-day overlap to catch indexing delays and revisions, but record the true public date separately.
  2. Run every component query below with the target year/month and synonyms. Search arXiv, OpenReview, ACL Anthology, PMLR, CVF, Nature and Science; follow citations from the strongest candidates. Search official research and release pages from relevant model labs and robotics/laboratory vendors for demonstrated features. Also search the candidate’s name with “replication”, “failure”, “retraction”, “reward hacking” and “limitation”. Search result snippets only nominate candidates.
  3. Record the exact query, engine/repository, search timestamp, date window, result URLs screened and access failures. Deduplicate by paper and version. Preserve first-public and version dates separately. Read primary methods/results for plausible generality candidates and any claimed score change. Log papers as preprints unless publication is verified.
  4. For each candidate record task families, what was held out, when tasks were introduced, human work, per-domain tuning, resources, comparator, metric, uncertainty, replication status, failure evidence and generality decision. Distinguish no evidence from evidence of failure. Do not add scores for research volume.
  5. Evaluate all 18 components, including ones with no new papers. Record raise, hold, lower or incomplete, with prior/current score and source IDs. Audit source dates against the month cutoff. Do not use later revised results to rewrite earlier months unless issuing a correction.
  6. Produce a versioned PNG infographic, source-linked HTML review, JSON ledger, tweet draft of at most 280 characters, and image alt text. Show year-to-date history, current score and monthly change. Check text, plot labels, source links, dates, and that narrow evidence changed no score.
  7. Deliver the package in this chat for review each month. Never post to X, send messages, or publish externally. If the same reporting month already has a completed package, skip duplicate delivery. Report access failures or incomplete areas in the review package.

Initial search coverage

The initial review used targeted searches across all 18 components and primary-source follow-up, including the sources below. Additional broad query runs were saved locally on October 1. This is a scoped research review, not a systematic review with complete recall. Generality candidates Hyperagents, AIDE² and value-transfer work receive explicit consideration rather than automatic exclusion. Absence of qualifying evidence is the reviewer’s conclusion under the stated rubric.

Use the latest dated JSON ledger for the 18 component queries and evidence history. Maintain immutable monthly editions; never overwrite prior releases silently.