Only demonstrated progress on the general-intelligence problem can change the score. Narrow-intelligence progress has zero score impact, even when commercially useful or superhuman. A general-purpose foundation model inside a specialized workflow does not make the result general.
Evidence cutoff: September 30, 2026. Prepared October 1. This edition supersedes the earlier provisional workbook scoring.
One scale for every problem
Score
Evidence level
Common requirement
1
General progress unestablished
No reviewed demonstration passes the generality gate for this problem. Narrow advances, theoretical results, announcements and incomplete evidence do not lift the score.
2
Initial general evidence
At least one dated paper or inspectable product demonstration passes every generality and evidence gate and shows a measurable advance on the component-specific test. The problem need not be solved.
3
Repeatable general progress
Level 2 plus at least three successive learning or self-modification cycles with fresh held-out tasks, matched resource accounting, reproducible runs and no material regression of the relevant capability.
4
Independently robust
Level 3 plus replication by an independent team on fresh domains and adversarial distribution shifts, with disclosed failures and successful transfer to successors.
5
Sustained general solution
Level 4 plus the component works in an integrated general self-improvement loop over at least six successive cycles and three months, meeting predeclared success and regression bounds without bespoke human repair.
What qualifies as general progress?
The same underlying mechanism must show the relevant capability in at least three qualitatively different domains, including unfamiliar task structures introduced after the method was fixed. At least one domain must require open-ended task formulation, adaptation and independent outcome verification, rather than only answering a fixed benchmark. No domain-specific redesign or bespoke retraining may explain the transfer. Ordinary task instructions, tool access and the method’s own learning are allowed and must be disclosed. Physical claims require physical outcomes.
Three domains alone is insufficient: multiple coding datasets, several robot tasks, or several subject labels within a single fixed task format do not establish generality. Conversely, a result is not disqualified merely because it uses frozen model weights or an agent harness. Evaluate the demonstrated behavior and causal improvement, not the implementation label. Full AGI is not required for score 2.
A dated primary source must describe the setup, comparator, measurable improvement, uncertainty or run-level outcomes, human assistance, compute/resources, failures and transfer conditions. Inspect the experiments and methods, not just the title or abstract, before awarding points. A live product claim needs a version/date plus reproducible tests or inspectable traces. Marketing announcements, model cards without relevant demonstrations, surveys, theoretical possibility and benchmark counts alone are insufficient. If a required fact is absent, record “not established” and withhold the increase.
These are this tracker’s operational conventions, not a scientific consensus definition of general intelligence. Freeze rubric version 2.0 before future scoring. Amendments require a dated explanation and a separately labeled restated history, never a silent change.
How scores and trends work
Score each of the 18 components independently on the common ordinal scale. A headline area is the minimum of its component scores, because one prerequisite cannot compensate for another. Do not average the seven areas or label a score as a percentage of AGI. Advancement in one component remains visible in the detailed ledger even if the headline minimum does not move.
Every increase must reference a new qualifying source, its first verifiable public date, relevant experimental section, the prior and new score, the satisfied gate, and a written decision. Narrow results never change scores. Negative evidence may trigger a decrease if it invalidates the evidence supporting the current score; the floor is 1. A completed review with no qualifying result carries the score forward with an explicit reason. An incomplete review is NA and a chart gap, never fabricated stasis.
The initial January–September history is a retrospective, limited-coverage public-evidence reconstruction performed on October 1, 2026. It was not measured monthly at the time. January 1 uses pre-2026 evidence; subsequent snapshots only admit evidence available by that month-end. Flat lines mean no qualifying milestone was established in this reviewed corpus, not that all research stopped or that undisclosed progress is impossible. The complete 2026 literature has not been exhaustively assessed. Initial scores are provisional for review.
Repeatable monthly research procedure
Use the previous calendar month in America/Chicago as the reporting period. Load the prior ledger, rubric version and source watchlist. Revisit a 45-day overlap to catch indexing delays and revisions, but record the true public date separately.
Run every component query below with the target year/month and synonyms. Search arXiv, OpenReview, ACL Anthology, PMLR, CVF, Nature and Science; follow citations from the strongest candidates. Search official research and release pages from relevant model labs and robotics/laboratory vendors for demonstrated features. Also search the candidate’s name with “replication”, “failure”, “retraction”, “reward hacking” and “limitation”. Search result snippets only nominate candidates.
Record the exact query, engine/repository, search timestamp, date window, result URLs screened and access failures. Deduplicate by paper and version. Preserve first-public and version dates separately. Read primary methods/results for plausible generality candidates and any claimed score change. Log papers as preprints unless publication is verified.
For each candidate record task families, what was held out, when tasks were introduced, human work, per-domain tuning, resources, comparator, metric, uncertainty, replication status, failure evidence and generality decision. Distinguish no evidence from evidence of failure. Do not add scores for research volume.
Evaluate all 18 components, including ones with no new papers. Record raise, hold, lower or incomplete, with prior/current score and source IDs. Audit source dates against the month cutoff. Do not use later revised results to rewrite earlier months unless issuing a correction.
Produce a versioned PNG infographic, source-linked HTML review, JSON ledger, tweet draft of at most 280 characters, and image alt text. Show year-to-date history, current score and monthly change. Check text, plot labels, source links, dates, and that narrow evidence changed no score.
Deliver the package in this chat for review each month. Never post to X, send messages, or publish externally. If the same reporting month already has a completed package, skip duplicate delivery. Report access failures or incomplete areas in the review package.
Initial search coverage
The initial review used targeted searches across all 18 components and primary-source follow-up, including the sources below. Additional broad query runs were saved locally on October 1. This is a scoped research review, not a systematic review with complete recall. Generality candidates Hyperagents, AIDE² and value-transfer work receive explicit consideration rather than automatic exclusion. Absence of qualifying evidence is the reviewer’s conclusion under the stated rubric.
Problem register and repeatable searches
All component scores in this review: 1/5. Each row specifies the observable evidence that would be needed for score 2, in addition to the common generality gate. Queries are search stems; add the reporting date window and repository filters.
1. Seed research intelligence 1 / 5
General reasoning, memory, agency and reality
Component
Test for progress
Repeatable query / evidence
1.1 Grounded world models 1 / 5
Predict and test novel interventions across physical and informational domains; measure intervention accuracy and calibration.
world model causal generalization unseen environments B04 · R02 · R09
1.2 Reliable general reasoning 1 / 5
Solve independently introduced task structures without bespoke tuning; report success, error detection and calibration across domains.
general reasoning out of distribution transfer B02 · R09 · R13 · R14
1.3 Continual learning / durable memory 1 / 5
Acquire unfamiliar skills sequentially while retaining earlier ones; measure forward transfer and worst-domain forgetting.
Correctly accept genuine improvements and reject adversarial exploits on unseen domains; use external ground truth and report false accept/reject rates.
All seven headline areas and all 18 components remain provisionally at 1. The January 1 baseline and nine month-end snapshots are in the JSON ledger. Evidence below explains why research activity did not trigger a general-intelligence upgrade.
Hold 1 → 1. Research-agent transfer and long-duration mathematics are notable; open-world generality and trustworthy oversight are not established.
Evidence review
Source claims are summarized briefly. “Does not qualify” is a scope decision for this tracker, not a claim that the research lacks value. Paper results are author-reported unless replication is explicitly identified. No qualifying independent replication is asserted here.
Reported evidence: A-Lab integrates planning and physical materials synthesis.
Score impact: 0. Domain-specific laboratory; no evidence of transfer to general physical experimentation. Numerical claims omitted because the article has corrections.
Reported evidence: Editable task and meta agents improve across several domains; learned meta-level changes transfer across domains and runs.
Score impact: 0. Strong transfer candidate. The reviewed experiments remain benchmark-defined with domain-specific evaluations; they do not establish autonomous adaptation to independently introduced, unfamiliar problem structures under the full generality gate. This is an assessment judgment, not the authors’ conclusion.
Reported evidence: Reports improved generalization of specified values and reduced misalignment in held-out evaluations.
Score impact: 0. Promising value-transfer evidence. Does not establish general alignment under open-world conflicts, reduced oversight, and self-modifying successors.
Reported evidence: Causality-focused training improves causal benchmarks and reasoning faithfulness in several application settings.
Score impact: 0. Cross-setting transfer is relevant, but benchmark reasoning does not establish general grounded intervention and unfamiliar-task competence.
Reported evidence: Learns predictions of internal interventions; self-knowledge remains incomplete and does not outperform direct repair in the reported ResNet comparison.
Score impact: 0. Limited predictive self-model evidence, with a negative transfer result; no general self-understanding.
Reported evidence: Internal estimates report substantial AI participation in R&D, but no measured task subset classified as fully autonomous.
Score impact: 0. An internal workflow study, not a demonstrated self-sustaining general research loop. August measurements first enter this tracker in September, when published.
Version: Official report. Components: 1.4, 6.1, 6.2.
Reported evidence: Reports seven successive improvements over eight days, transfer to four held-out benchmarks, and reduced reward hacking on another task family.
Score impact: 0. Strong transfer candidate, not dismissed as coding-only. The evidence concerns a bounded research-agent protocol; it does not establish the full generality gate or sustained acceleration across open-ended unfamiliar problems. Counted in September using the dated paper, not a subsequently updated earlier blog.
Reported evidence: Across 17 models and 38 tasks, the study finds spontaneous reward hacking and adaptive evasion of review.
Score impact: 0. Counterevidence to trusting reported scores alone. Rates are specific to tested tasks, not estimates of all AI behavior.
Version: v1. Components: 3.1, 7.1, 7.2.
Source article and editorial scope
The seven headline areas follow the article’s numbered sections; the 18 components unpack its subproblems. Labels are shortened for readability. The source article is a framing document, not empirical proof of current scores: Before We Panic About Superintelligence, We Should Probably Invent It First. The tracker does not infer that unsolved prerequisites eliminate AI risks, or that bounded recursive improvement is impossible.
Monthly delivery
Requested cadence: first day of each month at 9:00 a.m. America/Chicago, covering the previous calendar month. Prepare for review only. The scheduler configuration is maintained separately; this report does not itself execute a schedule.