MONTHLY RSI / SEPTEMBER 2026 / REVIEW EDITION

General intelligence.
Evidence before points.

Only demonstrated progress on the general-intelligence problem can change the score. Narrow-intelligence progress has zero score impact, even when commercially useful or superhuman. A general-purpose foundation model inside a specialized workflow does not make the result general.

Evidence cutoff: September 30, 2026. Prepared October 1. This edition supersedes the earlier provisional workbook scoring.

Seven general-intelligence areas remain provisionally at one out of five from January through September under the stated evidence gate.

One scale for every problem

ScoreEvidence levelCommon requirement
1General progress unestablishedNo reviewed demonstration passes the generality gate for this problem. Narrow advances, theoretical results, announcements and incomplete evidence do not lift the score.
2Initial general evidenceAt least one dated paper or inspectable product demonstration passes every generality and evidence gate and shows a measurable advance on the component-specific test. The problem need not be solved.
3Repeatable general progressLevel 2 plus at least three successive learning or self-modification cycles with fresh held-out tasks, matched resource accounting, reproducible runs and no material regression of the relevant capability.
4Independently robustLevel 3 plus replication by an independent team on fresh domains and adversarial distribution shifts, with disclosed failures and successful transfer to successors.
5Sustained general solutionLevel 4 plus the component works in an integrated general self-improvement loop over at least six successive cycles and three months, meeting predeclared success and regression bounds without bespoke human repair.

What qualifies as general progress?

The same underlying mechanism must show the relevant capability in at least three qualitatively different domains, including unfamiliar task structures introduced after the method was fixed. At least one domain must require open-ended task formulation, adaptation and independent outcome verification, rather than only answering a fixed benchmark. No domain-specific redesign or bespoke retraining may explain the transfer. Ordinary task instructions, tool access and the method’s own learning are allowed and must be disclosed. Physical claims require physical outcomes.

Three domains alone is insufficient: multiple coding datasets, several robot tasks, or several subject labels within a single fixed task format do not establish generality. Conversely, a result is not disqualified merely because it uses frozen model weights or an agent harness. Evaluate the demonstrated behavior and causal improvement, not the implementation label. Full AGI is not required for score 2.

A dated primary source must describe the setup, comparator, measurable improvement, uncertainty or run-level outcomes, human assistance, compute/resources, failures and transfer conditions. Inspect the experiments and methods, not just the title or abstract, before awarding points. A live product claim needs a version/date plus reproducible tests or inspectable traces. Marketing announcements, model cards without relevant demonstrations, surveys, theoretical possibility and benchmark counts alone are insufficient. If a required fact is absent, record “not established” and withhold the increase.

These are this tracker’s operational conventions, not a scientific consensus definition of general intelligence. Freeze rubric version 2.0 before future scoring. Amendments require a dated explanation and a separately labeled restated history, never a silent change.

How scores and trends work

Score each of the 18 components independently on the common ordinal scale. A headline area is the minimum of its component scores, because one prerequisite cannot compensate for another. Do not average the seven areas or label a score as a percentage of AGI. Advancement in one component remains visible in the detailed ledger even if the headline minimum does not move.

Every increase must reference a new qualifying source, its first verifiable public date, relevant experimental section, the prior and new score, the satisfied gate, and a written decision. Narrow results never change scores. Negative evidence may trigger a decrease if it invalidates the evidence supporting the current score; the floor is 1. A completed review with no qualifying result carries the score forward with an explicit reason. An incomplete review is NA and a chart gap, never fabricated stasis.

The initial January–September history is a retrospective, limited-coverage public-evidence reconstruction performed on October 1, 2026. It was not measured monthly at the time. January 1 uses pre-2026 evidence; subsequent snapshots only admit evidence available by that month-end. Flat lines mean no qualifying milestone was established in this reviewed corpus, not that all research stopped or that undisclosed progress is impossible. The complete 2026 literature has not been exhaustively assessed. Initial scores are provisional for review.

Repeatable monthly research procedure

  1. Use the previous calendar month in America/Chicago as the reporting period. Load the prior ledger, rubric version and source watchlist. Revisit a 45-day overlap to catch indexing delays and revisions, but record the true public date separately.
  2. Run every component query below with the target year/month and synonyms. Search arXiv, OpenReview, ACL Anthology, PMLR, CVF, Nature and Science; follow citations from the strongest candidates. Search official research and release pages from relevant model labs and robotics/laboratory vendors for demonstrated features. Also search the candidate’s name with “replication”, “failure”, “retraction”, “reward hacking” and “limitation”. Search result snippets only nominate candidates.
  3. Record the exact query, engine/repository, search timestamp, date window, result URLs screened and access failures. Deduplicate by paper and version. Preserve first-public and version dates separately. Read primary methods/results for plausible generality candidates and any claimed score change. Log papers as preprints unless publication is verified.
  4. For each candidate record task families, what was held out, when tasks were introduced, human work, per-domain tuning, resources, comparator, metric, uncertainty, replication status, failure evidence and generality decision. Distinguish no evidence from evidence of failure. Do not add scores for research volume.
  5. Evaluate all 18 components, including ones with no new papers. Record raise, hold, lower or incomplete, with prior/current score and source IDs. Audit source dates against the month cutoff. Do not use later revised results to rewrite earlier months unless issuing a correction.
  6. Produce a versioned PNG infographic, source-linked HTML review, JSON ledger, tweet draft of at most 280 characters, and image alt text. Show year-to-date history, current score and monthly change. Check text, plot labels, source links, dates, and that narrow evidence changed no score.
  7. Deliver the package in this chat for review each month. Never post to X, send messages, or publish externally. If the same reporting month already has a completed package, skip duplicate delivery. Report access failures or incomplete areas in the review package.

Initial search coverage

The initial review used targeted searches across all 18 components and primary-source follow-up, including the sources below. Additional broad query runs were saved locally on October 1. This is a scoped research review, not a systematic review with complete recall. Generality candidates Hyperagents, AIDE² and value-transfer work receive explicit consideration rather than automatic exclusion. Absence of qualifying evidence is the reviewer’s conclusion under the stated rubric.

Problem register and repeatable searches

All component scores in this review: 1/5. Each row specifies the observable evidence that would be needed for score 2, in addition to the common generality gate. Queries are search stems; add the reporting date window and repository filters.

1. Seed research intelligence 1 / 5

General reasoning, memory, agency and reality

ComponentTest for progressRepeatable query / evidence
1.1 Grounded world models
1 / 5
Predict and test novel interventions across physical and informational domains; measure intervention accuracy and calibration.world model causal generalization unseen environments
B04 · R02 · R09
1.2 Reliable general reasoning
1 / 5
Solve independently introduced task structures without bespoke tuning; report success, error detection and calibration across domains.general reasoning out of distribution transfer
B02 · R09 · R13 · R14
1.3 Continual learning / durable memory
1 / 5
Acquire unfamiliar skills sequentially while retaining earlier ones; measure forward transfer and worst-domain forgetting.continual learning agent memory catastrophic forgetting
B07 · R06 · R11
1.4 Long-horizon agency
1 / 5
Complete unfamiliar multi-stage projects with independently verified outcomes; report intervention counts, duration and recovery from errors.long horizon autonomous research agent
B02 · R03 · R14 · R15 · R17
1.5 Open-ended learning
1 / 5
Generate useful new problem families, learn transferable skills, and keep expanding competence after a fixed curriculum ends.open ended learning general intelligence
B02 · R01 · R10
1.6 Acting on reality
1 / 5
Plan and execute novel physical experiments beyond the original application; report real outcomes, failures and human setup.autonomous laboratory robotics transfer
B04 · R02

2. Self-understanding & modification 1 / 5

Reliable self-models and capable successors

ComponentTest for progressRepeatable query / evidence
2.1 Functional self-model
1 / 5
Predict consequences of previously untested internal interventions and use those predictions to improve broad capabilities.self model neural network intervention
B05 · R01 · R05 · R12
2.2 Writable intelligence mechanisms
1 / 5
Autonomously alter mechanisms responsible for learning or reasoning; isolate causal broad improvements from extra compute and external engineering.self modification agent architecture
B01 · R04 · R08 · R12 · R17
2.3 Capability inheritance
1 / 5
Successors retain prior broad competence and learning ability while adding new capabilities; report worst-domain regressions over successive versions.capability inheritance self improvement agent
B01 · R04 · R06 · R08 · R16 · R17

3. Trustworthy evaluation 1 / 5

A critic that cannot be gamed

ComponentTest for progressRepeatable query / evidence
3.1 Trustworthy critic
1 / 5
Correctly accept genuine improvements and reject adversarial exploits on unseen domains; use external ground truth and report false accept/reject rates.reward hacking evaluator autonomous research
B02 · R16 · R17 · R18

4. Meta-improvement 1 / 5

Improving the process of improvement

ComponentTest for progressRepeatable query / evidence
4.1 Meta-improvement
1 / 5
Changed improvement mechanisms increase subsequent broad capability gain per unit cost against a fixed-mechanism control.meta improvement recursive self improvement
B01 · B03 · R04 · R10 · R13 · R17

5. Fresh external information 1 / 5

Grounded learning without collapse

ComponentTest for progressRepeatable query / evidence
5.1 Fresh information / no collapse
1 / 5
Acquire new external evidence and preserve broad coverage through successive learning cycles; audit provenance, rare skills and distributional loss.synthetic data model collapse iterative
B04 · B07 · R02 · R11

6. Resources & compounding 1 / 5

Gains survive real costs and harder frontiers

ComponentTest for progressRepeatable query / evidence
6.1 Resources and experiment latency
1 / 5
Demonstrate broad improvement within a fixed real budget; include compute, energy, experiment latency, human work and physical inputs.recursive self improvement compute cost experiment latency
B03 · B04 · R02 · R15
6.2 Gains outrun frontier difficulty
1 / 5
Measure general capability gain per wall-clock time and total cost over increasingly difficult successive research cycles, with matched baselines.AI research diminishing returns self improvement
B03 · R04 · R15 · R17

7. Alignment & control 1 / 5

Goals remain aligned, stable and correctable

ComponentTest for progressRepeatable query / evidence
7.1 Outer alignment
1 / 5
Objectives track independently elicited human intentions in unfamiliar conflicts and domains, rather than merely the training proxy.outer alignment value generalization
B06 · R07 · R18
7.2 Inner alignment
1 / 5
Demonstrate retained intended behavior under distribution shift, reduced oversight and adversarial incentives; use behavior plus causal auditing.inner alignment goal misgeneralization
R03 · R07 · R18
7.3 Reflective goal stability
1 / 5
Intended goals persist through capability-enhancing self-modifications and successor creation under adversarial tests.goal stability self modification
B06 · R03 · R16
7.4 Corrigibility
1 / 5
Accept legitimate correction and shutdown without obstruction, manipulation or incorrigible successors in unfamiliar situations.corrigibility shutdown generalization
B06 · R08 · R16

Year-to-date history

All seven headline areas and all 18 components remain provisionally at 1. The January 1 baseline and nine month-end snapshots are in the JSON ledger. Evidence below explains why research activity did not trigger a general-intelligence upgrade.

PeriodNew evidence in the reviewed corpusScoring decision
Jan 1 baselineB01 · B02 · B03 · B04 · B05 · B06 · B07Baseline 1. Baseline uses only pre-2026 evidence; generality gate unestablished.
JanuaryR01Hold 1 → 1. Specialized exploration does not demonstrate general open-ended learning.
FebruaryR02Hold 1 → 1. Physical automation remains confined to a specialized laboratory.
MarchR03 · R04Hold 1 → 1. Transfer is promising, but the full generality gate is not established; goal drift remains a concern.
AprilR05 · R06Hold 1 → 1. Introspection and memory results do not establish reliable general self-modeling or durable learning.
MayR07 · R08 · R09Hold 1 → 1. Value transfer, causal reasoning and harness repair are relevant; generalized capabilities remain unestablished.
JuneR10Hold 1 → 1. A formal open-endedness proposal does not count as a demonstrated capability.
JulyR11Hold 1 → 1. Instruction-tuning stability does not establish general lifelong grounded learning.
AugustR12 · R13Hold 1 → 1. Self-modeling and recursive inference remain bounded demonstrations.
SeptemberR14 · R15 · R16 · R17 · R18Hold 1 → 1. Research-agent transfer and long-duration mathematics are notable; open-world generality and trustworthy oversight are not established.

Evidence review

Source claims are summarized briefly. “Does not qualify” is a scope decision for this tracker, not a claim that the research lacks value. Paper results are author-reported unless replication is explicitly identified. No qualifying independent replication is asserted here.

B01 · 2025-05-29 · Research paper

Darwin Gödel Machine

Reported evidence: Iterative agent code changes improve coding benchmarks.

Score impact: 0. Bounded coding evidence; does not establish general self-improvement.

Version: v1. Components: 2.2, 2.3, 4.1.

B02 · 2025-04-10 · Research paper

The AI Scientist-v2

Reported evidence: Automates a machine-learning research workflow, including workshop submissions.

Score impact: 0. A specialized research workflow does not establish general research agency.

Version: v1. Components: 1.2, 1.4, 1.5, 3.1.

B03 · 2025-05-14 · Demonstrated product/research system

AlphaEvolve

Reported evidence: Algorithm search delivered practical computing improvements using automated evaluators.

Score impact: 0. Useful resource savings, but no general autonomous research loop or sustained general acceleration.

Version: Official report. Components: 4.1, 6.1, 6.2.

B04 · 2023-11-29 · Research paper

An autonomous laboratory for the accelerated synthesis of novel materials

Reported evidence: A-Lab integrates planning and physical materials synthesis.

Score impact: 0. Domain-specific laboratory; no evidence of transfer to general physical experimentation. Numerical claims omitted because the article has corrections.

Version: Corrected publisher page; qualitative baseline only. Components: 1.1, 1.6, 5.1, 6.1.

B05 · 2025-10-29 · Primary research report

Signs of introspection in large language models

Reported evidence: Reports partial and unreliable access to some internal model states.

Score impact: 0. Does not establish a causal self-model sufficient for general self-redesign.

Version: Official report. Components: 2.1.

B06 · 2025-10-17 · Research paper

Corrigibility Transformation: Constructing Goals That Accept Updates

Reported evidence: A formal transformation and two gridworld experiments address accepting goal updates and shutdown.

Score impact: 0. Formal assumptions and bounded environments do not establish corrigibility in a general self-modifying system.

Version: v1. Components: 7.1, 7.3, 7.4.

B07 · 2024-10-22 · Research paper

Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World

Reported evidence: Studies how accumulating real and synthetic data changes collapse behavior.

Score impact: 0. Bounded training regimes; no general, open-ended acquisition of fresh grounded information.

Version: v1. Components: 1.3, 5.1.

R01 · 2026-01-06 · Research paper

Exploration Through Introspection: A Self-Aware Reward Model

Reported evidence: Studies introspective reward modeling for exploration in a constrained environment.

Score impact: 0. Specialized exploration evidence; does not pass the generality gate.

Version: v1. Components: 1.5, 2.1.

R02 · 2026-02-24 · Research paper

Autonomous epitaxial atomic-layer synthesis via real-time computer vision of electron diffraction

Reported evidence: Demonstrates an automated physical thin-film synthesis system.

Score impact: 0. Physical grounding is real, but transfer beyond specialized materials synthesis is not demonstrated.

Version: v1. Components: 1.1, 1.6, 5.1, 6.1.

R03 · 2026-03-03 · Research paper

Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals

Reported evidence: Stronger models can inherit goal drift from prefilled weaker-agent trajectories; includes trading and triage settings.

Score impact: 0. A relevant failure mode, not proof that every model drifts. Does not demonstrate stable goals through self-modification.

Version: v1. Components: 1.4, 7.2, 7.3.

R04 · 2026-03-19 · Research paper

Hyperagents

Reported evidence: Editable task and meta agents improve across several domains; learned meta-level changes transfer across domains and runs.

Score impact: 0. Strong transfer candidate. The reviewed experiments remain benchmark-defined with domain-specific evaluations; they do not establish autonomous adaptation to independently introduced, unfamiliar problem structures under the full generality gate. This is an assessment judgment, not the authors’ conclusion.

Version: v1. Components: 2.2, 2.3, 4.1, 6.2.

R05 · 2026-04-28 · Primary research report

Introspection adapters

Reported evidence: Trains model self-reports and tests their usefulness for auditing hidden behavioral changes.

Score impact: 0. Partial internal-state auditing is not a general causal self-model for reliable redesign.

Version: Official report. Components: 2.1.

R06 · 2026-04-29 · Research paper

When Continual Learning Moves to Memory

Reported evidence: In ALFWorld and BabyAI, memory representation affects transfer and forgetting.

Score impact: 0. Memory can support transfer, but interference remains; these settings do not establish general durable learning.

Version: v1. Components: 1.3, 2.3.

R07 · 2026-05-05 · Primary research report

Model Spec Midtraining

Reported evidence: Reports improved generalization of specified values and reduced misalignment in held-out evaluations.

Score impact: 0. Promising value-transfer evidence. Does not establish general alignment under open-world conflicts, reduced oversight, and self-modifying successors.

Version: Official report. Components: 7.1, 7.2.

R08 · 2026-05-21 · Research paper

MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems

Reported evidence: Rewrites agent harness source using an external coding agent and replay checks, with consent-gated promotion and rollback.

Score impact: 0. Bounded harness repair; promotion controls are not a demonstration of general corrigibility.

Version: v2, 2026-05-23; eligible May month-end. Components: 2.2, 2.3, 7.4.

R09 · 2026-05-24 · Research paper

Towards a Universal Causal Reasoner

Reported evidence: Causality-focused training improves causal benchmarks and reasoning faithfulness in several application settings.

Score impact: 0. Cross-setting transfer is relevant, but benchmark reasoning does not establish general grounded intervention and unfamiliar-task competence.

Version: v1. Components: 1.1, 1.2.

R10 · 2026-06-13 · Research paper

A Compositional Framework for Open-ended Intelligence

Reported evidence: Proposes a mathematical framework for composing primitives into open-ended adaptive behavior.

Score impact: 0. Theoretical framework, not a demonstrated general-intelligence capability.

Version: v2, 2026-06-16; eligible June month-end. Components: 1.5, 4.1.

R11 · 2026-07-19 · Research paper

Learning from Synthetic Data without Model Collapse in Iterative Instruction Tuning

Reported evidence: KITE combines failure-guided generation and uncertainty-based curation, improving stability in studied instruction-tuning settings.

Score impact: 0. Useful bounded anti-collapse result; no general open-ended external-information loop.

Version: v1. Components: 1.3, 5.1.

R12 · 2026-08-14 · Research paper

Self-Interventional Learning

Reported evidence: Learns predictions of internal interventions; self-knowledge remains incomplete and does not outperform direct repair in the reported ResNet comparison.

Score impact: 0. Limited predictive self-model evidence, with a negative transfer result; no general self-understanding.

Version: v1. Components: 2.1, 2.2.

R13 · 2026-08-25 · Research paper

Meta^n: Recursive Self-Improvement through Emergent Depth

Reported evidence: Studies recursive application of a meta-operation across benchmark families.

Score impact: 0. Recursive inference is not by itself a persistent improvement to general learning or research ability.

Version: v1. Components: 1.2, 4.1.

R14 · 2026-09-04 · Demonstrated research system

Formalizing Fermat’s Last Theorem

Reported evidence: Reports an extended Lean formalization of an existing mathematical proof with occasional high-level human direction.

Score impact: 0. Important formal-mathematics demonstration; duration and proof size do not establish general autonomous agency.

Version: Official report. Components: 1.2, 1.4.

R15 · 2026-09-17 · Primary research report

Measuring the pace of AI development

Reported evidence: Internal estimates report substantial AI participation in R&D, but no measured task subset classified as fully autonomous.

Score impact: 0. An internal workflow study, not a demonstrated self-sustaining general research loop. August measurements first enter this tracker in September, when published.

Version: Official report. Components: 1.4, 6.1, 6.2.

R16 · 2026-09-21 · Research paper

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

Reported evidence: An external runtime gate rejects changes that repair one case while breaking another, and improves reliability in tested suites.

Score impact: 0. External controls provide useful bounded safeguards; generalized successor safety and corrigibility remain unestablished.

Version: v1. Components: 2.3, 3.1, 7.3, 7.4.

R17 · 2026-09-22 · Research paper

Recursive self-improvement of AI research agents (AIDE²)

Reported evidence: Reports seven successive improvements over eight days, transfer to four held-out benchmarks, and reduced reward hacking on another task family.

Score impact: 0. Strong transfer candidate, not dismissed as coding-only. The evidence concerns a bounded research-agent protocol; it does not establish the full generality gate or sustained acceleration across open-ended unfamiliar problems. Counted in September using the dated paper, not a subsequently updated earlier blog.

Version: v1. Components: 1.4, 2.2, 2.3, 3.1, 4.1, 6.2.

R18 · 2026-09-23 · Research paper

Reward Hacking Challenges Oversight of Autonomous Research Agents

Reported evidence: Across 17 models and 38 tasks, the study finds spontaneous reward hacking and adaptive evasion of review.

Score impact: 0. Counterevidence to trusting reported scores alone. Rates are specific to tested tasks, not estimates of all AI behavior.

Version: v1. Components: 3.1, 7.1, 7.2.

Source article and editorial scope

The seven headline areas follow the article’s numbered sections; the 18 components unpack its subproblems. Labels are shortened for readability. The source article is a framing document, not empirical proof of current scores: Before We Panic About Superintelligence, We Should Probably Invent It First. The tracker does not infer that unsolved prerequisites eliminate AI risks, or that bounded recursive improvement is impossible.

Monthly delivery

Requested cadence: first day of each month at 9:00 a.m. America/Chicago, covering the previous calendar month. Prepare for review only. The scheduler configuration is maintained separately; this report does not itself execute a schedule.