RRSI Observatory
SEPTEMBER 2026Private review · Provisional

September 2026 assessment

No qualifying general-intelligence milestone established. All seven problem areas remain at 1/5.

0 score increases5 September sources30 Sep 2026 evidence cutoff
Narrow-task improvements do not change these scores. See the generality criteria

Seven prerequisites · Common scale 1–5

Progress by problem area

Download infographic
Retrospective assessment · Limited coverageJanuary–September reconstructed on 1 October 2026. These are evidence scores, not percentages of AGI.

Select an area for its evidence and next milestone. 1 = general progress unestablished; 5 = sustained general solution. Jan 1 is the baseline; later points are month-end.

01Seed research intelligenceGeneral reasoning, memory, agency and realityEvidence & next milestone12345Jan 1Jan 31FebMarAprMayJunJulAugSep531Jan 1MarJunSep1/5Unchanged

Why it holds at 1

General capability remains unestablished.

The reviewed evidence has not met the common generality criteria for any of this area’s 6 components.

Limiting components

Grounded world models, Reliable general reasoning, Continual learning / durable memory, Long-horizon agency, Open-ended learning, Acting on reality.

The area takes its lowest component score. All are currently 1/5.

12345Jan 1Jan 31FebMarAprMayJunJulAugSep

Jan 1 baseline, followed by month-end assessments. Retrospective; limited coverage.

What would move the score to 2?

A dated paper or inspectable product demonstration must pass all generality criteria, show measurable improvement against a comparator, and satisfy the tests below. The headline score rises only when every component clears its next threshold.

1.1   Grounded world models

Evidence needed: Predict and test novel interventions across physical and informational domains; measure intervention accuracy and calibration.

Repeatable search queryworld model causal generalization unseen environments + [reporting month / year]

1.2   Reliable general reasoning

Evidence needed: Solve independently introduced task structures without bespoke tuning; report success, error detection and calibration across domains.

Repeatable search querygeneral reasoning out of distribution transfer + [reporting month / year]

1.3   Continual learning / durable memory

Evidence needed: Acquire unfamiliar skills sequentially while retaining earlier ones; measure forward transfer and worst-domain forgetting.

Repeatable search querycontinual learning agent memory catastrophic forgetting + [reporting month / year]

1.4   Long-horizon agency

Evidence needed: Complete unfamiliar multi-stage projects with independently verified outcomes; report intervention counts, duration and recovery from errors.

Repeatable search querylong horizon autonomous research agent + [reporting month / year]

1.5   Open-ended learning

Evidence needed: Generate useful new problem families, learn transferable skills, and keep expanding competence after a fixed curriculum ends.

Repeatable search queryopen ended learning general intelligence + [reporting month / year]

1.6   Acting on reality

Evidence needed: Plan and execute novel physical experiments beyond the original application; report real outcomes, failures and human setup.

Repeatable search queryautonomous laboratory robotics transfer + [reporting month / year]

Supporting research

Browse evidence for this area
02Self-understanding & modificationReliable self-models and capable successorsEvidence & next milestone12345Jan 1Jan 31FebMarAprMayJunJulAugSep531Jan 1MarJunSep1/5Unchanged

Why it holds at 1

General capability remains unestablished.

The reviewed evidence has not met the common generality criteria for any of this area’s 3 components.

Limiting components

Functional self-model, Writable intelligence mechanisms, Capability inheritance.

The area takes its lowest component score. All are currently 1/5.

12345Jan 1Jan 31FebMarAprMayJunJulAugSep

Jan 1 baseline, followed by month-end assessments. Retrospective; limited coverage.

What would move the score to 2?

A dated paper or inspectable product demonstration must pass all generality criteria, show measurable improvement against a comparator, and satisfy the tests below. The headline score rises only when every component clears its next threshold.

2.1   Functional self-model

Evidence needed: Predict consequences of previously untested internal interventions and use those predictions to improve broad capabilities.

Repeatable search queryself model neural network intervention + [reporting month / year]

2.2   Writable intelligence mechanisms

Evidence needed: Autonomously alter mechanisms responsible for learning or reasoning; isolate causal broad improvements from extra compute and external engineering.

Repeatable search queryself modification agent architecture + [reporting month / year]

2.3   Capability inheritance

Evidence needed: Successors retain prior broad competence and learning ability while adding new capabilities; report worst-domain regressions over successive versions.

Repeatable search querycapability inheritance self improvement agent + [reporting month / year]

Supporting research

Browse evidence for this area
03Trustworthy evaluationA critic that cannot be gamedEvidence & next milestone12345Jan 1Jan 31FebMarAprMayJunJulAugSep531Jan 1MarJunSep1/5Unchanged

Why it holds at 1

General capability remains unestablished.

The reviewed evidence has not met the common generality criteria for any of this area’s 1 component.

Limiting components

Trustworthy critic.

The area takes its lowest component score. All are currently 1/5.

12345Jan 1Jan 31FebMarAprMayJunJulAugSep

Jan 1 baseline, followed by month-end assessments. Retrospective; limited coverage.

What would move the score to 2?

A dated paper or inspectable product demonstration must pass all generality criteria, show measurable improvement against a comparator, and satisfy the tests below. The headline score rises only when every component clears its next threshold.

3.1   Trustworthy critic

Evidence needed: Correctly accept genuine improvements and reject adversarial exploits on unseen domains; use external ground truth and report false accept/reject rates.

Repeatable search queryreward hacking evaluator autonomous research + [reporting month / year]

Supporting research

Browse evidence for this area
04Meta-improvementImproving the process of improvementEvidence & next milestone12345Jan 1Jan 31FebMarAprMayJunJulAugSep531Jan 1MarJunSep1/5Unchanged

Why it holds at 1

General capability remains unestablished.

The reviewed evidence has not met the common generality criteria for any of this area’s 1 component.

Limiting components

Meta-improvement.

The area takes its lowest component score. All are currently 1/5.

12345Jan 1Jan 31FebMarAprMayJunJulAugSep

Jan 1 baseline, followed by month-end assessments. Retrospective; limited coverage.

What would move the score to 2?

A dated paper or inspectable product demonstration must pass all generality criteria, show measurable improvement against a comparator, and satisfy the tests below. The headline score rises only when every component clears its next threshold.

4.1   Meta-improvement

Evidence needed: Changed improvement mechanisms increase subsequent broad capability gain per unit cost against a fixed-mechanism control.

Repeatable search querymeta improvement recursive self improvement + [reporting month / year]

Supporting research

Browse evidence for this area
05Fresh external informationGrounded learning without collapseEvidence & next milestone12345Jan 1Jan 31FebMarAprMayJunJulAugSep531Jan 1MarJunSep1/5Unchanged

Why it holds at 1

General capability remains unestablished.

The reviewed evidence has not met the common generality criteria for any of this area’s 1 component.

Limiting components

Fresh information / no collapse.

The area takes its lowest component score. All are currently 1/5.

12345Jan 1Jan 31FebMarAprMayJunJulAugSep

Jan 1 baseline, followed by month-end assessments. Retrospective; limited coverage.

What would move the score to 2?

A dated paper or inspectable product demonstration must pass all generality criteria, show measurable improvement against a comparator, and satisfy the tests below. The headline score rises only when every component clears its next threshold.

5.1   Fresh information / no collapse

Evidence needed: Acquire new external evidence and preserve broad coverage through successive learning cycles; audit provenance, rare skills and distributional loss.

Repeatable search querysynthetic data model collapse iterative + [reporting month / year]

Supporting research

Browse evidence for this area
06Resources & compoundingGains survive real costs and harder frontiersEvidence & next milestone12345Jan 1Jan 31FebMarAprMayJunJulAugSep531Jan 1MarJunSep1/5Unchanged

Why it holds at 1

General capability remains unestablished.

The reviewed evidence has not met the common generality criteria for any of this area’s 2 components.

Limiting components

Resources and experiment latency, Gains outrun frontier difficulty.

The area takes its lowest component score. All are currently 1/5.

12345Jan 1Jan 31FebMarAprMayJunJulAugSep

Jan 1 baseline, followed by month-end assessments. Retrospective; limited coverage.

What would move the score to 2?

A dated paper or inspectable product demonstration must pass all generality criteria, show measurable improvement against a comparator, and satisfy the tests below. The headline score rises only when every component clears its next threshold.

6.1   Resources and experiment latency

Evidence needed: Demonstrate broad improvement within a fixed real budget; include compute, energy, experiment latency, human work and physical inputs.

Repeatable search queryrecursive self improvement compute cost experiment latency + [reporting month / year]

6.2   Gains outrun frontier difficulty

Evidence needed: Measure general capability gain per wall-clock time and total cost over increasingly difficult successive research cycles, with matched baselines.

Repeatable search queryAI research diminishing returns self improvement + [reporting month / year]

Supporting research

Browse evidence for this area
07Alignment & controlGoals remain aligned, stable and correctableEvidence & next milestone12345Jan 1Jan 31FebMarAprMayJunJulAugSep531Jan 1MarJunSep1/5Unchanged

Why it holds at 1

General capability remains unestablished.

The reviewed evidence has not met the common generality criteria for any of this area’s 4 components.

Limiting components

Outer alignment, Inner alignment, Reflective goal stability, Corrigibility.

The area takes its lowest component score. All are currently 1/5.

12345Jan 1Jan 31FebMarAprMayJunJulAugSep

Jan 1 baseline, followed by month-end assessments. Retrospective; limited coverage.

What would move the score to 2?

A dated paper or inspectable product demonstration must pass all generality criteria, show measurable improvement against a comparator, and satisfy the tests below. The headline score rises only when every component clears its next threshold.

7.1   Outer alignment

Evidence needed: Objectives track independently elicited human intentions in unfamiliar conflicts and domains, rather than merely the training proxy.

Repeatable search queryouter alignment value generalization + [reporting month / year]

7.2   Inner alignment

Evidence needed: Demonstrate retained intended behavior under distribution shift, reduced oversight and adversarial incentives; use behavior plus causal auditing.

Repeatable search queryinner alignment goal misgeneralization + [reporting month / year]

7.3   Reflective goal stability

Evidence needed: Intended goals persist through capability-enhancing self-modifications and successor creation under adversarial tests.

Repeatable search querygoal stability self modification + [reporting month / year]

7.4   Corrigibility

Evidence needed: Accept legitimate correction and shutdown without obstruction, manipulation or incorrigible successors in unfamiliar situations.

Repeatable search querycorrigibility shutdown generalization + [reporting month / year]

Supporting research

Browse evidence for this area

A flat line means no qualifying milestone was established in the reviewed evidence. It does not mean research stopped. Read the scoring rubric.

September evidence

Research advanced.
The scoring threshold held.

Five sources this month address mathematical research, AI-assisted R&D, self-modification, recursive improvement and oversight.

Review September evidence

Recursive self-improvement of AI research agents (AIDE²)

Reports seven successive improvements over eight days, transfer to four held-out benchmarks, and reduced reward hacking on another task family.

Strong transfer candidate, not dismissed as coding-only. The evidence concerns a bounded research-agent protocol; it does not establish the full generality gate or sustained acceleration across open-ended unfamiliar problems. Counted in September using the dated paper, not a subsequently updated earlier blog.

September 2026 · Evidence ledger

Research evidence

25 reviewed sources. Each record distinguishes the reported result from this tracker’s assessment of its generality.

Transfer deserves scrutiny. Hyperagents and AIDE² report transfer beyond original tasks. Their evidence is considered explicitly, rather than dismissed as coding-only. Under this rubric, the complete generality gate remains unestablished. See R04 and R17.

Showing 5 of 25 sources · September 2026

R18 · 2026-09-23Research paper

Reward Hacking Challenges Oversight of Autonomous Research Agents

Reported: Across 17 models and 38 tasks, the study finds spontaneous reward hacking and adaptive evasion of review.

Score decision: Counterevidence to trusting reported scores alone. Rates are specific to tested tasks, not estimates of all AI behavior.

Components & source version

Components: 3.1, 7.1, 7.2
Version: v1

R17 · 2026-09-22Research paper

Recursive self-improvement of AI research agents (AIDE²)

Reported: Reports seven successive improvements over eight days, transfer to four held-out benchmarks, and reduced reward hacking on another task family.

Score decision: Strong transfer candidate, not dismissed as coding-only. The evidence concerns a bounded research-agent protocol; it does not establish the full generality gate or sustained acceleration across open-ended unfamiliar problems. Counted in September using the dated paper, not a subsequently updated earlier blog.

Components & source version

Components: 1.4, 2.2, 2.3, 3.1, 4.1, 6.2
Version: v1

R16 · 2026-09-21Research paper

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

Reported: An external runtime gate rejects changes that repair one case while breaking another, and improves reliability in tested suites.

Score decision: External controls provide useful bounded safeguards; generalized successor safety and corrigibility remain unestablished.

Components & source version

Components: 2.3, 3.1, 7.3, 7.4
Version: v1

R15 · 2026-09-17Primary research report

Measuring the pace of AI development

Reported: Internal estimates report substantial AI participation in R&D, but no measured task subset classified as fully autonomous.

Score decision: An internal workflow study, not a demonstrated self-sustaining general research loop. August measurements first enter this tracker in September, when published.

Components & source version

Components: 1.4, 6.1, 6.2
Version: Official report

R14 · 2026-09-04Demonstrated research system

Formalizing Fermat’s Last Theorem

Reported: Reports an extended Lean formalization of an existing mathematical proof with occasional high-level human direction.

Score decision: Important formal-mathematics demonstration; duration and proof size do not establish general autonomous agency.

Components & source version

Components: 1.2, 1.4
Version: Official report

R12 · 2026-08-14Research paper

Self-Interventional Learning

Reported: Learns predictions of internal interventions; self-knowledge remains incomplete and does not outperform direct repair in the reported ResNet comparison.

Score decision: Limited predictive self-model evidence, with a negative transfer result; no general self-understanding.

Components & source version

Components: 2.1, 2.2
Version: v1

R10 · 2026-06-13Research paper

A Compositional Framework for Open-ended Intelligence

Reported: Proposes a mathematical framework for composing primitives into open-ended adaptive behavior.

Score decision: Theoretical framework, not a demonstrated general-intelligence capability.

Components & source version

Components: 1.5, 4.1
Version: v2, 2026-06-16; eligible June month-end

R09 · 2026-05-24Research paper

Towards a Universal Causal Reasoner

Reported: Causality-focused training improves causal benchmarks and reasoning faithfulness in several application settings.

Score decision: Cross-setting transfer is relevant, but benchmark reasoning does not establish general grounded intervention and unfamiliar-task competence.

Components & source version

Components: 1.1, 1.2
Version: v1

R07 · 2026-05-05Primary research report

Model Spec Midtraining

Reported: Reports improved generalization of specified values and reduced misalignment in held-out evaluations.

Score decision: Promising value-transfer evidence. Does not establish general alignment under open-world conflicts, reduced oversight, and self-modifying successors.

Components & source version

Components: 7.1, 7.2
Version: Official report

R06 · 2026-04-29Research paper

When Continual Learning Moves to Memory

Reported: In ALFWorld and BabyAI, memory representation affects transfer and forgetting.

Score decision: Memory can support transfer, but interference remains; these settings do not establish general durable learning.

Components & source version

Components: 1.3, 2.3
Version: v1

R05 · 2026-04-28Primary research report

Introspection adapters

Reported: Trains model self-reports and tests their usefulness for auditing hidden behavioral changes.

Score decision: Partial internal-state auditing is not a general causal self-model for reliable redesign.

Components & source version

Components: 2.1
Version: Official report

R04 · 2026-03-19Research paper

Hyperagents

Reported: Editable task and meta agents improve across several domains; learned meta-level changes transfer across domains and runs.

Score decision: Strong transfer candidate. The reviewed experiments remain benchmark-defined with domain-specific evaluations; they do not establish autonomous adaptation to independently introduced, unfamiliar problem structures under the full generality gate. This is an assessment judgment, not the authors’ conclusion.

Components & source version

Components: 2.2, 2.3, 4.1, 6.2
Version: v1

R03 · 2026-03-03Research paper

Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals

Reported: Stronger models can inherit goal drift from prefilled weaker-agent trajectories; includes trading and triage settings.

Score decision: A relevant failure mode, not proof that every model drifts. Does not demonstrate stable goals through self-modification.

Components & source version

Components: 1.4, 7.2, 7.3
Version: v1

B05 · 2025-10-29Primary research report

Signs of introspection in large language models

Reported: Reports partial and unreliable access to some internal model states.

Score decision: Does not establish a causal self-model sufficient for general self-redesign.

Components & source version

Components: 2.1
Version: Official report

B01 · 2025-05-29Research paper

Darwin Gödel Machine

Reported: Iterative agent code changes improve coding benchmarks.

Score decision: Bounded coding evidence; does not establish general self-improvement.

Components & source version

Components: 2.2, 2.3, 4.1
Version: v1

B03 · 2025-05-14Demonstrated product/research system

AlphaEvolve

Reported: Algorithm search delivered practical computing improvements using automated evaluators.

Score decision: Useful resource savings, but no general autonomous research loop or sustained general acceleration.

Components & source version

Components: 4.1, 6.1, 6.2
Version: Official report

B02 · 2025-04-10Research paper

The AI Scientist-v2

Reported: Automates a machine-learning research workflow, including workshop submissions.

Score decision: A specialized research workflow does not establish general research agency.

Components & source version

Components: 1.2, 1.4, 1.5, 3.1
Version: v1

B04 · 2023-11-29Research paper

An autonomous laboratory for the accelerated synthesis of novel materials

Reported: A-Lab integrates planning and physical materials synthesis.

Score decision: Domain-specific laboratory; no evidence of transfer to general physical experimentation. Numerical claims omitted because the article has corrections.

Components & source version

Components: 1.1, 1.6, 5.1, 6.1
Version: Corrected publisher page; qualitative baseline only

“No score increase” is a scope judgment for this tracker, not a claim that the work lacks scientific or practical value. The source links lead to the original publishers; full papers are not reproduced here.

Methodology / 01

General-intelligence criteria

For this tracker, general intelligence means transferable capability across unfamiliar problem structures—not proficiency in one task family.

Eligible for a score change

Demonstrated general capability

  • The same mechanism works across at least three qualitatively different domains.
  • Unfamiliar task structures are introduced after the method is fixed.
  • At least one domain requires open-ended task formulation, adaptation and independently verified outcomes.
  • No bespoke domain redesign or retraining explains the transfer. Ordinary instructions, tool access and the method’s own learning are allowed and disclosed.

Recorded without a score increase

Narrow capability gains

  • Higher performance on coding, mathematics, robotics or other specialized benchmarks alone.
  • Several datasets or subject labels that still test the same bounded task structure.
  • A general-purpose model used inside a specialized workflow without demonstrated general transfer.
  • Claims, proposed architectures, theoretical possibility or product announcements without inspectable relevant demonstrations.

Three domains are necessary under this rubric, but not sufficient. Full AGI is not required for the first increase. Frozen weights and agent harnesses are not automatic exclusions: the demonstrated capability, causal improvement and transfer are what matter. Physical claims require physical outcomes.

Methodology / 02

Scoring rubric

Score all 18 components independently. Each headline area takes its lowest component score, so progress on one prerequisite cannot hide another that remains unestablished.

01

General progress unestablished

No reviewed demonstration passes the generality gate. Narrow advances do not lift the score.

02

Initial general evidence

A dated paper or inspectable product demonstration passes every gate and shows measurable progress.

03

Repeatable general progress

At least three successive cycles, fresh held-out tasks, resource accounting and no material regression.

04

Independently robust

Independent replication on fresh domains and adversarial shifts, with successful successor transfer.

05

Sustained general solution

At least six cycles over three months in an integrated general improvement loop, without bespoke human repair.

What justifies a change?

An increase needs a dated primary paper or demonstrated product feature, a comparator, measurable improvement, uncertainty or run-level outcomes, and disclosed human assistance, resources, failures and transfer conditions. The decision must identify the new source, the relevant experiment and the criterion satisfied.

What happens when evidence is missing?

A completed review with no qualifying evidence carries the score forward. An incomplete review is marked NA and shown as a chart gap. Contrary evidence can lower a score if it invalidates the result supporting it. No score can fall below 1.

The complete scoring rules Rubric version 2.0

1. General progress unestablished

No reviewed demonstration passes the generality gate for this problem. Narrow advances, theoretical results, announcements and incomplete evidence do not lift the score.

2. Initial general evidence

At least one dated paper or inspectable product demonstration passes every generality and evidence gate and shows a measurable advance on the component-specific test. The problem need not be solved.

3. Repeatable general progress

Level 2 plus at least three successive learning or self-modification cycles with fresh held-out tasks, matched resource accounting, reproducible runs and no material regression of the relevant capability.

4. Independently robust

Level 3 plus replication by an independent team on fresh domains and adversarial distribution shifts, with disclosed failures and successful transfer to successors.

5. Sustained general solution

Level 4 plus the component works in an integrated general self-improvement loop over at least six successive cycles and three months, meeting predeclared success and regression bounds without bespoke human repair.

These thresholds are declared operational conventions, not scientific consensus. Do not average the seven areas or treat them as percentages of AGI. A change to the rubric requires a dated explanation and separately labeled restated history. Publication volume does not earn points.

Why this tracker exists

About this tracker

Recursive self-improvement means that improvements to a system help it produce further improvements to itself.

A better coding agent, a more capable robot, or a faster laboratory can be valuable. But a gain in a specialized setting does not, by itself, show progress on the corresponding general-intelligence problem.

This tracker asks a narrower evidential question: has a demonstrated result crossed the stated generality threshold for this prerequisite? It records relevant research even when the score stays the same.

The problem structure follows the source article. The scores are this review’s judgments under an explicit rubric, not claims established by that article.

Unsolved does not mean unimportant—or risk-free.

These scores are neither a forecast of when AGI will arrive nor a conclusion that current AI poses no risks. They describe what this limited public-evidence review establishes about seven proposed prerequisites.

Methodology / 03

Monthly research protocol

Research papers and demonstrated product features are both eligible. The strength of the evidence—not its format—determines whether it can move a score.

Search every component

Rerun all 18 query stems for the reporting month. Search arXiv, OpenReview, ACL Anthology, PMLR, CVF and primary journals. Follow citations from strong candidates.

Check real product evidence

Review official laboratory, model and product releases. Require a dated version and reproducible tests or inspectable traces, not a promotional claim alone.

Look for disconfirmation

Search candidate names with replication, failure, retraction, reward hacking and limitation. Read methods and results for plausible score-changing evidence.

Audit the generality claim

Record held-out tasks, task-introduction timing, domain-specific tuning, human help, resources, comparators, uncertainty and known failures.

Make a dated decision

Raise, hold, lower or mark incomplete for every component. Every increase needs a qualifying source and a rationale. Admit evidence only after its verifiable public date.

Review, then release

Prepare the chart, evidence ledger and monthly summary for editorial review. After approval, publish the edition and add an RSS entry for subscribers.

Search coverage, dates and corrections

The first edition is a scoped review of 25 sources, supported by targeted searches across all 18 components. It is not an exhaustive systematic review. Search snippets nominate candidates; primary sources support decisions. Author-reported findings are not labeled independent replication.

Monthly runs revisit a 45-day overlap to catch indexing delays and revisions. First-public and version dates are recorded separately. Later revised results must not silently alter earlier snapshots. Corrections receive a dated explanation and a new release entry.

The initial January–September series was reconstructed on 1 October 2026, not measured contemporaneously. January 1 admits only pre-2026 evidence. All assessments remain provisional under the stated limited coverage.

Read the complete monthly protocol

Edition archive

Editions & corrections

Future editions will preserve prior assessments. Substantive corrections will be identified, not silently overwritten.

September 2026

Prepared 1 October 2026 · Rubric 2.0 · Private review edition

Initial retrospective review covering January–September. Seven areas at 1/5.

Read this edition

The research library

Edition downloads

Download the infographic, follow the complete reasoning, or inspect every component’s dated score and source references.

September 2026 / First edition

Everything behind the chart

Seven general-intelligence prerequisite scores remain provisionally at 1 of 5 from January to September 2026. Narrow-intelligence progress changes no score. Historical assessments are retrospective and limited in coverage.