STUDY AREA · OPEN RESOURCES

RSI: AI that helps improve AI

How does a system propose changes, test results and retain what works? Explore recursive self-improvement through references, examples and criteria for assessing claims.

Improving an answer, improving an agent and improving the ability to create new agents are different things.

Start here

COURSE · PT / EN / ES

RSI v6.2

18 lessons in six modules, with exercises, review and materials for applying LOOP-R and recording approved knowledge. Available in Portuguese, English and Spanish.

Open the course in English

PROJECT · PT / EN / ES

RSI guide

A map of the topic, mechanisms, applications and limitations. The project collects research and educational content; it is not a ready-to-run autonomous RSI system.

Open the guide

Alongside the reference guide and RSI v6.2 course, the collection includes the Copiloto and Dream-RSI projects and the Alerta IA 2028 course.

More RSI projects and courses

Explore applications, research and critical reading. Each resource states what is available and in which language.

RESEARCH · GUIDE PT / EN / ES

Dream-RSI · google-rsi

Independent educational guide on learning from experiment histories. Includes five figures from the authors, an interactive comparison and a comprehension activity. The lab and eight sessions are future plans, not an implementation of the paper.

Explore Dream-RSI

COURSE · PT / EN / ES

Alerta IA 2028

Course and introductory explainer in three tracks: the improvement cycle, evidence and evaluation limits. Presents 2028 scenarios as hypotheses to examine, without a guaranteed timeline.

Open Alerta IA 2028

FRAMEWORK AND COURSE · PT / EN / ES

LOOP‑R

The Execute → Measure → Critique → Propose → Test → Validate → Promote → Repeat cycle, ready to use in Claude Code: nine assistants with separate roles, a record of every version, a cost ceiling and a command to roll back. No worse version replaces the current one by the system's own decision. The course has five tracks and 21 lessons for owners and managers without a technical background.

Open the LOOP-R framework

Take the LOOP-R course

What changes at each level

Revise the answer

The model critiques and rewrites an output. This can improve a task without changing its parameters or development process.

Improve the system

Instructions, memory, tools or agent code change. Compare versions using held-out tasks and preserve the ability to roll back.

Investigate recursion

The improvement helps produce future improvements. This mechanism requires evidence: more attempts or a higher score do not demonstrate unlimited growth.

From an idea to a verifiable test

Start with a small task: answer questions using a policy, create questions from a text or extract action items from notes. The course organizes this work using LOOP-R, a conceptual proposal from the project.

  1. Define the task, the correct reference and limits on data, spending and actions.
  2. Record the initial version and separate development cases from cases held out for validation.
  3. Propose one change at a time and compare using the same criteria, including errors, cost and review effort.
  4. Validate before adopting. Keep results and versions, define who decides and maintain a rollback path.

A known evaluation can be exploited. Preserve test independence, check for fabricated data and examine out-of-sample results. Do not confuse local gains with general autonomy.

References for deeper study

Selected from the course research, consulted on September 25, 2026. Links lead to the original publications in English. Results apply to the study conditions; they were not reproduced in this project.

  • Self-Refine

    Generation, critique and iterative refinement of answers, without requiring additional training in the proposed procedure.

  • Reflexion

    Verbal feedback and episodic memory guide later attempts without updating model weights.

  • Darwin Gödel Machine

    Agents that modify their code and maintain an archive of variants. Experimental evidence from specific evaluations.

  • AlphaEvolve · Google DeepMind

    Program search with automated evaluation. Component improvements are not a universal measure of intelligence.

  • Self-Rewarding Language Models

    The model participates in generating rewards during iterative training. This differs from requesting a critique in a conversation.

  • Reward tampering · Anthropic

    Experiments on reward tampering highlight the need to protect the evaluation process.

  • The Curse of Recursion

    Studies degradation under training conditions involving generated data. This does not imply that all synthetic data is harmful.