Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement
The rapid advancement of large language models (LLMs) has significantly enhanced the capabilities of agents across various tasks. However, existing agentic systems, whether based on fixed pipeline algorithms or pre-defined meta-learning frameworks, cannot search the whole agent design space due to the restriction of human-designed components, and thus might miss the more optimal agent design. In this paper, we introduce G ̈odel Agent, a selfevolving framework inspired by the G ̈odel machine, enabling agents to recursively improve themselves without relying on predefined routines or fixed optimization algorithms. G ̈odel Agent leverages LLMs to dynamically modify its own logic and behavior, guided solely by high-level objectives through prompting. Experimental results on multiple domains demonstrate that implementation of G ̈odel Agent can achieve continuous self-improvement, surpassing manually crafted agents in performance, efficiency, and generalizability.
Introduction. As large language models (LLMs) (OpenAI et al., 2024; Dubey et al., 2024) demonstrate increasingly strong reasoning and planning capabilities, LLMdriven agentic systems have achieved remarkable performance in a wide range of tasks (Wang et al., 2024a). Substantial effort has been invested in manually designing sophisticated agentic systems using human priors in different application areas. Recently, there has been a significant interest in creating self-evolving agents, that not only greatly reduce human labor but also produce better solutions. Given that human effort can only cover a small search space of agent design, it is reasonable to expect that a self-evolving agent with the freedom to explore the full design space has the potential to produce a more optimal solution. There is a large body of work proposing agents capable of self-refinement. Some agents are designed to iterate over a fixed routine consisting of a list of fixed modules, while some of the modules are capable of taking self- or environment feedback to refine their actions (Chen et al., 2023b; Qu et al., 2024a; Tang et al., 2025). This type of agent, referred to as Hand-Designed Agent, is depicted as having the lowest degree of freedom in Figure 2. More automated agents have been designed to be able to update their routines or modules in some pre-defined meta-learning routine, for example, natural language gradients (Zhou et al., 2024), meta agent (Hu et al., 2024), or creating and collecting demonstrations (Khattab et al., 2023). This type of agent, known as Meta-Learning Optimized Agents, is depicted as having the middle degree of freedom in Figure 2. However, there are inevitably some human priors involved in these agent designs that cannot be improved during the inference time. In this paper, we propose G ̈odel Agent to eliminate the human design prior, which is an automated LLM agent that can freely decide its own routine, modules, and even the way to update them. It is inspired by the self-referential G ̈odel machine (Schmidhuber, 2003), which was proven to be able to find the global optimal solutions. Self-reference means the property of a system that can analyze and modify its own code, including the parts responsible for the analysis and modification processes (Astrachan, 1994). Therefore, it can achieve what’s known as ”recursive self-improvement”, where it iteratively updates itself to become more efficient and effective at achieving its predefined goals. In this case, as shown in Figure 1, G ̈odel Agent can analyze and modify its own code, including the code for analyzing and modifying itself, and thus can search the full agent design space, which is depicted as having the highest degree of freedom in Figure 2. G ̈odel Agent can theoretically make increasingly better modifications over time through recursively self-update (Wang, 2018). In this paper, we choose to implement it by letting it manipulate its own runtime memory, i.e., the agent is able to retrieve its current code in the runtime memory and modify it by monkey patching (Bimal, 2012), which dynamically modifies classes or modules during execution. To allow it to update the logic of the running main function, unlike the loop-iterative approach of traditional agents, we implement the main function as a recursive function. In this function, LLM analyzes and makes a series of decisions, including reading and modifying its own code from runtime memory (self-awareness1 and self-modification), and interacting with the environment to gather feedback. The agent then proceeds to the subsequent recursive depth and continues to optimize itself. To validate the effectiveness of G ̈odel Agent, we conduct experiments on multiple domains including coding, science, math, and reasoning. Our results demonstrate that G ̈odel Agent achieves significant performance gain across various tasks, surpassing various widely-used agents that require human design. The same implementation of G ̈odel Agent can easily adapt to different tasks by only specifying the environment description and feedback mechanism. Additionally, the case study of the optimization progress reveals that G ̈odel Agent can provide novel insights into agent design. Our codes are released to facilitate future research2. In summary, our contributions are as follows: • We propose the first fully self-referential agent framework, G ̈odel Agent, and implement it using monkey patching. It autonomously engages in self-awareness, self-modification, and recursive self-improvement.
• Experiments shows that G ̈odel Agent is superior to the previous agent frameworks in terms of performance, flexibility, cost, and potential.
• We analyze G ̈odel Agent ’s optimization process, including its self-referential abilities and the optimized agentic systems, aiming to deepen our understanding of both LLMs and agents.
• Our framework offers a promising direction for developing flexible and capable agents through recursive self-improvement.
Related work. Hand-Designed Agent Systems Researchers have designed numerous agent systems tailored to various tasks based on predefined heuristics and prior knowledge. These systems often employ techniques such as prompt engineering (Chen et al., 2023a; Schulhoff et al., 2024), chain-of-thought reasoning and planning (Wei et al., 2022; Yao et al., 2022), as well as reflection (Shinn et al., 2024; Madaan et al., 2024), code generation (Wang et al., 2023a; Vemprala et al., 2024), tool use (Nakano et al., 2021; Qu et al., 2024a), retrieval-augmented generation (Lewis et al., 2020; Zhang et al., 2024b), and multi-agent collaboration (Xu et al., 2023; Wu et al., 2023; Qian et al., 2023; Hong et al., 2023). Once crafted by human designers, these systems remain static and do not adapt or evolve over time. Meta-Learning Optimized Agent Systems Some researchers have explored methods for enhancing agents through fixed learning algorithms (Zhou et al., 2024; Hu et al., 2024). For example, certain frameworks store an agent’s successful or failed strategies in memory based on environmental feedback (Liu et al., 2023; Hu et al., 2023; Qian et al., 2024), while others automatically optimize agent prompts (Khattab et al., 2023; Zhang et al., 2024a; Khattab et al., 2023). Some studies focus on designing prompts that enable agents to autonomously refine specific functions (Zhang et al.). However, these meta-algorithms are also designed manually and remain unchanged once deployed, limiting the agents’ ability. Recursive Self-Improvement The concept of recursive self-improvement has a long history (Good, 1966; Schmidhuber, 1987). G ̈odel machine (Schmidhuber, 2003) introduced the notion of a proof searcher that executes a selfmodification, thereby enabling the machine to enhance itself. In the early days, there were also some discussions of self-improving agents that were not based on LLM (Hall, 2007; Steunebrink and Schmidhuber, 2012). More recently, Zelikman et al. (2023) applied recursive self-improvement to code generation, where the target of improvement was the optimizer itself. Some work (Havrilla et al., 2024; Qu et al., 2024b; Kumar et al., 2024) also explores recursive self-improvement by finetuning models to introspect and correct previous mistakes. G ̈odel Agent represents the first selfreferential agent based on LLM. This approach is more flexible, removing human-designed constraints.
Method. In this section, we first describe the formal definitions for previous agent methods with a lower degree of freedom, including hand-design and metalearning optimized agents, as a background. Then we introduce our proposed G ̈odel Agent, a selfreferential agent that can recursively update its own code, evolving over training. Let E ∈S denote a specific environment state, where S denotes the set of all possible environments the agent will encounter. For example, an environment can be a mathematical problem with ground truth solutions. We denote the policy that an agent follows to solve a problem in the current environment by π ∈Π, where Π is the set of all possible policies the agent can follow. A hand-designed agent, as shown in the left panel of Figure 2, is not capable of updating its policy and following the same policy π all the time, regardless of environmental feedback. In contrast, a meta-learning optimized agent updates its policy based on a meta-learning algorithm I at training time based on the feedback it receives from the environment, as shown in the middle panel of Figure 2. The environment feedback is usually defined as a utility function U : S×Π →R, which maps an environment and a policy to a realvalued performance score. The main training algorithm of a meta-learning optimized agent can then be written as follows: recursive design of the agent requires the specification of an initial agent algorithm (π0, I0), detailed as follows:
• A initial agent policy π0 to perform the desired task within the environment E. For example, it can be chain-of-thought prompting of an LLM.
• A self-referential learning algorithm I0 for recursively querying an LLM to rewrite its own code based on the environmental feedback.
We then further specify a possible initialization of the self-referential learning algorithm I0 = (f0, o0), using a mutual recursion between a decision-making function f0, and an action function o0:
• The decision-making function f0, implemented by an LLM, determines a sequence of appropriate actions a1, a2, ..., an ∈A based on the current environment E, the agent’s algorithm (πt, It), and the goal g.
• The action function o0, executes the selected action and updates the agent’s policy accordingly.
The set of actions A for the action function o to execute needs to include the following four actions:
• self inspect: Introspect and read the agent’s current algorithm (πt, It).
• interact: Interact with the environment by calling the utility function U to assess the performance of the current policy πt.
• self update: Alter and update (πt, It) with an LLM and produce (πt+1, It+1).
• continue improve: If no other actions can be taken, recursively invoke the decision algorithm f to produce new actions.
The agent code is updated to (πt+1, It+1) after the current execution of (πt, It) is finished. Both the agent algorithm (π, I) and the action set A are not static and can be expanded and modified by the agent itself at the training time. Algorithm 1 illustrates the described algorithm for the G ̈odel Agent. Each recursive call enables the agent to refine its logic and become progressively more efficient.
There are various ways to initiate a G ̈odel Agent. Any specific agent instance during the recursive optimization process can be viewed as an instantiation of the G ̈odel Agent. Our implementation leverages runtime memory interaction techniques to enable self-awareness and self-modification, as illustrated in Figure 3. These techniques include dynamic memory reading and writing (monkey patching) to facilitate recursive self-improvement. Additionally, we have incorporated several auxiliary tools to accelerate the convergence of the G ̈odel Agent ’s optimization process.
The core functionalities of our G ̈odel Agent are outlined below: Self-Awareness via Runtime Memory Inspection G ̈odel Agent achieves self-awareness by inspecting runtime memory, particularly local and global variables in Python. This capability allows the agent to extract and interpret the variables, functions, and classes that constitute both the environment and the agent itself, according to the modular structure of the system.
Discussion. Table 3 draws an analogy between human selfreference and the potential for self-referential capabilities in artificial agents. Inspired by this analogy, we believe that self-reference constitutes a foundational and indispensable attribute for the development of AGI, and that future agents should inherently be self-referential. As foundation models grow in power, agents can more effectively enhance their own capabilities, ultimately evolving beyond the boundaries (or limitations) of human design. Furthermore, when an agent adjusts its own code based on feedback, this is akin to an executable version of test-time computing. In the context of LLMs, test-time computing typically involves generating additional tokens during inference, which then serve as a prefix to the final answer. This is because LLMs process information solely through text, making this their primary method for increasing computational effort at test time. For agents, however, their ability to call tools and execute code allows for far more diverse forms of test-time computing. G ̈odel Agent actualizes these more diverse forms of test-time computing precisely by modifying its own runtime code during test time.
There is significant room for improvement in the effectiveness, efficiency, and robustness of the G ̈odel Agent’s self-improvement capabilities, which requires better initial designs. The following are some promising directions for enhancement: 1) Enhanced Optimization Modules: Utilize human priors to design more effective optimization modules, such as genetic algorithms and reinforcement learning frameworks. 2) Expanded Modifiability: Broaden the scope of permissible modifications, allowing the agent to design and execute code that can fine-tune its own LLM modules. 3) Improved Environmental Feedback and Task Sequencing: Implement more sophisticated environmental feedback mechanisms and carefully curated task sequences during the initial optimization phase to prime the agent’s capabilities. Once the agent demonstrates sufficient competence, it can then be exposed to real-world environments. In addition, there are several other directions worth exploring and analyzing: Collective Intelligence Investigate the interactions among multiple G ̈odel Agents. Agents could consider other agents as part of their environment, modeling them using techniques such as game theory. This approach treats these agents as predictable components of the environment, enabling the study of properties related to this specific subset of the environment. Agent and LLM Characteristics Use the G ̈odel Agent’s self-improvement process as a means to study the characteristics of agents or LLMs. For example, can an agent genuinely become aware of its own existence, or does it merely analyze and improve its state as an external observer? This line of inquiry could yield insights into the nature of self-awareness in artificial systems. Theoretical Analysis Explore whether G ̈odel Agent can achieve theoretical optimality and what the upper bound of its optimization might be. Determine whether the optimization process could surpass the agent’s own understanding, and if so, at what point this might occur. Safety Considerations Although the current behavior of FMs remains controllable, as their capabilities grow, fully self-modifying agents will require human oversight and regulation. It may become necessary to limit the scope and extent of an agent’s self-modifications, ensuring that modifications occur only within a controlled environment.
Conclusion. We propose G ̈odel Agent, a self-referential framework that enables agents to recursively improve themselves, overcoming the limitations of handdesigned agents and meta-learning optimized agents. G ̈odel Agent can dynamically modify its logic based on high-level objectives. Experimental results demonstrate its superior performance, efficiency, and adaptability compared to traditional agents. This research lays the groundwork for a new paradigm in autonomous agent development, where LLMs, rather than human-designed constraints, define the capabilities of AI systems.
Limitations. As the first self-referential agent, G ̈odel Agent has to construct all task-related code autonomously, which poses significant challenges. Consequently, this work does not compare directly with the most complex existing agent systems, such as Open- Devin (Wang et al., 2024b), which have benefited from extensive manual engineering efforts. This makes it unrealistic to expect it to outperform systems that have taken researchers several months or even years to develop. The experiments presented in this paper are intended to demonstrate the feasibility of recursive self-improvement. Additionally, as the agent system becomes increasingly complex through self-optimization, it may require exponentially more intelligence to understand itself. Consequently, a system capable of complete self-referential at the outset may lose this capability as it evolves (Yampolskiy, 2015). The exact point at which the agent can no longer comprehend and improve itself has not been thoroughly explored. Investigating this phenomenon, both experimentally and theoretically, could provide valuable insights into the limitations of recursive selfimprovement. A more robust and advanced implementation of the G ̈odel Agent is anticipated, with numerous potential improvements outlined in Section 7.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What limits recursive self-improvement in autonomous AI systems?- Does keeping the utility function external limit true self-reference?
- What makes evolving the benchmark different from evolving the optimizer itself?
- Do evolutionary discovery systems like FunSearch count as bounded or open-ended improvement?
- Do evolutionary archives let agents improve themselves without formal proof?
- Why is self-amplification a property of AI-R&D systems rather than isolated agents?
- How do evolutionary archives enable open-ended self-improvement without formal proofs?
- How does compiling natural language goals into executable code enable objective evolution?
- Can objective search escape the limitations of fixed-objective central planning?
- Can AI systems generate and refine their own objective functions?
- How do current AI models perform when asked to specify their own goals?
- How would a bi-level agent restructure objective functions during discovery?
- Can humans fully understand why the AI search strategy succeeded here?