Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation

Paper · arXiv 2310.02304 · Published October 3, 2023
Frontier AI Risk & RSI

Several recent advances in AI systems solve problems by providing a “scaffolding” program that structures multiple calls to language models (LMs) to generate better outputs. A scaffolding program is written in a programming language such as Python. In this work, we use a language-model-infused scaffolding program to improve itself. We start with a seed “improver” that improves an input program according to a given utility function by querying an LM several times and returning the best solution. We then run this seed improver to improve itself. Across a small set of downstream tasks, the resulting improved improver generates programs with significantly better performance than its seed improver. A variety of self-improvement strategies are proposed by the language model, including beam search, genetic algorithms, and simulated annealing. Since the language models themselves are not altered, this is not full recursive self-improvement. Nonetheless, it demonstrates that a modern language model, GPT-4 in our experiments, is capable of writing code that can call itself to improve itself. We consider concerns around the development of self-improving technologies and evaluate the frequency with which the generated code bypasses a sandbox.

Introduction. A language model (LM) can be queried to optimize virtually any objective describable in natural language. However, a program that makes multiple, structured calls to an LM can often produce outputs with higher objective values (Yao et al., 2022; 2023; Zelikman et al., 2023; Chen et al., 2022b). We refer to these as “scaffolding” programs, typically written (by humans) in a programming language such as Python. Our key observation is that, for any distribution over optimization problems and any fixed LM, designing a scaffolding program is itself an optimization problem.

In this work, we introduce the Self-Taught Optimizer (STOP), a method in which code that applies an LM to improve arbitrary solutions is applied recursively to improve itself within a defined scope. Our approach begins with a seed ‘improver’ scaffolding program that uses the LM to improve a solution to some downstream task. As the system iterates, the LM refines this improver. We quantify the performance of our self-optimizing framework with downstream algorithmic tasks, observing improvements when the LM applies its self-improvement strategies over increasing iterations. Thus, STOP shows how LMs can act as their own meta-optimizers. We also investigate the kinds of self-improvement strategies

Related work. Language Model Scaffolding. Many prompting strategies and scaffolds have been developed to enable more systematic reasoning in LMs (Wei et al., 2022b; Yao et al., 2022; 2023; Zelikman et al., 2023; Chen et al., 2022b; Zhou et al., 2022a; Khattab et al., 2022; Jiang et al., 2022; Sel et al., 2023; Besta et al., 2023; Poesia et al., 2023). For example, scratchpads and chain-of-thought rely on communicating to the model that it should work through a problem step-by-step (Nye et al., 2021; Wei et al., 2022b). Tree-of-Thoughts algorithmically scaffolds the model to consider branching paths of reasoning steps (Yao et al., 2023). Graph of thoughts extends this, allowing other graph operations (where nodes are reasoning steps), such as aggregation (Besta et al., 2023). Other work has focused on letting models reason with access to an interpreter such as Program of Thoughts prompting (Chen et al., 2022b), Program-aided Language Models (Gao et al., 2023), Reflexion (Shinn et al., 2023), or ReAct (Yao et al., 2022), while yet others abstracted this scaffolding structure such as Demonstrate- Search-Predict (DSP) (Khattab et al., 2022), Language Model Cascades (Dohan et al., 2022), or Cognitive Architectures (Sumers et al., 2023). Each work can be viewed as the result of researchers asking, “Given an imperfect LM, how can we provide structure to help it solve problems?” We instead ask if LMs can design that structure and improve it using itself. Surprisingly, GPT-4 proposes scaffolding techniques introduced after its training cutoff.

Language Models as Prompt Engineers. Work has also explored LMs’ ability to optimize prompts, such as the Automatic Prompt Engineer (APE) (Zhou et al., 2022b) or, recently, OPRO (Yang et al., 2023) and Promptbreeder (Fernando et al., 2023). Note that, for these, the goal has consistently been to scaffold the LM to produce a prompt but not to scaffold it to produce a better scaffolding (beyond prompting-only scaffolds like zero-shot chain-of-thought), nor to produce a recursively applicable scaffolding. In other words, these works can be understood as proposing particular scaffolds for prompt engineering but not for scaffold proposal. But, we share the inspiration of LMs improving their reasoning without fine-tuning.

Language Model Self-Improvement. Prior work, such as STaR (Zelikman et al., 2022), demonstrated that LMs can learn to solve harder problems by learning from their reasoning chains by filtering based on incorrect answers (as well as Huang et al. 2022, which explored the specific case where a majority vote is used as the filter and Uesato et al. 2022, which emphasized the value of checking the accuracy of the reasoning itself). Inspired by self-play in games, Haluptzok et al.

Method. Figure 3 provides a visual schematic of the self-improvement pipeline envisaged in Section 3, while Algorithm 1 provides Self-Taught Optimizer (STOP) pseudocode. The key observation is that the selection of I is an optimization problem itself, to which we can recursively apply improvement. STOP begins with an initial seed improver I0. We define the t-th improver as the output of t self-improvement rounds with meta-utility ˆu: It ≜It−1( ˆu, It−1, L). This is iterated for a prespecified number of iterations T, per available resources.

Intuition. By using ˆu, STOP selects improver based on a downstream utility improvement. This approach is motivated by the intuitions that 1) improvers that are good at improving downstream solutions may be more likely to be good scaffolding programs and thus to be good at self-improvement, and 2) selecting for single-round improvements may lead to better multi-round improvements. In practice, we allow the utilities and LM to impose budget constraints and initial solutions to be generated by humans or a model. Moreover, the cost is essentially O((budgetu + budgetL) ∗budget ˆu), where budget specifies the number of times an improver can use a function, with these asymptotics defined with respect to the budget parameters.

Designing the seed improver. Our chosen seed improver (Figure 2) simply prompts the LM to generate candidate improvements of an initial solution and then returns the best solution according to the utility function. We chose this simple form to provide nontrivial improvement for a generic downstream task while 1) encouraging the LM to be as “creative” as possible, 2) minimizing initial prompt complexity, since self-improvement introduces additional complexity due to nested references to code strings inside of prompts, and 3) minimizing the prompt token count and therefore the costs of LM queries. We considered other seed prompt variants but heuristically found that this version maximized the novelty of GPT-4-proposed improver improvements.

Describing the utility. To effectively convey the details of the utility function to the LM, we provide the utility to the improver in two forms, as a callable function and as a utility description string containing the essential elements of the utility source code (see Appendices E and F for examples). This choice was made for the following reasons. The description allows us to clearly convey budgetary constraints (e.g., on runtime or function calls) imposed by the utility to the LM. We first attempted to describe budgetary instructions in the seed improver prompt, but, as we discuss in Section 6.2, this led to the removal of such instructions and attempts at reward-hacking in later iterations. The downside of our approach is that it separates the constraints from the code to be optimized by the LM, which may decrease the likelihood that it will be used by the LM (Liu et al., 2023b). Finally, we observe empirically that replacing the source code with a purely English description of the utility leads to a reduced frequency of non-trivial improvement.

Discussion. We first describe STOP-proposed self-improvement strategies, with examples detailed in Appendix B and visualized in Figure 1. While each strategy was implemented by STOP, not all were ultimately selected as improvements, and some used an earlier iteration of the seed improver than in Figure 2 (see Figure A.19). Nonetheless, a variety of self-improvement strategies were selected as improved improvers, including the example given in Figure 5.

Ethics Statement: Concerns about Developing STOP Concerns about the consequences of RSI have been raised since its first mention. Minsky (1966) wrote, “Once we have devised programs with a genuine capacity for self-improvement, a rapid evolutionary process will begin... It is hard to say how close we are to this threshold, but once it is crossed, the world will not be the same.” This is a particularly contentious topic recently, with intensified concern over negative consequences (Ambartsoumean & Yampolskiy, 2023; Gabriel & Ghazavi, 2021).

It is thus important to carefully weigh the risks and benefits of studying RSI and specifically the small advance we introduce. First, STOP does not alter the black-box LM and hence is not full RSI. Moreover, at this point, we do not believe the scaffolding systems STOP creates are superior to those hand-engineered by experts. If this is the case, then STOP is not (currently) enabling additional AI misuse. At the same time, it facilitates the study of aspects of RSI code generation such as sandbox avoidance and reward hacking. As Christiano (2023) argues, advances in scaffolding and agent models have the advantage of interpretability compared to advances in LMs.

Conclusion. In this work, we introduced STOP, a framework for recursively optimizing code generation using LMs as meta-optimizers. We demonstrated that LMs like GPT-4 are capable of improving code that leverages the LM itself. We found that, across a variety of algorithmic tasks, STOP generates improvers that boost the performance of downstream code. While the model does not optimize its weights or underlying architecture, this work indicates that self-optimizing LMs do not require that. However, this is itself a motivation: the capabilities of future LMs may be misunderstood if strong scaffolding strategies are not tested. Understanding how LMs can improve their scaffoldings can help researchers understand and mitigate the potential for misuse of more powerful LMs. Lastly, STOP may allow researchers to investigate techniques for mitigating undesirable self-improvement strategies.

Limitations. A fundamental limitation of our approach is that the LM itself is not improved. Furthermore, our meta-utility measures improver quality only indirectly via improvements in down- Published as a conference paper at COLM 2024 stream task utility. Unlike in some prior work (e.g., Schmidhuber, 2003), any improvement attempt may result in worse performance, which can lead to further deterioration. Another limitation is that STOP requires an efficiently-evaluatable (and describable) utility function, which may not be available for every task. Correspondingly, as STOP’s cost grows substantially faster than the cost of the optimized improver, it may be expensive to run.

Our improvement framework also maintains a single improver at each step, while some approaches may benefit from maintaining a population. While this is not a strict limitation in that an improver could itself sample from a population of implementations, it likely imposes a bias. Moreover, a deeper analysis of alternative seed improvers and their tradeoffs would be valuable future work. Lastly, our experiments depend on a closed large LM that may be deprecated in the future, which harms interpretability and long-term reproducibility. Based on the GPT-3.5 results, it is unlikely that STOP would consistently work with any open-source LM at the time of writing (Touvron et al., 2023; Jiang et al., 2023).

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do models learn from self-generated outputs without cascading failures? Why does self-revision amplify confidence in wrong model answers? Can AI systems discover fundamental improvements to their own architectures? How does fine-tuning trade off accuracy against reasoning quality? What limits language model accuracy in evaluating ideas? What limits recursive self-improvement in autonomous AI systems? How do training data quality and composition affect downstream model performance? How do curriculum design and feedback approaches affect model learning?