# Project export: CodeRepair

This document was generated by HackStack to give an AI agent context about a hackathon project. Sections are labeled with their provenance; content marked as truncated was cut to keep this document small.

## Project metadata

- Hackathon: TreeHacks 2024
- Tagline: Tool to automatically identify and fix bugs in codebases using GPT-4 to use debugger tools and automatically erase bugs. Scales to codebase of any size.
- Devpost: https://devpost.com/software/coderepair
- GitHub: https://github.com/jasondu7297/coderepair.ai
- Result: winner (YCombinator: YC / Yoneda Labs Prize (Office Hours with YC Group Partner))
- Team: 3 GitHub contributor(s) — Jason Du (25 commits), Anthony Wang (5 commits), AaranArulraj (1 commits)

## Devpost submission (written by the team)

### Inspiration

Many aspects of the programming workflow have been automated by the introduction of LLMs—especially with regard to code synthesis with products such as Github Copilot, AlphaCode and more. Other companies such as CodeGen and Sweep AI have begun tackling more niche subtasks, such as automatically solving tickets, conducting unit testing, catching security vulnerabilities, and more. However the complexity of coding tasks that can be solved with language models such as GPT4 is still very limited. Specifically, LLMs are used to scan and analyze codebases at a surface level to suggest code changes, identify errors, and fill code gaps. On the other hand, when humans debug code, they leverage tools such as GDB to inspect the state of a program’s internal variables and logic to garner a more complete picture of the root cause of the bug. We put these same developer tools in the hands of GPT4 to augment its current capacity for debugging code in hopes that it will be more capable in addressing runtime errors, algorithmic errors, and more.

### What it does

CodeRepair identifies and fixes errors that cause runtime failures. Anything from segmentation faults and logic errors, to unit test failures and integration test failures… We allow GPT4 to conduct a far deeper analysis of the code’s behavior by giving it control over the GDB debugging tool. You think that’s it? We were also able to provide monitored access to shell commands so GPT can dynamically explore the codebase, circumventing its normal context length limitations. As far as we know, this is the FIRST application of LLMs where developer tools like GDB were provided to further beef up the robust model.

### How we built it

We envisioned an agentic workflow where GPT4 can determine its own decisions while requesting for additional knowledge in the form of shell command and GDB command outputs from our program. More specifically, CodeRepair is based on Python and the OpenAI API, using the GPT4 Turbo model.

### Challenges we ran into

Designing the backend architecture of the GPT agent proved to be a much larger challenge than any of us expected. We spent a significant portion of our time testing a solution that eventually appeared unnecessarily complicated and highly undoable in the time constraints. One of the largest challenges we ran into was scalability. This problem was twofold: How do we deal with really large codebases that would exceed GPT context length? Come watch our live demo to see this unfold! How do we deal with really large agent traces? We didn’t want the length of the [(action, output), (action, output), …] chain produced by a single agent to grow too large due to performs concerns. This was a large consideration for us while developing and motivated our original approach which was to create multiple subagents, all of which were instantiated and orchestrated by a single main agent. This approach proved to overcomplicate the task, as getting the orchestrating agent to perform well was very difficult.

### Accomplishments we're proud of

Watching GPT effectively use gdb and intelligently step through the code base was super cool and something that we are proud of. We are not aware of any other software that gives a language model that capability so it was definitely great to watch our program doing this.

### What we learned

We learned the importance of providing GPT with deliberate instructions. One of the most difficult things was getting GPT to produce the output in a correct structure so that it could be reliably used by later parts of our program. One of the things we found is that in order to do this we really had to spoon feed to GPT exactly what it was and wasn’t allowed do.

### What's next

There is a lot in store for CodeRepair that we are super excited about. One particular thing we are really eager to work on next is to generalize our program to work for all languages rather than being bottlenecked by the debuggers available to us. https://github.com/jasondu7297/coderepair.ai

## README (from the GitHub repository)

No README available.

## Detected evidence (automated analysis)

Indexed codebase: 1167 recognized source files, 1870 KB.
- C (language) — detected in the code
- C++ (language) — detected in the code
- OpenAI (technology) — detected in the code
- Python (language) — detected in the code

## Codebase structure (from repository index)

### Files (120 of 1188)

```
.gitignore
data/test_reduced.json
pyproject.toml
scripts/coderepair-lib.sh
setup.py
src/agent.py
src/command.py
src/jobs.py
src/main.py
tests/code_force_tests/code_force_test_prompts/description0.txt
tests/code_force_tests/code_force_test_prompts/description1.txt
tests/code_force_tests/code_force_test_prompts/description2.txt
tests/code_force_tests/code_force_test_prompts/description3.txt
tests/code_force_tests/code_force_test_prompts/description4.txt
tests/code_force_tests/code_force_test_prompts/description5.txt
tests/code_force_tests/code0_0.cc
tests/code_force_tests/code0_106.cc
tests/code_force_tests/code0_107.cc
tests/code_force_tests/code0_108.cc
tests/code_force_tests/code0_110.cc
tests/code_force_tests/code0_115.cc
tests/code_force_tests/code0_12.cc
tests/code_force_tests/code0_120.cc
tests/code_force_tests/code0_125.cc
tests/code_force_tests/code0_126.cc
tests/code_force_tests/code0_128.cc
tests/code_force_tests/code0_131.cc
tests/code_force_tests/code0_133.cc
tests/code_force_tests/code0_137.cc
tests/code_force_tests/code0_140.cc
tests/code_force_tests/code0_143.cc
tests/code_force_tests/code0_144.cc
tests/code_force_tests/code0_146.cc
tests/code_force_tests/code0_148.cc
tests/code_force_tests/code0_150.cc
tests/code_force_tests/code0_153.cc
tests/code_force_tests/code0_16.cc
tests/code_force_tests/code0_161.cc
tests/code_force_tests/code0_163.cc
tests/code_force_tests/code0_164.cc
tests/code_force_tests/code0_166.cc
tests/code_force_tests/code0_168.cc
tests/code_force_tests/code0_17.cc
tests/code_force_tests/code0_171.cc
tests/code_force_tests/code0_175.cc
tests/code_force_tests/code0_176.cc
tests/code_force_tests/code0_178.cc
tests/code_force_tests/code0_181.cc
tests/code_force_tests/code0_185.cc
tests/code_force_tests/code0_193.cc
tests/code_force_tests/code0_196.cc
tests/code_force_tests/code0_197.cc
tests/code_force_tests/code0_2.cc
tests/code_force_tests/code0_20.cc
tests/code_force_tests/code0_200.cc
tests/code_force_tests/code0_203.cc
tests/code_force_tests/code0_204.cc
tests/code_force_tests/code0_205.cc
tests/code_force_tests/code0_208.cc
tests/code_force_tests/code0_209.cc
tests/code_force_tests/code0_211.cc
tests/code_force_tests/code0_212.cc
tests/code_force_tests/code0_217.cc
tests/code_force_tests/code0_219.cc
tests/code_force_tests/code0_223.cc
tests/code_force_tests/code0_226.cc
tests/code_force_tests/code0_23.cc
tests/code_force_tests/code0_232.cc
tests/code_force_tests/code0_233.cc
tests/code_force_tests/code0_238.cc
tests/code_force_tests/code0_24.cc
tests/code_force_tests/code0_241.cc
tests/code_force_tests/code0_244.cc
tests/code_force_tests/code0_252.cc
tests/code_force_tests/code0_26.cc
tests/code_force_tests/code0_263.cc
tests/code_force_tests/code0_265.cc
tests/code_force_tests/code0_266.cc
tests/code_force_tests/code0_269.cc
tests/code_force_tests/code0_271.cc
tests/code_force_tests/code0_272.cc
tests/code_force_tests/code0_275.cc
tests/code_force_tests/code0_277.cc
tests/code_force_tests/code0_278.cc
tests/code_force_tests/code0_280.cc
tests/code_force_tests/code0_286.cc
tests/code_force_tests/code0_288.cc
tests/code_force_tests/code0_294.cc
tests/code_force_tests/code0_296.cc
tests/code_force_tests/code0_301.cc
tests/code_force_tests/code0_303.cc
tests/code_force_tests/code0_310.cc
tests/code_force_tests/code0_311.cc
tests/code_force_tests/code0_313.cc
tests/code_force_tests/code0_314.cc
tests/code_force_tests/code0_316.cc
tests/code_force_tests/code0_318.cc
tests/code_force_tests/code0_321.cc
tests/code_force_tests/code0_322.cc
tests/code_force_tests/code0_326.cc
tests/code_force_tests/code0_327.cc
tests/code_force_tests/code0_33.cc
tests/code_force_tests/code0_330.cc
tests/code_force_tests/code0_334.cc
tests/code_force_tests/code0_340.cc
tests/code_force_tests/code0_348.cc
tests/code_force_tests/code0_349.cc
tests/code_force_tests/code0_352.cc
tests/code_force_tests/code0_357.cc
tests/code_force_tests/code0_360.cc
tests/code_force_tests/code0_362.cc
tests/code_force_tests/code0_364.cc
tests/code_force_tests/code0_370.cc
tests/code_force_tests/code0_374.cc
tests/code_force_tests/code0_378.cc
tests/code_force_tests/code0_38.cc
tests/code_force_tests/code0_380.cc
tests/code_force_tests/code0_382.cc
tests/code_force_tests/code0_386.cc
tests/code_force_tests/code0_39.cc
[1068 more files omitted for size]
```

### Dependencies

- pyproject.toml: annotated-types@>=0.6.0, argparse@>=1.4.0, openai@>=1.12.0, pydantic@>=2.6.1, pydantic_core@>=2.16.2, typing_extensions@>=4.9.0

### Recent commits (newest first)

- FInal changes to workflow and add stdin
- Debugging and fixing gdb + bash interactions
- migrated from using message history to appending cmd_history to prompt manually
- update program name for argparsing
- Move tests to correct directory
- Low-hanging fruit: Script for program, update gitignore
- added testcase prompts
- Restructure and cleanup tests
- Add python argparsing
- fix testcasesVF with .cc input output
- Update command logic for BASH
- Restructure directories
- Bug fixes
- Merge pull request #3 from jasondu7297/jasondu/major-refactor
- Merge pull request #2 from jasondu7297/jasondu/debugging
- Finish debugging!
- Remove pydantic due to difficulty in handling with subprocess and openai
- debugging and removing classmethods
- Update infra
- codeforces bugged solutions

## Key source files (fetched from GitHub, selected and truncated for size)

### pyproject.toml

```
[project]
name = "code-repair"
version = "0.0.1"
readme = "README.md"

dependencies = [
  "annotated-types>=0.6.0",
  "argparse>=1.4.0",
  "openai>=1.12.0",
  "pydantic>=2.6.1",
  "pydantic_core>=2.16.2",
  "typing_extensions>=4.9.0",

]

[tool.pyright]
pythonVersion = "3.10"
include = [
  "src/"
]

```

### src/main.py

```python
import sys
from argparse import ArgumentParser, Namespace
from jobs import Job, ATTEMPTS_LIMIT_EXCEEDED

def parse_args() -> Namespace:
    parser = ArgumentParser(prog='coderepair')

    parser.add_argument(
        "-e",
        "--executable",
        help="Path to the executable to run your test suite",
    )
    parser.add_argument(
        "-c",
        "--compile-cmd",
        help="Command to recompile your executable as inputed with -e/--executable flag",
    )

    parser.add_argument(
        "-i",
        "--ipt",
        help="std input to pass to tested program",
    )

    args = parser.parse_args()

    if len(vars(args)) < 2:
        parser.error("both the arguments -e/--executable and -c/--compile-cmd are required")

    return args

def main():
    args = parse_args()
    job = Job(
        exec_path=args.executable,
        compile_cmd=args.compile_cmd,
        ipt = args.ipt)

    try:
        job.execute()
    except ATTEMPTS_LIMIT_EXCEEDED as e:
        print(f"Program debugging failed with error message: {e.message}")
    

if __name__ == "__main__":
    main()

```

### setup.py

```python
from setuptools import find_namespace_packages, setup

setup(
    name="code-repair",
    packages=find_namespace_packages(include=["src.*"]),
)

```

### src/command.py

```python
import json
import subprocess
from enum import Enum
from typing import Any

class CommandType(Enum):
    BASH = 0
    GDB = 1
    WRITE = 2
    END = 3

class Command():
    cmd_str: str
    cmd_type: CommandType
    params: list[str] = []

    def __init__(self, cmd_raw: str, use_gdb: bool):
        if cmd_raw.upper().startswith("BASH"):
            self.cmd_str = cmd_raw[len("BASH COMMAND:"):].strip()
            self.cmd_type = CommandType.BASH
        elif cmd_raw.upper().startswith("GDB"):
            self.cmd_str = cmd_raw[len("GDB COMMAND:"):].strip()
            self.cmd_type = CommandType.GDB
        elif cmd_raw.upper().startswith("WRITE"):
            self.cmd_str = cmd_raw[len("WRITE"):].strip()
            self.cmd_type = CommandType.WRITE
        elif cmd_raw.upper().startswith("FINISH!"):
            self.cmd_type = CommandType.END
        else:
            raise ValueError("GPT response not of valid type")

    def execute(self) -> Any:
        kwargs = {
            "capture_output": True,
            "text": True,
        }

        if self.cmd_type == CommandType.BASH:
            kwargs["shell"] = True
        
        response = subprocess.run(self.cmd_str, **kwargs)
        return response.stdout
```

### scripts/coderepair-lib.sh

```shell
#!/bin/bash

# Bash library for controlling the CodeRepair build and environment

# ASSUMPTIONS:
#   - repo root does not move

############################################
# Variables
############################################

export SCRIPT_DIR="$( cd -- "$( dirname -- "${BASH_SOURCE[0]}" )" &> /dev/null && pwd )"
export REPO_ROOT="$(cd ${SCRIPT_DIR} && git rev-parse --show-superproject-working-tree --show-toplevel | head -1)"/src

############################################
# Help
############################################

coderepair-help() {
    cat <<'_EOF_'
-------------------------------------------------------------
coderepair-lib: a collection of bash functions to ease development
-------------------------------------------------------------

    Running CodeRepair:
        coderepair run: ------------- Run the CodeRepair debugging program
    
    Building and Installing:
        coderepair install: --------- Install Python dependencies and setup project

    Other:
        coderepair -h|--help: ------------ Prints this help message

Happy developing!

_EOF_

    [[ ${#} -eq 0 ]]
}

############################################
# Running CodeRepair
############################################

coderepair-run() {
    local HELP="""\
Run the CodeRepair debugging program

Syntax: coderepair-run [-h]
----------------------------------------------
    -h                  Print this help message
"""

    python3 "${REPO_ROOT}"/main.py "${@}"
}

############################################
# Aliasing CodeRepair
############################################

coderepair() {
    case "$1" in
        "-h"|"--help") coderepair-help ;;
        "run") coderepair-run "${@:2}" ;;
        *) echo "Unknown command. Run 'coderepair -h' for usage." ;;
    esac
}
```

### src/jobs.py

```python
import logging
import subprocess
from typing import List
from agent import Agent, AgentConfig

class Job():
    exec_path: str
    compile_cmd: str

    # For any job, there will be a minimum of 1 attempt
    _remaining_attempts: int = 9

    def __init__(self, exec_path: str, compile_cmd: str, ipt: str):
        self.exec_path = exec_path
        self.compile_cmd = compile_cmd
        self.input = ipt

    def execute(self) -> int:
        try:
            result = subprocess.run(self.exec_path, input=self.input, capture_output=True, text=True, check=True, shell=True)
        except subprocess.CalledProcessError as e:
            '''
            When check=True, a CalledProcessError exception is thrown if subprocess.run results in non-zero exit code
            If CalledProcessError exception thrown, init GPTJob to execute with GPT assistance
            '''
            print("Program failed with output: ", e.stdout)
            print("Program failed with error: ", e.stderr)

            gptJob = GPTJob(
                exec_path=self.exec_path,
                compile_cmd=self.compile_cmd,
                error_trace=e.stderr,
                ipt=self.input
            )
            # Initiate early stoppage to prevent indefinite iterations
            # Throw ATTEMPTS_LIMIT_EXCEEDED
            if self._remaining_attempts == 0:
                raise ATTEMPTS_LIMIT_EXCEEDED

            response_msg = gptJob.execute()
            recompile = gptJob.recompile()
            run_tests = gptJob.run_tests()

            if recompile == 0 and run_tests == 0:
                print("YAYYYYY! " + response_msg)
                return 0

            self._remaining_attempts -= 1

    def terminate(self):
        pass

class GPTJobMetadata():
    changes: dict
    failed: bool

class GPTJob():
    '''
    For each execution of GPTJob, we should retain the suggested change, the actual output, and expected output
    '''
    history: list[GPTJobMetadata] = []
    
    exec_path: str
    compile_cmd: str
    error_trace: str

    def __init__(self, exec_path: str, compile_cmd: str, error_trace: str, ipt: str):
        self.exec_path = exec_path
        self.compile_cmd = compile_cmd
        self.error_trace = error_trace
        self.input = ipt

    def execute(self) -> bool:
        agentConfig = AgentConfig(openai_key='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', executable_path=self.exec_path)
        gptAgent = Agent(agentConfig, self.error_trace)

        return gptAgent.spawn_gpt()            

    def recompile(self) -> int:
        response = subprocess.run(self.compile_cmd, shell=True, capture_output=True, text=True)
        return response.returncode
    
    def run_tests(self) -> int:
        response = subprocess.run(self.exec_path, input=self.input, shell=True, capture_output=True, text=True)
        return response.returncode

class ATTEMPTS_LIMIT_EXCEEDED(Exception):
    def __init__(self):
        self.message = f"Exceeded limit of maximum attempts"
        super().__init__(self.message)
```

### src/agent.py

```python
from typing import Any, Dict, List
from command import Command, CommandType
import subprocess
import openai
import logging

class AgentConfig():
    openai_key: str
    executable_path: str
    use_gdb: bool

    def __init__(self, openai_key: str, executable_path: str, use_gdb = True):
        self.openai_key = openai_key
        self.executable_path = executable_path
        self.use_gdb = True

class Agent():
    cmd_history: List[str] = []
    error_trace: str
    executable: str

    use_gdb: bool
    gdb_instance: object

    prompt: str

    def __init__(self, config: AgentConfig, error_trace: str):
        self.executable = config.executable_path
        self.error_trace = error_trace
        self.use_gdb = config.use_gdb

        self.update_prompt()
        openai.api_key = config.openai_key
    
    def spawn_gpt(self) -> Any:
        if self.use_gdb:
            self.gdb_instance = spawn_persisting_gdb(self.executable)

        '''
        Define the behaviour of GPT4 model
        '''

        should_continue = True
        while should_continue:
            response = openai.chat.completions.create(
                model="gpt-4-turbo-preview",
                messages=[
                    {
                        "role": "user",
                        "content": self.prompt,
                    },
                ],
            )
            response_msg = response.choices[0].message

            if response_msg.content:
                cmd_str = response_msg.content.strip()                
                serialized_cmd = Command(cmd_str, self.use_gdb)

                print(cmd_str)

                if serialized_cmd.cmd_type == CommandType.END:
                    if self.use_gdb:
                        terminate_persisting_gdb(self.gdb_instance)

                    print("Finishing code repair session based on OpenAI suggestion:", response_msg.content)
                    return response_msg

                elif serialized_cmd.cmd_type == CommandType.GDB:
                    self.gdb_instance.stdin.write(serialized_cmd.cmd_str + "\n")
                    self.gdb_instance.stdin.flush()

                    execution_result = self.gdb_instance.stdout.readline().strip()
                    self.cmd_history.append("Command: " + serialized_cmd.cmd_str + "\nOutput: " + execution_result)
                    self.update_prompt()
                    continue

                execution_result = serialized_cmd.execute()          
                self.cmd_history.append("Command: " + serialized_cmd.cmd_str + "\nOutput: " + execution_result)
                self.update_prompt()

    def update_prompt(self) -> None:
        self.prompt = f"""        
Traceback Error: {self.error_trace}

Previous commands that have been run and their respective outputs: 
""" + ('\n'.join(self.cmd_history) if self.cmd_history else "None") + """
You are to assume that we are already in a GDB debugging session. You are to analyze and change source code only. You are not to change any other code such as ASM, machine code, and file metadata. Any bash commands used should be to find and replace error-inducing lines of code.

REMEMBER: You are at the root-level of the codebase and all code related to the errors are within that scope. All commands will be run from your location, which is root-level.

You are to respond with exactly one GDB command OR one bash command—nothing more, nothing less. Under no circumstances should you deviate from this instruction, unless you are absolutely certain that the debugging process has concluded AND THAT WE HAVE used bash to fix the bugs in the code. In that case, and that case only, you may indicate completion by stating “FINISH!”, followed by a concise explaination of the bug and where it is in the same line, and info on how you fixed it. Do not attempt to provide explanations, elaborations, or multiple commands. Your sole output must be a single, precise GDB command or bash command relevant to the given context. Failure to adhere to these guidelines is not an option.

The bash command is to be executed using 

BASH COMMAND: <command>

The GDB command is to be executed using 

GDB COMMAND: <gdb command>
"""

def spawn_persisting_gdb(executable: str) -> object:
    return subprocess.Popen(['gdb', executable], stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True)

def terminate_persisting_gdb(gdb_instance: object) -> None:
    gdb_instance.terminate()

```

### tests/example_tests/segfault.c

```c
#include <stdio.h>

int main() {
    int a = 10;
    int b = 20;
    int segmentation = 20;
    int d = 20;
    int *ptr = NULL;
    *ptr = 10;  // Dereferencing a NULL pointer will cause a segmentation fault
    return 0;
}
```

### tests/code_force_tests/code0_26.cc

```c++
#include <bits/stdc++.h>
int t;
long long n, m, k;
int main() {
  scanf("%d", &t);
  while (t--) {
    scanf("%lld%lld%lld", &n, &m, &k);
    if (k - 2 <= 0 || n - 1 > m || m > (n * (n - 1) >> 1ll))
      puts("NO");
    else if (m != (n * (n - 1ll) >> 1ll) && k == 3)
      puts("NO");
    else
      puts("YES");
  }
  return 0;
}


```

### tests/code_force_tests/code0_554.cc

```c++
#include <bits/stdc++.h>
using namespace std;
int main() {
  int t;
  cin >> t;
  while (t--) {
    int n, m, k;
    cin >> n >> m >> k;
    long long int upper = (n * (n - 1)) / 2;
    if (m > upper || (k == 1)) {
      cout << "NO\n";
    } else if (m < upper && k <= 3) {
      cout << "NO\n";
    } else {
      cout << "YES\n";
    }
  }
}


```

[1158 more indexed source files omitted to keep this export small. The full file list is in the Codebase structure section above.]