# Project export: Aipeiron

This document was generated by HackStack to give an AI agent context about a hackathon project. Sections are labeled with their provenance; content marked as truncated was cut to keep this document small.

## Project metadata

- Hackathon: Cal Hacks 10.0
- Tagline: Automating accounting, financial modeling, and taxes with the power of AI. Connect to your bank, and accounting will be done behind the scenes. Financial models generated, and tax questions answered.
- Devpost: https://devpost.com/software/aipeiron
- GitHub: https://github.com/KesavaViswanadha/AipeironChatbot
- Demo: https://calhacks10-seven.vercel.app/api/auth/signin?callbackUrl=%2Fsetup
- Video: https://www.youtube.com/embed/lfqpJIAjOsk?enablejsapi=1&hl=en_US&rel=0&start=&version=3&wmode=transparent
- Team: 1 GitHub contributor(s) — KesavaViswanadha (2 commits)

## Devpost submission (written by the team)

### Inspiration

Back when I was a freshman, I started working at Rangeview, an aerospace manufacturing startup. There I handled operations, which included accounting. I had absolutely no clue how to do taxes or accounting so I decided to hire an accounting firm. But the process was all very confusing to me. They would ask me for financial documents, or ask to do historical bookkeeping for our LLC, but no matter how many people I asked, I never had a solid grasp of what accounting was. So I decided to drop all of my classes and sign up for some accounting ones. A month into those classes, I realized what the accounting firms were doing was easy, and that I shouldn't be paying them so much money to do a task I could easily do myself, so I fired them and took up the task of being rangeviews sole accountant. Accounting today is broken. Companies with less than 100k in their bank account are expected to pay thousands of dollars a month for an accounting firm, and that’s not to mention the thousands more that they’ll charge you once you register to do your taxes or other special tasks. This is because accounting is a foreign word to most founders. Many of them don’t even know if accounting is legally required, let alone the processes that go behind the scenes of the accounting firm they hired. And thus, accounting firms can get away with low effort and sub-par results. Something has to change.

### What it does

Aipeiron to automates accounting, financial modeling, and compliance with the power of AI. With Aipeiron, any company will be able to get the insights that an accounting firm could provide 100x faster and 50x cheaper just as effectively. And the founder has full control and ownership every step of the way. Doing so is simple: by connecting your bank account, transactions flow automatically from your bank into our finetuned GPT-4 model, which then outputs categorized transactions, exactly what an accountant would do. From that data, you can generate financial models and budgets, or chat with our chatbot trained on the IRS’s instructions to talk about how you would file taxes, because taxes are confusing.

### How we built it

We used NextJS14. The account creation step was done with NextAuth and Google OAuth. The Bank connection and free flow of transactions was done through Plaid. The plaid transactions then flow into our GPT-4 model, and out comes a categorized transaction. We did this through MindsDB. Our database is Postgres, and we used Prisma as an ORM and Supabase to host it on the cloud. Everything from company and user information to the Plaid data, to the classified transactions, have all been stored in Postgres. We wrote our chatbot in Python by generating text embeddings by writing a parser

### Challenges we ran into

One challenge we ran into was using MindsDB and getting it up and running. I had a 50 message long slack thread about this! However, the main challenge was figuring out how to make three people work together. If we all worked on the full stack, the git pushes and pulls would get confusing, and honestly we’d all slow each other down, so we decided to divide up the tasks in a way that everyone was working on something separate, such as the plaid integration for one person, the chatbot for another, and the classifying transactions and storing the data on the cloud and displaying the data to the user for another. Lastly, figuring out how to deploy to Vercel was also a little bit confusing.

### Accomplishments we're proud of

We built this project literally from scratch from the moment Calhacks started. We’re pretty proud of that, and even though our project is very messy and “hacked” together, we got all the things we wanted to up and running.

### What we learned

We learned about data pipelines such as mindsdb, we learned how to use plaid to pass financial transactions back and forth, we learned about vector embeddings, we learned about creating charts and graphs in React. It was very fun.

### What's next

We want to keep building this idea, and turn it into an actual company. The code right now works, but lacks refine. We will build the product fully, launch it, test it out with various companies, see if we can raise money, and try to turn it into an actual startup.

## README (from the GitHub repository)

No README available.

## Detected evidence (automated analysis)

Indexed codebase: 4 recognized source files, 16 KB.
- JavaScript (language) — detected in the code
- OpenAI (technology) — detected in the code
- Python (language) — detected in the code
- Supabase (technology) — detected in the code
- Next.js (technology) — claimed on Devpost, not found in the code
- PostgreSQL (technology) — claimed on Devpost, not found in the code
- React (technology) — claimed on Devpost, not found in the code
- TypeScript (language) — claimed on Devpost, not found in the code

## Codebase structure (from repository index)

### Files (9 of 9)

```
Python_chatbot/.DS_Store
Python_chatbot/document_chunker.py
Python_chatbot/output_query_vectara.json
Python_chatbot/output_query.json
Python_chatbot/output.json
Python_chatbot/package.json
Python_chatbot/query_bot.py
Python_chatbot/server.js
Python_chatbot/vectara_chatbot.py
```

### Dependencies

- Python_chatbot/package.json: @supabase/supabase-js@^2.38.4, openai@^4.14.1

### Recent commits (newest first)

- lksdjf
- Add files via upload

## Key source files (fetched from GitHub, selected and truncated for size)

### Python_chatbot/package.json

```
{
  "name": "python_chatbot",
  "version": "1.0.0",
  "description": "",
  "main": "index.js",
  "scripts": {
    "test": "echo \"Error: no test specified\" && exit 1"
  },
  "keywords": [],
  "author": "",
  "license": "ISC",
  "dependencies": {
    "@supabase/supabase-js": "^2.38.4",
    "openai": "^4.14.1"
  },
  "type": "module"
}

```

### Python_chatbot/server.js

```javascript
import { createClient } from "@supabase/supabase-js";
import OpenAIApi from "openai";
import { spawn } from 'child_process';
import { promises as fs } from 'fs';
import { createInterface } from 'readline';

/**const supabaseClient = createClient("https://vswuyfouifxhjyloqbdl.supabase.co", "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6InZzd3V5Zm91aWZ4aGp5bG9xYmRsIiwicm9sZSI6ImFub24iLCJpYXQiOjE2OTg1Mzk3NDQsImV4cCI6MjAxNDExNTc0NH0.LoSboreByG0r1ly6ePCk3Ve2Ewf-v_ko4iXRoxojiz8");
const openai = new OpenAIApi({
    apiKey: "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"
  });*/
//const { spawn } = require('child_process');
//const fs = require('fs').promises; // Use the promise-based version of the 'fs' module

async function callDocChunker(filePath) {
  return new Promise((resolve, reject) => {
    const pythonProcess = spawn('python3', ['document_chunker.py', filePath]);

    pythonProcess.on('close', async (code) => {
      if (code === 0) {
        try {
          // Read the results from the file written by the Python script
          const dataString = await fs.readFile('output.json', { encoding: 'utf8' });
          const result = JSON.parse(dataString);
          resolve(result);
        } catch (e) {
          reject(e);
        }
      } else {
        reject(`child process exited with code ${code}`);
      }
    });
  });
}

async function callQueryBot(this_question) {
  return new Promise((resolve, reject) => {
    const pythonProcess = spawn('python3', ['query_bot.py', this_question]);

    pythonProcess.on('close', async (code) => {
      if (code === 0) {
        try {
          // Read the results from the file written by the Python script
          const dataString = await fs.readFile('output_query.json', { encoding: 'utf8' });
          const result = JSON.parse(dataString);
          resolve(result);
        } catch (e) {
          reject(e);
        }
      } else {
        reject(`child process exited with code ${code}`);
      }
    });
  });
}

async function callQueryBotVectara(this_question) {
  return new Promise((resolve, reject) => {
    const pythonProcess = spawn('python3', ['vectara_chatbot.py', this_question]);

    pythonProcess.on('close', async (code) => {
      if (code === 0) {
        try {
          // Read the results from the file written by the Python script
          const dataString = await fs.readFile('output_query_vectara.json', { encoding: 'utf8' });
          const result = JSON.parse(dataString);
          resolve(result);
        } catch (e) {
          reject(e);
        }
      } else {
        reject(`child process exited with code ${code}`);
      }
    });
  });
}

// function embedText(inputText) {
//     try {
//         var result = ""
//         return new Promise ((resolve) => {
//             openai
//                 .createEmbedding ({
//                     model: "text-embedding-ada-002",
//                     input: inputText,
//                 })
//                 .then ((res) => {
//                     //console. log (res.data ['data'] [0] ["embedding"])
//                     result = res.data["data"][0]["embedding"];
//                 });
//             setTimeout (() => {
//                 resolve (result);
//             }, 2000);
//         });
//     } catch (error) {
//         console.error(err);
//     }
// }

const readline = createInterface({
  input: process.stdin,
  output: process.stdout
});

async function generateEmbeddings() {
    //const configuration = new Configuration({ apiKey: "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"});

    const documents = await callDocChunker("/Users/kesvis/skylerProject/calhacks10/Python_chatbot/i941taxdoc.pdf");

    try {
      const userResponse = new Promise((resolve) => {
          readline.question('Please enter a prompt? ', (answer) => {
              resolve(answer);
          });
      });

      // Wait for the user to respond
      const this_prompt = await userResponse;

      // Call the bot with the user's input and wait for the response
      const answer_to_prompt = await callQueryBot(this_prompt);
      const answer_to_prompt_vectara = await callQueryBotVectara(this_prompt);
      // Log the answer and close the readline interface
      console.log(answer_to_prompt_vectara);
      //console.log(answer_to_prompt);
      readline.close();

  } catch (error) {
      // Handle any errors that may occur during the document chunking or vectara call
      console.error('An error occurred in generateEmbeddings:', error);
      readline.close();
  }
 
    
    
    /**console.log(answer_to_prompt);
    
    const answer_to_prompt_vectara = await callQueryBotVectara("How do I fill out form 941 properly if I am a large business in a hurricane?");
    console.log(answer_to_prompt_vectara);*/
 
  }
  

generateEmbeddings();
   // for (const document of documents) {
    //     const input = document.replace(/\n/g, '');
    //     embedText(input).then(async (result) => {

    //         const { data, error } = await supabaseClient
    //             .from("documents")
    //             .insert({
    //                 content: document,
    //                 embedding: result
    //             });
    //         setTimeout(() => {}, 500);
    //     });
        /**const embeddingResponse = await openai.createEmbedding({
            model: "text-embedding-ada-002",
            input
        })

        const [{ embedding }] = embeddingResponse.data.data;

        await supabaseClient.from('document').insert({
            content: document,
            embedding
        })*/
```

### Python_chatbot/query_bot.py

```python
from document_chunker import DocChunker
import pinecone
import numpy as np
import spacy
from PyPDF2 import PdfReader
import nltk
import sys
import json
import os
import openai
from scipy.spatial import distance
import plotly.express as px
from sklearn.cluster import KMeans
from umap import UMAP

#GET THE VECTOR FROM JAVASCRIPT
pinecone.init(api_key="b0ec4895-ab56-43b3-baf0-a404a9e28e20", environment="gcp-starter")
openai.api_key = "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"
if not("data-embeddings" in pinecone.list_indexes()):
    pinecone.create_index("data-embeddings", dimension=1536, metric="euclidean")
this_table = pinecone.Index("data-embeddings")

def getQueryAns(this_question):
    response = openai.Embedding.create(model= "text-embedding-ada-002", input=[this_question])
    this_ans = this_table.query(vector=response["data"][0]["embedding"], top_k=2, include_values=True)
    prompt = "You are a Certified Public Accountant who was asked: " + this_question + ". Please answer using this as context:" + this_ans['matches'][0]['id'] + ", " + this_ans['matches'][1]['id']
    response = openai.Completion.create(
    engine="davinci",
    prompt=prompt,
    max_tokens=500
    )

    return response.choices[0].text.strip()

def main(this_question):
    with open('output_query.json', 'w') as f:  # Write the results to a file
        json.dump(getQueryAns(this_question), f)

if __name__ == "__main__":
    main(sys.argv[1])
```

### Python_chatbot/vectara_chatbot.py

```python
""" This is an example of calling Vectara API via python using http/rest as communication protocol.
"""

import argparse
import json
import logging
import requests
import numpy as np
import spacy
from PyPDF2 import PdfReader
import nltk
import sys
import json
import os
import openai
from scipy.spatial import distance
import plotly.express as px
from sklearn.cluster import KMeans
from umap import UMAP
import pinecone

def _get_query_json(customer_id: int, corpus_id: int, query_value: str):
    """ Returns a query json. """
    query = {}
    query_obj = {}

    query_obj["query"] = query_value
    query_obj["num_results"] = 10

    corpus_key = {}
    corpus_key["customer_id"] = customer_id
    corpus_key["corpus_id"] = corpus_id

    query_obj["corpus_key"] = [ corpus_key ]
    query["query"] = [ query_obj ]
    return json.dumps(query)


def query(customer_id: int, corpus_id: int, query_address: str, api_key: str, query: str):
    """This method queries the data.
    Args:
        customer_id: Unique customer ID in vectara platform.
        corpus_id: ID of the corpus to which data needs to be indexed.
        query_address: Address of the querying server. e.g., api.vectara.io
        api_key: A valid API key with query access on the corpus.

    Returns:
        (response, True) in case of success and returns (error, False) in case of failure.

    """
    post_headers = {
        "customer-id": f"{customer_id}",
        "x-api-key": api_key
    }

    response = requests.post(
        f"https://{query_address}/v1/query",
        data=_get_query_json(customer_id, corpus_id, query),
        verify=True,
        headers=post_headers)

    if response.status_code != 200:
        logging.error("Query failed with code %d, reason %s, text %s",
                       response.status_code,
                       response.reason,
                       response.text)
        return response, False
    return response, True


def main(this_question):
    full_ans = query(2120125989, 1, "api.vectara.io", "zqt_fl6OJXFubZFMvUHJWOo3DUoEPbfX82Ff1SXz7w", this_question)
    if not full_ans[1]:
        print("failed.")
    this_ans = full_ans[0].json()
   #print(this_ans)
    with open('output_query_vectara.json', 'w') as f:  # Write the results to a file
        print()
        json.dump(this_ans['responseSet'][0]['response'][0]['text'] + " " + this_ans['responseSet'][0]['response'][1]['text'] + " " + this_ans['responseSet'][0]['response'][2]['text'], f)

if __name__ == "__main__":
    main(sys.argv[1])
    # logging.basicConfig(
    #     format="%(asctime)s %(levelname)-8s %(message)s", level=logging.INFO)

    # parser = argparse.ArgumentParser(
    #             description="Vectara rest example (With API Key authentication.")

    # parser.add_argument("--customer-id", type=int, help="Unique customer ID in Vectara platform.")
    # parser.add_argument("--corpus-id",
    #                     type=int,
    #                     help="Corpus ID to which data will be indexed and queried from.")

    # parser.add_argument("--serving-endpoint", help="The endpoint of querying server.",
    #                     default="api.vectara.io")
    # parser.add_argument("--api-key", help="API key retrieved from Vectara console.")
    # parser.add_argument("--query", help="Query to run against the corpus.", default="Test query")

    # args = parser.parse_args()

    # if args:
    #     error, status = query(args.customer_id,
    #                           args.corpus_id,
    #                           args.serving_endpoint,
    #                           args.api_key,
    #                           args.query)

```

### Python_chatbot/document_chunker.py

```python
import numpy as np
import spacy
from PyPDF2 import PdfReader
import nltk
import sys
import json
import os
import openai
from scipy.spatial import distance
import plotly.express as px
from sklearn.cluster import KMeans
from umap import UMAP
import pinecone

class DocChunker:
    # Load the Spacy model
    
    def __init__(self):
        # BEWARE OF THE FACT THAT INACTIVE DATABASES ON PINECONE GET GARBAGE
        # COLLECTED AFTER A DAY SO WEIRD STUFF MIGHT HAPPEN!!!!!!!
        pinecone.init(api_key="b0ec4895-ab56-43b3-baf0-a404a9e28e20", environment="gcp-starter")
        openai.api_key = "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"
        if not ("data-embeddings" in pinecone.list_indexes()):
            pinecone.create_index("data-embeddings", dimension=1536, metric="euclidean")
        self.this_table = pinecone.Index("data-embeddings")

        self.nlp = spacy.load('en_core_web_sm')
        self.clusters_lens = []
        self.final_texts = []
        self.final_embeddings = []
        nltk.download('punkt')

    # Extracting Text from PDF
    def extract_text_from_pdf(self, file_path):
        with open(file_path, 'rb') as file:
            pdf = PdfReader(file)
            text = " ".join(page.extract_text() for page in pdf.pages)
        return text
    
    def process_pdf(self, file_path):
        return self.final_chunks(self.extract_text_from_pdf(file_path))

    def process(self, text):
        doc = self.nlp(text)
        sents = list(doc.sents)
        vecs = np.stack([sent.vector / sent.vector_norm for sent in sents])

        return sents, vecs

    def cluster_text(self, sents, vecs, threshold):
        clusters = [[0]]
        for i in range(1, len(sents)):
            if np.dot(vecs[i], vecs[i-1]) < threshold:
                clusters.append([])
            clusters[-1].append(i)
        
        return clusters

    def clean_text(self, text):
        # Add your text cleaning process here
        return text


    def get_embedding(self, text_to_embed):
        # Embed a line of text
        response = openai.Embedding.create(model= "text-embedding-ada-002", input=[text_to_embed])
        # Extract the AI output embedding as a list of floats
        embedding = response["data"][0]["embedding"]
        
        return embedding

    def final_chunks(self, text):
        # Process the chunk
        threshold = 0.2
        sents, vecs = self.process(text)

        # Cluster the sentences
        clusters = self.cluster_text(sents, vecs, threshold)

        for cluster in clusters:
            cluster_txt = self.clean_text(' '.join([sents[i].text for i in cluster]))
            cluster_len = len(cluster_txt)
            
            # Check if the cluster is too short
            if cluster_len < 60:
                continue
            
            # Check if the cluster is too long
            elif cluster_len > 512:
                threshold = 0.5
                sents_div, vecs_div = self.process(cluster_txt)
                reclusters = self.cluster_text(sents_div, vecs_div, threshold)
                
                for subcluster in reclusters:
                    div_txt = self.clean_text(' '.join([sents_div[i].text for i in subcluster]))
                    div_len = len(div_txt)
                    
                    if div_len < 60 or div_len > 512:
                        continue
                    
                    self.clusters_lens.append(div_len)
                    self.final_texts.append(div_txt)
                    
            else:
                # ascii_text = cluster_txt.encode('ascii', 'ignore').decode('ascii')
                # # Replace various forms of newlines and adjacent spaces
                # cleaned_txt = ascii_text.replace('\r\n', ' ').replace('\n', ' ').replace('\r', ' ')
                # # Further strip any leading/trailing whitespace that could have been left over
                # cleaned_txt = cleaned_txt.strip()
                # split_text = cleaned_txt.split("\n")
                # clean_text = "".join(split_text)
                
                self.clusters_lens.append(cluster_len)
                self.final_texts = [x.replace('\n', '').encode('ascii', 'ignore').decode('ascii') for x in (self.final_texts + [cluster_txt])]
                self.final_embeddings.append(self.get_embedding(self.final_texts[-1]))

                if len(self.final_texts) > 120:
                    ##self.final_texts = [x.replace('\n', '') for x in self.final_texts]
                    #print(self.final_texts)
                    #print(list(zip(self.final_texts, self.final_embeddings)))
                    self.this_table.upsert(list(zip(self.final_texts, self.final_embeddings)))
                    self.final_texts = []
                    self.final_embeddings = []
                
        return self.final_texts

def main(file_path):
    doc_chunker = DocChunker()
    results = doc_chunker.process_pdf(file_path)
    print(results[20])
    with open('output.json', 'w') as f:  # Write the results to a file
        json.dump(results, f)

if __name__ == "__main__":
    main(sys.argv[1])
```