# Project export: InShort

This document was generated by HackStack to give an AI agent context about a hackathon project. Sections are labeled with their provenance; content marked as truncated was cut to keep this document small.

## Project metadata

- Hackathon: UC Berkeley AI Hackathon 2025
- Tagline: Learn about how Congress will affect you, in real time.
- Devpost: https://devpost.com/software/inshort
- GitHub: https://github.com/siddharthhtyagi/InShort
- Demo: https://github.com/tede12/InShort
- Video: https://www.youtube.com/embed/Q7qhtpakh5w?enablejsapi=1&hl=en_US&rel=0&start=&version=3&wmode=transparent
- Team: 4 GitHub contributor(s) — Suhaas Surapaneni (8 commits), tede12 (7 commits), siddharthtyagi2211 (7 commits), Andy Phu (3 commits)

## Devpost submission (written by the team)

### Inspiration

A recent survey showed that 70% of Americans lacked basic civic literacy on topics related to the US Government. In a democracy, it’s not just a privilege but almost a necessity to form opinions that help shape the government we follow. source Today’s generation has a new way of absorbing information, and it’s through short forms of content (Reels/Tiktoks/etc). So the idea of reading through a newspaper, let alone an entire bill, feels completely out of the question. But these same bills have the power to affect our day-to-day, and make a lasting impact on our lives. So why shouldn’t there be a way for our generation to stay informed in a format that actually fits how we consume content? The Solution We built an iOS app that meets you where you are. InShort uses AI to learn what matters to you: your lifestyle, your interests, your needs, and then explains the bills that affect you most, simply and clearly. It’s like TikTok but for laws and political awareness. InShort is designed for convenience. You only need to tell us the causes you care about, and we’ll take care of notifying you. If you feel strongly about a pending law, we give you the ability to contact your lawmaker with 1 click.

### How we built it

We started off with an initial MVP using iOS native frameworks, and then we used Congress API to scrape over 500+ Congressional bills from which we made a vector database so that we can establish a recommender system for people to find the bills that are most relevant to them. Then we added an AI chatbot using Groq personalized to the user, and familiar with Congressional law. For our recommender system, we fine-tuned the parameters for the cosine similarity search function to improve relevance for the user. We hosted the backend for these AI features on an OVH server. Then we established a notification system that notifies people about the bills that they would care about. Lastly, we improved user experience by providing people with a forum for public discourse that uses reinforcement learning to suggest more of what you would engage with. An algorithm suited for you. And finally, we added the option to contact your lawmaker with a click of a button. Who we are We met on the Slack Intro channel, where we quickly connected over a shared passion for using technology to make civic engagement easier. What started as a group chat with 4 strangers, quickly turned into hours of brainstorming, building and problem solving together! For some of us, this was our first hackathon!

### Challenges we ran into

We started a step behind. Our team spent the first few hours just brainstorming, throwing out ideas, scrapping them, and trying to find something we could all genuinely believe in. We didn’t want to build just another project for the sake of it, we wanted something that would stick. That took time. We also had issue being that most of us have never worked on an IOS application before, and setting up a back-end server in within hackathon time constraints was pretty hard. How we overcame them Once we locked in on the idea, we divided work based on strengths, stayed up all night debugging the backend and connecting it to the frontend, and somehow put it all together just in time. (everything was broken at 4 AM). Integration was honestly the hardest part, between the AI API, the recommender system, and the app UI, getting it to feel seamless was rough. But we made it work. Next steps! InShort was limited by the scope of this Hackathon, but we hope to partner with nonprofits like Pew research in the future to spread civic engagement for people who grew up in the digital generation. Furthermore, this would generally be helpful for everyday people to keep track of the changes that matter.

## README (from the GitHub repository)

# InShort: Personalized Bill Recommendations

InShort is a mobile application designed to help users discover and understand legislative bills that are relevant to their personal interests and location. Using a powerful AI backend, the app delivers personalized bill recommendations and custom-tailored summaries, making complex legislation accessible and engaging.

![Architecture Diagram](https://mermaid.ink/svg/eyJjb2RlIjoiZ3JhcGggVERcbiAgICBzdWJncmFwaCBGcm9udGVuZCAtIGlPUyBBcHBcbiAgICAgICAgQVtTd2lmdFVJIFZpZXdzXSAtLT4gQltWaWV3TW9kZWxzXTtcbiAgICAgICAgQiAtLT4gQ1tTZXJ2aWNlcyAtIEJpbGxTZXJ2aWNlLCBVc2VyU2VydmljZV07XG4gICAgZW5kXG5cbiAgICBzdWJncmFwaCBCYWNrZW5kIC0gUHl0aG9uIFNlcnZlclxuICAgICAgICBFW0Zhc3RBUEkgRW5kcG9pbnQgL3JlY29tbWVuZGF0aW9ucy9dIC0tPiBGW0JpbGxSZWNvbW1lbmRlcl07XG4gICAgICAgIEYgLS0-IEdbUGluZWNvbmUgVmVjdG9yIERCXTtcbiAgICAgICAgRiAtLT4gSFtPcGVuQUkgZm9yIEVtYmVkZGluZ3NdO1xuICAgICAgICBGIC0tPiBJW0dyb3EgZm9yIFN1bW1hcmllc107XG4gICAgZW5kXG5cbiAgICBzdWJncmFwaCBEYXRhIEZsb3dcbiAgICAgICAgSltVc2VyIFByb2ZpbGUgQ2hhbmdlc10gLS0-IEI7XG4gICAgICAgIEMgLS0gSFRUUCBSZXF1ZXN0IC0tPiBFO1xuICAgICAgICBFIC0tIEpTT04gUmVzcG9uc2UgLS0-IEM7XG4gICAgZW5kXG5cbiAgICBzdHlsZSBGcm9udGVuZCBmaWxsOiNFNkY3RkYsc3Ryb2tlOiNCM0Q5RkYsc3Ryb2tlLXdpZHRoOjJweFxuICAgIHN0eWxlIEJhY2tlbmQgZmlsbDojRThGNUU5LHN0cm9rZTojQTVENkE3LHN0cm9rZS13aWR0aDoycHhcbiAgICBzdHlsZSBEYXRhIEZsb3cgZmlsbDojRkZGOEUxLHN0cm9rZTojRkZFQ0IzLHN0cm9rZS13aWR0aDoycHgiLCJtZXJtYWlkIjp7InRoZW1lIjoiZGVmYXVsdCJ9LCJ1cGRhdGVFZGl0b3IiOmZhbHNlLCJhdXRvU3luYyI6dHJ1ZSwidXBkYXRlRGlhZ3JhbSI6ZmFsc2V9)

---

## Features

-   **Personalized Recommendations**: Leverages a vector database (Pinecone) to find bills that semantically match a user's unique profile, including their interests, location, and occupation.
-   **AI-Generated Summaries**: Uses a Large Language Model (Groq) to generate concise, easy-to-understand summaries of bills, personalized to be relevant to the user.
-   **Dynamic Profile Updates**: Seamlessly updates recommendations when a user changes their profile interests.
-   **SwiftUI Frontend**: A modern, reactive iOS application built with SwiftUI and the MVVM pattern.
-   **FastAPI Backend**: A robust and efficient Python backend serving the recommendation engine.

---

## Architecture

The project is a monorepo containing two main components: a SwiftUI frontend and a Python backend.

### Frontend (iOS App)

-   **Language**: Swift
-   **UI Framework**: SwiftUI
-   **Architecture**: Model-View-ViewModel (MVVM)
    -   **Views**: SwiftUI views define the UI and bind to ViewModel properties.
    -   **ViewModels**: Contain the presentation logic and state for the views.
    -   **Models**: Represent the data structures of the app (e.g., `Bill`, `UserProfile`).
    -   **Services**: Handle networking (`BillService`, `UserService`) and other shared logic (`NotificationService`).
-   **Data Persistence**: `UserDefaults` is used to persist the user's profile locally.
-   **Concurrency**: `Combine` and `async/await` are used for managing asynchronous operations and state updates.

### Backend (Python Server)

-   **Framework**: FastAPI
-   **Database**: Pinecone (Vector Database for semantic search)
-   **AI Services**:
    -   **OpenAI**: Used to generate vector embeddings for text data.
    -   **Groq**: Used to generate personalized bill summaries with a fast LLM.
-   **Core Logic**: The `BillRecommender` class in `RAG/` encapsulates the logic for querying the vector database and formatting results.

---

## Getting Started

### Prerequisites

-   macOS with Xcode installed.
-   Python 3.8+
-   `pip` for Python package management.

### 1. Backend Setup

First, set up and run the Python server.

1.  **Navigate to the Backend Directory**:
    ```bash
    cd InShort
    ```

2.  **Create an Environment File**:
    Create a file named `.env` in the `InShort/` directory and add your API keys:
    ```
    PINECONE_API_KEY="YOUR_PINECONE_KEY"
    OPENAI_API_KEY="YOUR_OPENAI_KEY"
    GROQ_API_KEY="YOUR_GROQ_KEY"
    ```

3.  **Install Dependencies**:
    ```bash
    pip install -r requirements.txt
    ```

4.  **Populate the Database**:
    Run the upsert script to populate your Pinecone index with bill data. Make sure your Pinecone index is configured for 384 dimensions to match the `text-embedding-3-small` model.
    ```bash
    python3 RAG/pcupsert.py
    ```

5.  **Run the Server**:
    ```bash
    uvicorn api:app --reload
    ```
    The server will be running at `http://127.0.0.1:8000`.

### 2. Frontend Setup

With the backend running, you can now launch the iOS application.

1.  **Navigate to the Frontend Project**:
    ```bash
    cd ../InShortFrontEnd
    ```

2.  **Open in Xcode**:
    Open the `InShort.xcodeproj` file in Xcode.
    ```bash
    open InShort.xcodeproj
    ```

3.  **Build and Run**:
    Select an iOS Simulator (e.g., iPhone 15 Pro) or a physical device and press the "Run" button (or `Cmd+R`). The app will launch and connect to your local backend.

---

## Project Structure

```
.
├── InShort/                  # Python Backend
│   ├── api.py                # FastAPI application endpoints
│   ├── requirements.txt      # Python dependencies
│   ├── .env.example          # Example environment file
│   └── RAG/
│       ├── billRecommender.py  # Core recommendation logic
│       └── pcupsert.py         # Script to populate Pinecone
│
└── InShortFrontEnd/          # iOS Frontend
    └── InShort/
        ├── InShort.xcodeproj   # Xcode Project
        └── InShort/
            ├── Models/         # Data models (Bill, UserProfile)
            ├── ViewModels/     # ViewModel layer (NewsViewModel, etc.)
            ├── Views/          # SwiftUI views
            └── Services/       # Networking and data services
```

## Prerequisites

1. **Congress.gov API Key**: Get a free API key from [Congress.gov API](https://api.congress.gov/)
   - Visit https://api.congress.gov/
   - Sign up for a free account
   - Generate your API key
   - Already there in google docs file

2. **Groq API Key**: Get an API key from [Groq](https://console.groq.com/)
   - Visit https://console.groq.com/
   - Sign up for an account
   - Generate your API key
   - Already there in google docs file

3. **Python Dependencies**: Install the required packages
   ```bash
   pip install -r requirements.txt
   ```

## Project Structure

- `full_bill_scraper.py` - Scrapes comprehensive bill data from Congress.gov
- `inshort_summarizer.py` - Generates personalized AI summaries using Groq
- `inshort_bills.json` - Bill data (generated by scraper)
- `requirements.txt` - Python dependencies
- `README.md` - This file

## Quick Start

### Step 1: Set up API Keys

Set your API keys as environment variables:

```bash
# Set Congress.gov API key
export CONGRESS_API_KEY="your_congress_api_key_here"

# Set Groq API key
export GROQ_API_KEY="your_groq_api_key_here"
```

### Step 2: Scrape Bill Data

Run the bill scraper to collect comprehensive bill data:

```bash
python3 full_bill_scraper.py
```

This will:
- Fetch the most recent 100 bills from Congress.gov
- Collect full details including sponsors, actions, amendments, text, and more
- Save the data to `inshort_bills.json`

### Step 3: Generate AI Summaries

Run the AI summarizer to create personalized summaries:

```bash
python3 inshort_summarizer.py
```

This will:
- Load the bill data from `inshort_bills.json`
- Generate personalized summaries for different user profiles
- Show how the same bill affects different users differently

## Detailed Usage

### Bill Scraper Options

The `full_bill_scraper.py` script collects comprehensive bill data:

```bash
# Basic usage (collects 100 bills)
python3 full_bill_scraper.py

# The script automatically:
# - Fetches bills from the 119th Congress
# - Collects full details for each bill
# - Saves data to inshort_bills.json
```

### AI Summarizer Features

The `inshort_summarizer.py` script generates personalized summaries for:

1. **Sarah (25yo, Texas, recent graduate)** - Focus on student loans,

[README truncated for size]

## Detected evidence (automated analysis)

Indexed codebase: 14 recognized source files, 121 KB.
- FastAPI (technology) — detected in the code
- Hugging Face (technology) — detected in the code
- LangChain (technology) — detected in the code
- OpenAI (technology) — detected in the code
- Python (language) — detected in the code
- PyTorch (technology) — detected in the code
- PostgreSQL (technology) — claimed on Devpost, not found in the code
- Swift (language) — claimed on Devpost, not found in the code
- Tailwind CSS (technology) — claimed on Devpost, not found in the code
- TypeScript (language) — claimed on Devpost, not found in the code
- Vercel (technology) — claimed on Devpost, not found in the code

## Codebase structure (from repository index)

### Files (19 of 19)

```
.gitignore
api.py
debug_upsert.py
example.env
fetch_bills.py
full_bill_scraper.py
inShort_agent.py
inshort_bills.json
inshort_summarizer.py
RAG/billRecommender.py
RAG/init.sh
RAG/pcsearch.py
RAG/pcstart.py
RAG/pcupsert.py
README.md
requirements.prod.txt
requirements.txt
test_bill_recommender.py
upsert_bills.py
```

### Dependencies

- requirements.txt: annotated-types@==0.7.0, anyio@==4.9.0, certifi@==2025.6.15, charset-normalizer@==3.4.2, click@==8.2.1, distro@==1.9.0, fastapi@==0.115.13, filelock@==3.18.0, fsspec@==2025.5.1, groq@==0.28.0, h11@==0.16.0, hf-xet@==1.1.5, httpcore@==1.0.9, httpx@==0.28.1, huggingface-hub@==0.33.0, idna@==3.10, Jinja2@==3.1.6, jiter@==0.10.0, joblib@==1.5.1, jsonpatch@==1.33, jsonpointer@==3.0.0, langchain@==0.3.26, langchain-core@==0.3.66, langchain-openai@==0.3.24, langchain-text-splitters@==0.3.8, langgraph@==0.4.8, langgraph-checkpoint@==2.1.0, langgraph-prebuilt@==0.2.2, langgraph-sdk@==0.1.70, langsmith@==0.4.1, MarkupSafe@==3.0.2, mpmath@==1.3.0, networkx@==3.5, numpy@==2.3.1, openai@==1.90.0, orjson@==3.10.18, ormsgpack@==1.10.0, packaging@==24.2, pillow@==11.2.1, pinecone@==7.2.0, pinecone-client@==6.0.0, pinecone-plugin-assistant@==1.7.0, pinecone-plugin-interface@==0.0.7, pydantic@==2.11.7, pydantic_core@==2.33.2, python-dateutil@==2.9.0.post0, python-dotenv@==1.1.0, PyYAML@==6.0.2, regex@==2024.11.6, requests@==2.32.4, requests-toolbelt@==1.0.0, safetensors@==0.5.3, scikit-learn@==1.7.0, scipy@==1.15.3, sentence-transformers@==4.1.0, setuptools@==80.9.0, six@==1.17.0, sniffio@==1.3.1, SQLAlchemy@==2.0.41, starlette@==0.46.2, sympy@==1.14.0, tenacity@==9.1.2, threadpoolctl@==3.6.0, tiktoken@==0.9.0, tokenizers@==0.21.1, torch@==2.7.1, tqdm@==4.67.1, transformers@==4.52.4, typing_extensions@==4.14.0, typing-inspection@==0.4.1, urllib3@==2.5.0, uvicorn@==0.34.3, xxhash@==3.5.0, zstandard@==0.23.0

### Recent commits (newest first)

- chatbot fixes
- data fixes for json stuff
- fixes
- more fixes
- fixed deprecation warning
- api.py updated for agent chatbot
- added agent chatbot to api.py
- Merge branch 'main' of github.com:siddharthhtyagi/InShort
- cosponsor tool added to chatbot agent
- Merge branch 'main' of https://github.com/siddharthhtyagi/InShort
- Initial commit
- update uvicorn port, remove unused imports, and refine requirements files
- update .gitignore, fix user_profile serialization, add example.env, and update dependencies
- fix packaging issue
- fix req
- billrecommender works p good
- fine-tuning agent
- add pinecone
- Merge branch 'main' of https://github.com/siddharthhtyagi/InShort
- Database with 500 bills

## Key source files (fetched from GitHub, selected and truncated for size)

### requirements.txt

```
annotated-types==0.7.0
anyio==4.9.0
certifi==2025.6.15
charset-normalizer==3.4.2
click==8.2.1
distro==1.9.0
fastapi==0.115.13
filelock==3.18.0
fsspec==2025.5.1
groq==0.28.0
h11==0.16.0
hf-xet==1.1.5
httpcore==1.0.9
httpx==0.28.1
huggingface-hub==0.33.0
idna==3.10
Jinja2==3.1.6
jiter==0.10.0
joblib==1.5.1
jsonpatch==1.33
jsonpointer==3.0.0
langchain==0.3.26
langchain-core==0.3.66
langchain-openai==0.3.24
langchain-text-splitters==0.3.8
langgraph==0.4.8
langgraph-checkpoint==2.1.0
langgraph-prebuilt==0.2.2
langgraph-sdk==0.1.70
langsmith==0.4.1
MarkupSafe==3.0.2
mpmath==1.3.0
networkx==3.5
numpy==2.3.1
openai==1.90.0
orjson==3.10.18
ormsgpack==1.10.0
packaging==24.2
pillow==11.2.1
pinecone==7.2.0
pinecone-client==6.0.0
pinecone-plugin-assistant==1.7.0
pinecone-plugin-interface==0.0.7
pydantic==2.11.7
pydantic_core==2.33.2
python-dateutil==2.9.0.post0
python-dotenv==1.1.0
PyYAML==6.0.2
regex==2024.11.6
requests==2.32.4
requests-toolbelt==1.0.0
safetensors==0.5.3
scikit-learn==1.7.0
scipy==1.15.3
sentence-transformers==4.1.0
setuptools==80.9.0
six==1.17.0
sniffio==1.3.1
SQLAlchemy==2.0.41
starlette==0.46.2
sympy==1.14.0
tenacity==9.1.2
threadpoolctl==3.6.0
tiktoken==0.9.0
tokenizers==0.21.1
torch==2.7.1
tqdm==4.67.1
transformers==4.52.4
typing-inspection==0.4.1
typing_extensions==4.14.0
urllib3==2.5.0
uvicorn==0.34.3
xxhash==3.5.0
zstandard==0.23.0

```

### test_bill_recommender.py

```python
#!/usr/bin/env python3
"""
Test script for the updated BillRecommender with Groq integration and personalization
"""

import os
from dotenv import load_dotenv
from RAG.billRecommender import BillRecommender

def test_personalized_bill_recommender():
    """Test the bill recommender with personalized summaries"""
    
    # Load environment variables
    load_dotenv()
    
    # Check for required API keys
    pinecone_api_key = os.getenv("PINECONE_API_KEY")
    groq_api_key = os.getenv("GROQ_API_KEY")
    
    if not pinecone_api_key:
        print("Error: PINECONE_API_KEY not found in environment variables")
        return
    
    if not groq_api_key:
        print("Error: GROQ_API_KEY not found in environment variables")
        print("Please set it with: export GROQ_API_KEY='your_api_key_here'")
        return
    
    print("✅ API keys found")
    print("🔍 Testing Personalized BillRecommender with Groq integration...")
    print("=" * 80)
    
    # Define different user profiles to test personalization
    user_profiles = [
        {
            "name": "Alex",
            "age": 28,
            "location": "New York",
            "interests": ["tech policy", "privacy", "startups"],
            "occupation": "software engineer"
        },
        {
            "name": "Jennifer",
            "age": 42,
            "location": "Colorado",
            "interests": ["environmental protection", "renewable energy", "public lands"],
            "occupation": "environmental consultant"
        },
        {
            "name": "David",
            "age": 58,
            "location": "Ohio",
            "interests": ["manufacturing", "trade policy", "infrastructure"],
            "occupation": "factory manager"
        }
    ]
    
    # Test each user profile
    for i, user_profile in enumerate(user_profiles, 1):
        print(f"\n🧑‍💼 Test {i}: Personalized recommendations for {user_profile['name']}")
        print(f"   Age: {user_profile['age']}, Location: {user_profile['location']}")
        print(f"   Occupation: {user_profile['occupation']}")
        print(f"   Interests: {', '.join(user_profile['interests'])}")
        print("-" * 80)
        
        try:
            # Initialize recommender with user profile
            recommender = BillRecommender(
                pinecone_api_key, 
                index_name="bills-index", 
                user_profile=user_profile
            )
            
            # Get personalized recommendations
            recommendations = recommender.recommend_bills(
                user_profile['interests'], 
                top_k=2, 
                min_score=0.1
            )
            print(recommendations)
            
        except Exception as e:
            print(f"❌ Error during test {i}: {e}")
        
        print("\n" + "=" * 80)

if __name__ == "__main__":
    test_personalized_bill_recommender() 
```

### upsert_bills.py

```python
import os
import json
from pinecone.grpc import PineconeGRPC as Pinecone
from pinecone import ServerlessSpec
from dotenv import load_dotenv
import openai
from tqdm import tqdm

# --- Configuration ---
load_dotenv()
PINECONE_API_KEY = os.getenv("PINECONE_API_KEY")
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
INDEX_NAME = "bills-index"
MODEL_DIMENSIONS = 384  # For text-embedding-3-small with dimensions=384

# --- Helper Functions ---

def generate_embedding(text: str, model: str = "text-embedding-3-small", dimensions: int = MODEL_DIMENSIONS) -> list[float]:
    """Generate embedding for a given text using OpenAI."""
    try:
        client = openai.OpenAI(api_key=OPENAI_API_KEY)
        response = client.embeddings.create(input=text, model=model, dimensions=dimensions)
        return response.data[0].embedding
    except Exception as e:
        print(f"Error generating embedding: {e}")
        return []

# --- Main Script ---

def main():
    """Main function to upsert bill data into Pinecone."""
    if not PINECONE_API_KEY or not OPENAI_API_KEY:
        print("Error: API keys for Pinecone or OpenAI not found in environment variables.")
        return

    # Initialize Pinecone
    pc = Pinecone(api_key=PINECONE_API_KEY)

    # Create index if it doesn't exist
    if INDEX_NAME not in pc.list_indexes().names():
        print(f"Creating index '{INDEX_NAME}'...")
        pc.create_index(
            name=INDEX_NAME,
            dimension=MODEL_DIMENSIONS,
            metric="cosine",
            spec=ServerlessSpec(cloud="aws", region="us-east-1")
        )
        print("Index created successfully.")
    else:
        print(f"Index '{INDEX_NAME}' already exists.")

    index = pc.Index(INDEX_NAME)

    # Load bill data
    try:
        with open("inshort_bills.json", 'r') as f:
            bills = json.load(f)
    except FileNotFoundError:
        print("Error: 'inshort_bills.json' not found. Please run the scraper first.")
        return

    # Prepare data for upsert
    print(f"Preparing {len(bills)} bills for upsert...")
    vectors_to_upsert = []
    for bill in tqdm(bills, desc="Generating embeddings"):
        # Combine relevant text fields for a comprehensive embedding
        content_to_embed = f"{bill.get('title', '')} {bill.get('summary', '')}"
        
        # Generate embedding
        embedding = generate_embedding(content_to_embed)
        if not embedding:
            print(f"Skipping bill {bill.get('id', 'N/A')} due to embedding generation failure.")
            continue

        # Prepare metadata
        metadata = {
            "title": bill.get("title", "N/A"),
            "summary": bill.get("summary", "No summary available."),
            "sponsor": bill.get("sponsor", {}).get("fullName", "N/A"),
            "congress": bill.get("congress", "N/A"),
            "bill_number": bill.get("number", "N/A"),
            "type": bill.get("type", "N/A"),
            "latest_action": bill.get("latestAction", {}).get("text", "N/A")
        }
        
        vectors_to_upsert.append({
            "id": bill["id"],
            "values": embedding,
            "metadata": metadata
        })

    # Batch upsert to Pinecone
    if not vectors_to_upsert:
        print("No vectors to upsert.")
        return
        
    print(f"Upserting {len(vectors_to_upsert)} vectors to Pinecone in batches...")
    batch_size = 100
    for i in tqdm(range(0, len(vectors_to_upsert), batch_size), desc="Upserting batches"):
        batch = vectors_to_upsert[i:i + batch_size]
        try:
            index.upsert(vectors=batch)
        except Exception as e:
            print(f"Error upserting batch {i//batch_size + 1}: {e}")

    print("\nUpsert complete!")
    print(index.describe_index_stats())

if __name__ == "__main__":
    main()

```

### debug_upsert.py

```python
import json
from pinecone import Pinecone, ServerlessSpec
from sentence_transformers import SentenceTransformer

# Initialize Pinecone
pc = Pinecone(
    api_key="pcsk_GNsF9_TsFND8SaQnpYVPFjibH7YjRGP2x2R6ALf4sfgFhArFKwnMt52m1zzFPATZGQ2sq"
)

# Initialize embedding model
model = SentenceTransformer("all-MiniLM-L6-v2")


def create_index_if_not_exists(index_name, dimension=384):
    """Create Pinecone index if it doesn't exist"""
    if index_name not in pc.list_indexes().names():
        pc.create_index(
            name=index_name,
            dimension=dimension,
            metric="cosine",
            spec=ServerlessSpec(cloud="aws", region="us-east-1"),
        )
        print(f"Created index: {index_name}")
    else:
        print(f"Index {index_name} already exists")

    return pc.Index(index_name)


def process_bill_data(bill_data):
    """Extract relevant text from bill data for embedding"""
    bill = bill_data.get("bill", {})

    # Extract key information
    title = bill.get("title", "")
    bill_number = f"{bill.get('type', '')}-{bill.get('number', '')}"
    sponsor_info = ""
    if bill.get("sponsors"):
        sponsor = bill["sponsors"][0]
        sponsor_info = f"Sponsored by {sponsor.get('fullName', '')}"

    policy_area = bill.get("policyArea", {}).get("name", "")
    latest_action = bill.get("latestAction", {}).get("text", "")
    
    # Extract summary if available
    summary = ""
    if 'summaries_details' in bill_data:
        summaries = bill_data['summaries_details'].get('summaries', [])
        if summaries:
            summary = summaries[0].get('text', '')

    # Create comprehensive text for embedding
    text_content = f"""
    Title: {title}
    Bill Number: {bill_number}
    {sponsor_info}
    Policy Area: {policy_area}
    Latest Action: {latest_action}
    Summary: {summary}
    """

    # Create metadata
    metadata = {
        "bill_number": bill_number,
        "title": title,
        "sponsor": sponsor_info,
        "policy_area": policy_area,
        "congress": bill.get("congress", ""),
        "introduced_date": bill.get("introducedDate", ""),
        "latest_action": latest_action,
        "origin_chamber": bill.get("originChamber", ""),
        "type": bill.get("type", ""),
        "summary": summary,  # Include summary in metadata
    }

    return text_content.strip(), metadata


def debug_upsert_bills_to_pinecone():
    """Debug version to identify upsert issues"""
    try:
        # Read JSON file
        print("Reading JSON file...")
        with open("inshort_bills.json", "r") as f:
            bills_data = json.load(f)

        print(f"Found {len(bills_data)} bills")

        # Create/connect to index
        index = create_index_if_not_exists("bills-index")
        
        # Check initial stats
        print("\nInitial index stats:")
        print(index.describe_index_stats())

        # Process just the first 5 bills for debugging
        print("\nProcessing first 5 bills for debugging...")
        
        for i, bill_data in enumerate(bills_data[:5]):
            try:
                print(f"\nProcessing bill {i+1}...")
                
                # Process bill data
                text_content, metadata = process_bill_data(bill_data)
                print(f"  Title: {metadata.get('title', 'Unknown')[:50]}...")
                print(f"  Text content length: {len(text_content)}")
                
                # Generate embedding
                embedding = model.encode(text_content).tolist()
                print(f"  Embedding dimension: {len(embedding)}")
                
                # Create vector for upsert
                bill_id = f"bill_{i}"
                vector_data = {"id": bill_id, "values": embedding, "metadata": metadata}
                print(f"  Vector ID: {bill_id}")
                
                # Try to upsert single vector
                print(f"  Attempting to upsert...")
                try:
                    result = index.upsert(vectors=[vector_data])
                    print(f"  ✅ Upsert successful: {result}")
                except Exception as upsert_error:
                    print(f"  ❌ Upsert failed: {upsert_error}")
                    continue
                
                # Check stats after each upsert
                stats = index.describe_index_stats()
                print(f"  Index stats after upsert: {stats['total_vector_count']} vectors")

            except Exception as e:
                print(f"  ❌ Error processing bill {i}: {e}")
                continue

        # Final stats check
        print("\nFinal index stats:")
        final_stats = index.describe_index_stats()
        print(final_stats)

    except Exception as e:
        print(f"Error: {e}")


if __name__ == "__main__":
    debug_upsert_bills_to_pinecone() 
```

### inShort_agent.py

```python
from typing import Annotated, Any
from typing_extensions import TypedDict

from langchain_openai import ChatOpenAI
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from langgraph.graph import StateGraph, START, END
from langgraph.prebuilt import ToolNode
from langgraph.graph.message import add_messages
from langgraph.checkpoint.memory import MemorySaver

import fetch_bills as fb

class State(TypedDict):
    messages: Annotated[list, add_messages]

def create_agent_graph():
    """
    Creates and returns a compiled LangGraph agent using the available tools.
    """
    graph_builder = StateGraph(State)

    llm = ChatOpenAI(model_name="gpt-4o")
    agent = llm.bind_tools([
        fb.fetch_congress_bills,
        fb.search_bills_by_keyword,
        fb.get_bill_details,
        fb.get_bill_cosponsors,
        fb.get_bill_summaries
    ])

    def chatbot(state: State):
        message = agent.invoke(state["messages"])
        return {"messages": [message]}

    def route_tools(state: State):
        if isinstance(state, list):
            ai_message = state[-1]
        elif messages := state.get("messages", []):
            ai_message = messages[-1]
        else:
            raise ValueError(f"No messages found in input state to tool_edge: {state}")

        if hasattr(ai_message, "tool_calls") and len(ai_message.tool_calls) > 0:
            tool_name = ai_message.tool_calls[0]["name"]
            if tool_name in [
                "fetch_congress_bills",
                "search_bills_by_keyword",
                "get_bill_details",
                "get_bill_cosponsors",
                "get_bill_summaries"
            ]:
                return tool_name
        return END

    # Create ToolNodes
    fetch_congress_bills_node = ToolNode(tools=[fb.fetch_congress_bills])
    search_bills_by_keyword_node = ToolNode(tools=[fb.search_bills_by_keyword])
    get_bill_details_node = ToolNode(tools=[fb.get_bill_details])
    get_bill_cosponsors_node = ToolNode(tools=[fb.get_bill_cosponsors])
    get_bill_summaries_node = ToolNode(tools=[fb.get_bill_summaries])

    # Add nodes to the graph
    graph_builder.add_node("chatbot", chatbot)
    graph_builder.add_node("fetch_congress_bills", fetch_congress_bills_node)
    graph_builder.add_node("search_bills_by_keyword", search_bills_by_keyword_node)
    graph_builder.add_node("get_bill_details", get_bill_details_node)
    graph_builder.add_node("get_bill_cosponsors", get_bill_cosponsors_node)
    graph_builder.add_node("get_bill_summaries", get_bill_summaries_node)

    # Add edges
    graph_builder.add_conditional_edges("chatbot", route_tools)
    graph_builder.add_edge("fetch_congress_bills", "chatbot")
    graph_builder.add_edge("search_bills_by_keyword", "chatbot")
    graph_builder.add_edge("get_bill_details", "chatbot")
    graph_builder.add_edge("get_bill_cosponsors", "chatbot")
    graph_builder.add_edge("get_bill_summaries", "chatbot")
    graph_builder.add_edge(START, "chatbot")

    memory = MemorySaver()
    graph = graph_builder.compile(checkpointer=memory)
    return graph

def get_system_message(user_profile: dict[str, Any]):
    return SystemMessage(content=f"You are a helpful assistant. "
                         f"You aid the user in finding information about US Congress bills. "
                         f"You can use the following tools to help the user: "
                         f"fetch_congress_bills, search_bills_by_keyword, "
                         f"get_bill_details, get_bill_cosponsors, get_bill_summaries. "
                         f"Try to get a lot of bills for more information. "
                         f"For each response, you should get the title, summary, "
                         f"you should explain how the bill affects the user's life based on their profile, "
                         f"and you should include a link to the bill. "
                         f"Explain the bill in a way that is easy to understand and not too technical. "
                         f"If the bill is not relevant to the user's profile, you should say so. "
                         f"If the bill is relevant to the user's profile, you should say so. "
                         f"Explain how the user's life will be impacted by each bill. "
                         f"The user's profile is {user_profile}")

def run_agent(graph: StateGraph, config: dict, user_input: str, user_profile: dict[str, Any]):
    system_msg = get_system_message(user_profile)
    final_response = None
    for event in graph.stream({"messages": [system_msg, HumanMessage(content=user_input)]},
                              config, stream_mode="values"):
        message = event["messages"][-1]
        if isinstance(message, AIMessage) and not message.tool_calls:
            final_response = message.content
    return final_response

USER_PROFILE = {
    "name": "John Doe",
    "age": 30,
    "gender": "male",
    "location": "San Francisco, CA",
    "interests": ["technology", "politics", "finance", ],
    "political_affiliation": "Democrat",
    "political_views": "Liberal",
    "political_party": "Democratic Party",
    "political_ideology": "Liberal",
}

if __name__ == "__main__":
    graph = create_agent_graph()
    config = {"configurable": {"thread_id": "1"}}
    user_input = input("User: ")
    response = run_agent(graph, config, user_input, USER_PROFILE)
    print(f"Agent: {response}")
```

### inshort_summarizer.py

```python
#!/usr/bin/env python3
"""
InShort Bill Summarizer using Groq

This script takes the collected bill data and uses Groq to generate
personalized summaries for different user profiles.

Usage:
    python3 inshort_summarizer.py
"""

import json
import os
import time
from typing import List, Dict, Optional
from groq import Groq

# Initialize Groq client
# You'll need to set your GROQ_API_KEY environment variable
client = Groq(api_key=os.getenv("GROQ_API_KEY"))

def load_bills(filename: str = "inshort_bills.json") -> List[Dict]:
    """Load bills from JSON file"""
    try:
        with open(filename, 'r') as f:
            bills = json.load(f)
        print(f"Loaded {len(bills)} bills from {filename}")
        return bills
    except FileNotFoundError:
        print(f"Error: {filename} not found. Please run the bill scraper first.")
        return []
    except Exception as e:
        print(f"Error loading bills: {e}")
        return []

def extract_bill_text(bill_data: Dict) -> str:
    """Extract relevant text from bill data for summarization"""
    bill = bill_data.get('bill', {})
    
    # Get basic info
    title = bill.get('title', 'N/A')
    bill_type = bill.get('type', 'N/A')
    bill_number = bill.get('number', 'N/A')
    congress = bill.get('congress', 'N/A')
    
    # Get summary if available
    summary = ""
    if 'summaries_details' in bill_data:
        summaries = bill_data['summaries_details'].get('summaries', [])
        if summaries:
            summary = summaries[0].get('text', '')
    
    # Get latest action
    latest_action = bill.get('latestAction', {})
    action_text = latest_action.get('text', '')
    action_date = latest_action.get('actionDate', '')
    
    # Get policy area
    policy_area = bill.get('policyArea', {}).get('name', 'N/A')
    
    # Get sponsors info
    sponsors_info = ""
    if 'sponsors_details' in bill_data:
        sponsors = bill_data['sponsors_details'].get('sponsors', [])
        if sponsors:
            sponsor = sponsors[0]
            sponsors_info = f"Sponsored by {sponsor.get('fullName', 'Unknown')} ({sponsor.get('party', 'Unknown')}-{sponsor.get('state', 'Unknown')})"
    
    # Get subjects
    subjects = []
    if 'subjects_details' in bill_data:
        subjects_data = bill_data['subjects_details'].get('subjects', [])
        if isinstance(subjects_data, list):
            subjects = [subject.get('name', '') for subject in subjects_data if isinstance(subject, dict)]
        elif isinstance(subjects_data, str):
            subjects = [subjects_data]
    
    # Compile all text
    text_parts = [
        f"Bill: {bill_type}{bill_number} ({congress}th Congress)",
        f"Title: {title}",
        f"Policy Area: {policy_area}",
        f"Subjects: {', '.join(subjects)}",
        f"Sponsor: {sponsors_info}",
        f"Latest Action: {action_text} (Date: {action_date})",
        f"Summary: {summary}"
    ]
    
    return "\n".join(text_parts)

def generate_personalized_summary(bill_text: str, user_profile: Dict) -> str:
    """Generate personalized summary using Groq"""
    
    prompt = f"""
You are InShort, an AI that turns complex government bills into personalized, easy-to-understand summaries.

BILL INFORMATION:
{bill_text}

USER PROFILE:
- Age: {user_profile['age']}
- Location: {user_profile['location']}
- Interests: {', '.join(user_profile['interests'])}
- Occupation: {user_profile['occupation']}

TASK:
Create a personalized summary of this bill that answers:
1. What does this bill do in plain English? (2-3 sentences)
2. How does it specifically affect this user? (1-2 sentences)
3. When do the changes take effect? (if mentioned)
4. Should they care? (Yes/No and why)

FORMAT:
- Keep it conversational and friendly
- Use "you" to address the user directly
- If the bill doesn't affect them, explain why they might still care
- Maximum 4 sentences total
- End with a simple "This affects you" or "This doesn't affect you directly"

EXAMPLE:
"This bill lowers insulin prices for Medicare recipients. If you're over 65, it could save you $200/month starting next year. This affects you."

RESPONSE:
"""
    
    try:
        response = client.chat.completions.create(
            model="llama3-8b-8192",  # Fast and effective for this task
            messages=[
                {"role": "system", "content": "You are InShort, a helpful AI that explains government bills in simple terms."},
                {"role": "user", "content": prompt}
            ],
            temperature=0.7,
            max_tokens=200
        )
        
        return response.choices[0].message.content.strip()
        
    except Exception as e:
        print(f"Error generating summary: {e}")
        return "Unable to generate summary at this time."

def create_user_profiles() -> List[Dict]:
    """Create sample user profiles for testing"""
    return [
        {
            "name": "Sarah",
            "age": 25,
            "location": "Texas",
            "interests": ["student loans", "healthcare", "climate change"],
            "occupation": "recent graduate"
        },
        {
            "name": "Robert",
            "age": 65,
            "location": "Florida",
            "interests": ["medicare", "social security", "veterans"],
            "occupation": "retired"
        },
        {
            "name": "Maria",
            "age": 35,
            "location": "California",
            "interests": ["housing", "education", "immigration"],
            "occupation": "teacher"
        },
        {
            "name": "David",
            "age": 45,
            "location": "New York",
            "interests": ["taxes", "business", "finance"],
            "occupation": "small business owner"
        }
    ]

def main():
    """Main function to process bills and generate summaries"""
    
    # Check if GROQ_API_KEY is set
    if not os.getenv("GROQ_API_KEY"):
        print("Error: GROQ_API_KEY environment variable not set.")
        print("Please set it with: e
[truncated — 1430 more characters]
```

### full_bill_scraper.py

```python
#!/usr/bin/env python3
"""
Full Congress.gov Bills Scraper for InShort

This script pulls complete bill details from congress.gov including all context,
sponsors, actions, amendments, text, and other comprehensive information.

Usage:
    python3 full_bill_scraper.py --max-bills 100
"""

import requests
import json
import time
import argparse
from typing import List, Dict, Optional

API_KEY = "PObLUqeVATUsVD34EGwQagrnuQgBExKjtu1XR4Y6"

class FullBillScraper:
    """Scraper for complete bill details"""
    
    BASE_URL = "https://api.congress.gov/v3"
    
    def __init__(self, api_key: str):
        self.api_key = api_key
        self.session = requests.Session()
        self.session.headers.update({
            'User-Agent': 'InShortBillScraper/1.0'
        })
    
    def get_bills_list(self, offset: int = 0, limit: int = 50) -> Dict:
        """Get list of bills"""
        params = {
            'api_key': self.api_key,
            'format': 'json',
            'offset': offset,
            'limit': limit
        }
        
        url = f"{self.BASE_URL}/bill"
        response = self.session.get(url, params=params)
        
        if response.status_code == 200:
            return response.json()
        else:
            raise Exception(f"API request failed: {response.status_code} - {response.text}")
    
    def get_full_bill_details(self, congress: int, bill_type: str, bill_number: str) -> Optional[Dict]:
        """Get complete bill details including all endpoints"""
        
        bill_type_lower = bill_type.lower()
        base_url = f"{self.BASE_URL}/bill/{congress}/{bill_type_lower}/{bill_number}"
        
        # Get basic bill information
        basic_params = {
            'api_key': self.api_key,
            'format': 'json'
        }
        
        try:
            # 1. Basic bill info
            response = self.session.get(base_url, params=basic_params)
            if response.status_code != 200:
                print(f"  Failed to get basic info for {bill_type}{bill_number}: {response.status_code}")
                return None
            
            bill_data = response.json()
            
            # 2. Get actions
            actions_response = self.session.get(f"{base_url}/actions", params=basic_params)
            if actions_response.status_code == 200:
                bill_data['actions_details'] = actions_response.json()
            
            # 3. Get sponsors
            sponsors_response = self.session.get(f"{base_url}/sponsors", params=basic_params)
            if sponsors_response.status_code == 200:
                bill_data['sponsors_details'] = sponsors_response.json()
            
            # 4. Get cosponsors
            cosponsors_response = self.session.get(f"{base_url}/cosponsors", params=basic_params)
            if cosponsors_response.status_code == 200:
                bill_data['cosponsors_details'] = cosponsors_response.json()
            
            # 5. Get amendments
            amendments_response = self.session.get(f"{base_url}/amendments", params=basic_params)
            if amendments_response.status_code == 200:
                bill_data['amendments_details'] = amendments_response.json()
            
            # 6. Get subjects
            subjects_response = self.session.get(f"{base_url}/subjects", params=basic_params)
            if subjects_response.status_code == 200:
                bill_data['subjects_details'] = subjects_response.json()
            
            # 7. Get summaries
            summaries_response = self.session.get(f"{base_url}/summaries", params=basic_params)
            if summaries_response.status_code == 200:
                bill_data['summaries_details'] = summaries_response.json()
            
            # 8. Get titles
            titles_response = self.session.get(f"{base_url}/titles", params=basic_params)
            if titles_response.status_code == 200:
                bill_data['titles_details'] = titles_response.json()
            
            # 9. Get text (if available)
            text_response = self.session.get(f"{base_url}/text", params=basic_params)
            if text_response.status_code == 200:
                bill_data['text_details'] = text_response.json()
            
            # 10. Get related bills
            related_response = self.session.get(f"{base_url}/related", params=basic_params)
            if related_response.status_code == 200:
                bill_data['related_details'] = related_response.json()
            
            # 11. Get committees
            committees_response = self.session.get(f"{base_url}/committees", params=basic_params)
            if committees_response.status_code == 200:
                bill_data['committees_details'] = committees_response.json()
            
            # 12. Get cbo cost estimates
            cbo_response = self.session.get(f"{base_url}/cbo-cost-estimates", params=basic_params)
            if cbo_response.status_code == 200:
                bill_data['cbo_details'] = cbo_response.json()
            
            return bill_data
            
        except Exception as e:
            print(f"  Error getting full details for {bill_type}{bill_number}: {e}")
            return None

def get_full_bills(max_bills: int = 100) -> List[Dict]:
    """Get complete details for bills"""
    
    scraper = FullBillScraper(API_KEY)
    full_bills = []
    offset = 0
    limit = 10  # Small batches to avoid rate limits
    
    print(f"Fetching complete details for {max_bills} bills for InShort...")
    
    while len(full_bills) < max_bills:
        print(f"\nFetching bills {offset} to {offset + limit}...")
        
        try:
            # Get list of bills
            response = scraper.get_bills_list(offset, limit)
            bills = response.get('bills', [])
            
            if not bills:
                print("No more bills to process")
                break
            
            for bill in bills:
          
[truncated — 3097 more characters]
```

### RAG/init.sh

```shell
source venv/bin/activate
```

### RAG/pcstart.py

```python
# pip install pinecone python-dotenv
import os
from dotenv import load_dotenv
from pinecone import Pinecone

# Load environment variables from .env file
load_dotenv()

pc = Pinecone(api_key=os.getenv("PINECONE_API_KEY"))
index = pc.Index("bills")


# index.upsert() - commented out incomplete call

```

[5 more indexed source files omitted to keep this export small. The full file list is in the Codebase structure section above.]