# Project export: Trialscope AI

This document was generated by HackStack to give an AI agent context about a hackathon project. Sections are labeled with their provenance; content marked as truncated was cut to keep this document small.

## Project metadata

- Hackathon: Cal Hacks 12.0
- Tagline: Your AI-driven clinical trial intelligence platform that reviews, benchmarks, and regenerates protocol drafts into USDM-ready, FDA-aligned docs: Reducing amendments, delays, and cost overruns.
- Devpost: https://devpost.com/software/trialscope-ai
- GitHub: https://github.com/Hilo-Hilo/cal-hacks-new
- Video: https://www.youtube.com/embed/FNYlGTxhrr4?enablejsapi=1&hl=en_US&rel=0&start=&version=3&wmode=transparent
- Result: winner (Regeneron: Grand Prize)
- Team: 2 GitHub contributor(s) — Hanson Wen (39 commits), Cyan Ding (16 commits)

## Devpost submission (written by the team)

### Inspiration

The members of the Trialscope AI team came into CalHacks with all but a singular question: "Why do SO MANY promising discoveries at the bench fail to reach the patients at the bedside?" In fact, roughly 9 in 10 clinical developments fail between starting Phase I trials and receiving regulatory approval. While many of these failures stem from biological uncertainty, a surprisingly large proportion are lost not in the lab, but in clinical trial design and operations. While wet-lab innovation races ahead, trial design still lives in sprawling word documents/PDFs - Even at leading biopharma companies. These protocols span hundreds of pages, presenting scattered trial design information. When foundational design choices are made inside such unstructured, manual systems, trials become vulnerable to avoidable operational risks: misaligned endpoints, impractical timelines, or regulatory gaps that can compromise even the most promising science. The result? Delayed trials, avoidable amendments, and millions of dollars in wasted effort. Enter TrialScope AI. Our mission? Controlling the controllables, by making clinical trial design as intelligent as the science it tests, narrowing the chasm between therapeutic discovery and approval.

### What it does

TrialScope AI transforms messy, unstructured trial drafts into structured and regulator-aligned designs, followed by regenerating improved versions using AI. Upload any Phase II–III trial draft PDF doc. Upload any Phase II–III trial draft PDF doc. Convert it into a machine-readable USDM structure (Schedule of Activities, endpoints, arms, eligibility, etc.) Convert it into a machine-readable USDM structure (Schedule of Activities, endpoints, arms, eligibility, etc.) Generate insights on factors that may slow down trial progress using data from 1M+ historical clinical studies, benchmarking performance metrics such as duration, procedural burden, and amendment likelihood. Generate insights on factors that may slow down trial progress using data from 1M+ historical clinical studies, benchmarking performance metrics such as duration, procedural burden, and amendment likelihood. Identify missing regulatory elements by cross-referencing FDA guidance documents, while highlighting compliance gaps and potential design inefficiencies. Identify missing regulatory elements by cross-referencing FDA guidance documents, while highlighting compliance gaps and potential design inefficiencies. Benchmark trial performance against studies of similar drugs, mechanisms, and phases, providing justification on how design choices (e.g., endpoints, visit frequency, population scope) align with successful precedents. Benchmark trial performance against studies of similar drugs, mechanisms, and phases, providing justification on how design choices (e.g., endpoints, visit frequency, population scope) align with successful precedents. Regenerate an improved, citation-linked draft and export it as USDM-ready JSON/XML for CRO or CTMS integration. Regenerate an improved, citation-linked draft and export it as USDM-ready JSON/XML for CRO or CTMS integration.

### How we built it

On Friday, we built a simple prototype of next.js frontend with a text box that communicated with a backend via an API to allow for natural language queries to the clinicaltrials.gov api. We spent the rest of the day prototyping a full-stack platform to include features such as finding similar studies to demonstrate average metrics such as cohort size, number of endpoints, etc. We also trained an xgboost ML model to predict the probability of a given protocol to go overtime, as we felt this unanticipated weeks to months led to large expenditures for companies. On Friday, we built a simple prototype of next.js frontend with a text box that communicated with a backend via an API to allow for natural language queries to the clinicaltrials.gov api. We spent the rest of the day prototyping a full-stack platform to include features such as finding similar studies to demonstrate average metrics such as cohort size, number of endpoints, etc. We also trained an xgboost ML model to predict the probability of a given protocol to go overtime, as we felt this unanticipated weeks to months led to large expenditures for companies. On Saturday morning, we met with Henry Wei of Regeneron to discuss the direction we were planning on going into and decided we wanted to build a tool to extract insights from clinical trials going from preclinical to phase 1. There, we decided to pivot to stage 2-3 clinical trial protocols, as that time-stage had the most potential to save researchers time through operation parameter changes. We also decided to include a tool to use LLMs to do compliance oversight on a protocol draft utilizing FDA guideline documents. Finally, we decided it would be critical to utilize the USDM format of trial design to standardize all internal operations. On Saturday morning, we met with Henry Wei of Regeneron to discuss the direction we were planning on going into and decided we wanted to build a tool to extract insights from clinical trials going from preclinical to phase 1. There, we decided to pivot to stage 2-3 clinical trial protocols, as that time-stage had the most potential to save researchers time through operation parameter changes. We also decided to include a tool to use LLMs to do compliance oversight on a protocol draft utilizing FDA guideline documents. Finally, we decided it would be critical to utilize the USDM format of trial design to standardize all internal operations. Saturday night, we met with Henry for the second time and gained a lot of insights as to the potential of LLM-improved clinical trials. Throughout the night, we added features such as the LLM oversight tool, a modified xgboost model to predict average time of study, additional graphs to visualize insights, and a USDM rewriting tool that showed side-by-side differences. Saturday night, we met with Henry for the second time and gained a lot of insights as to the potential of LLM-improved clinical trials. Throughout the night, we added features such as the LLM oversight tool, a modified xgboost model to predict average time of study, additional graphs to visualize insights, and a USDM rewriting tool that showed side-by-side differences. Going into Sunday morning, we had a functional product that could extract a plethora of actionable insights from protocol drafts for researchers targeting phase 2-3 clinical trial success. Going into Sunday morning, we had a functional product that could extract a plethora of actionable insights from protocol drafts for researchers targeting phase 2-3 clinical trial success.

### Challenges we ran into

The main challenge we ran into was narrowing down the precise problem scope. As a group that has had minimal experience running clinical trials, we ran mock “customer discovery/insights” interviews with experts in the field (Shout Out Henry Wei, M.D. of Regeneron!) to refine the problem/needs statement and inform our solution landscape + implementation. Pivoting our technical solution based on each interview required a flexible mindset and agile implementation. The main challenge we ran into was narrowing down the precise problem scope. As a group that has had minimal experience running clinical trials, we ran mock “customer discovery/insights” interviews with experts in the field (Shout Out Henry Wei, M.D. of Regeneron!) to refine the problem/needs statement and inform our solution landscape + implementation. Pivoting our technical solution based on each interview required a flexible mindset and agile implementation. There were some technical challenges, mainly on the front-end back-end interfacing, as well as resolving merge conflicts between team members. We were able to resolve these challenges with AI assistance. There were some technical challenges, mainly on the front-end back-end interfacing, as well as resolving merge conflicts between team members. We were able to resolve these challenges with AI assistance.

### Accomplishments we're proud of

We built a complete, production-ready pipeline capable of turning raw, unstructured clinical trial PDFs into optimized, regulator-aligned designs. The system processes multi-hundred-page protocols and benchmarks them against over half a million historical studies in under ten minutes. It implements CDISC USDM v3.0—the same data model used by top pharmaceutical companies—allowing seamless integration into CRO and CTMS workflows. The platform unites semantic embeddings, Claude API, MCP tools, XGBoost modeling, and rule-based validation into one coherent, automated framework. We built a complete, production-ready pipeline capable of turning raw, unstructured clinical trial PDFs into optimized, regulator-aligned designs. The system processes multi-hundred-page protocols and benchmarks them against over half a million historical studies in under ten minutes. It implements CDISC USDM v3.0—the same data model used by top pharmaceutical companies—allowing seamless integration into CRO and CTMS workflows. The platform unites semantic embeddings, Claude API, MCP tools, XGBoost modeling, and rule-based validation into one coherent, automated framework. Our most significant accomplishment wasn’t just the technical complexity, but the practicality of what we built. TrialScope AI functions as a true operational tool, capable of fitting into real biopharma workflows and addressing inefficiencies that cost companies millions. Within just 36 hours, we delivered a prototype that could realistically reduce delays, standardize design choices, and make the process of clinical trial planning as intelligent and data-driven as the science it aims to validate. Our most significant accomplishment wasn’t just the technical complexity, but the practicality of what we built. TrialScope AI functions as a true operational tool, capable of fitting into real biopharma workflows and addressing inefficiencies that cost companies millions. Within just 36 hours, we delivered a prototype that could realistically reduce delays, standardize design choices, and make the process of clinical trial planning as intelligent and data-driven as the science it aims to validate.

### What we learned

Solving a hackathon required full immersion in the problem space - especially as we were working specifically to address needs in the clinical trials space. This required a full-on crash course of the landscape, as further understanding of the technicalities of each clinical trial phase resulted in us challenging and augmenting our own assumptions: whether it was the type of need to address, our entry point into the clinical trial timeline or the modality in which we implemented our solutions. Solving a hackathon required full immersion in the problem space - especially as we were working specifically to address needs in the clinical trials space. This required a full-on crash course of the landscape, as further understanding of the technicalities of each clinical trial phase resulted in us challenging and augmenting our own assumptions: whether it was the type of need to address, our entry point into the clinical trial timeline or the modality in which we implemented our solutions. Our process went beyond simply building a prototype: rather, it involved thinking like a real bio-AI software company. We dived into processes like defining our market position, identifying unmet needs, and translating technical insights into product strategy. This hands-on experience revealed how interdisciplinary problem-solving operates in practice, especially at the intersection of biology and AI. Our process went beyond simply building a prototype: rather, it involved thinking like a real bio-AI software company. We dived into processes like defining our market position, identifying unmet needs, and translating technical insights into product strategy. This hands-on experience revealed how interdisciplinary problem-solving operates in practice, especially at the intersection of biology and AI. The hackathon became less about competition and more about understanding the actual workflow, validation process, and communication style of startups in this space. It showed us how to move from concept to execution within real constraints, and provided the foundation for pursuing further research, development, and innovation in this field beyond CalHacks. The hackathon became less about competition and more about understanding the actual workflow, validation process, and communication style of startups in this space. It showed us how to move from concept to execution within real constraints, and provided the foundation for pursuing further research, development, and innovation in this field beyond CalHacks.

### What's next

Short-Term (Next 3 Months) Multi-version protocol generation with A/B testing and citation tracking for every AI-generated recommendation. Track-change visualization between original and optimized drafts. Expand regulatory coverage with 50+ additional FDA guidance documents, ICH standards, and EMA compliance. Improve ML models to predict enrollment success, dropout risk, and time-to-first-patient-in. Add collaboration tools: multi-user access, role-based permissions, comment threads, and version control. Medium-Term (6–12 Months) Pilot partnerships with biotech and pharma companies to validate platform performance and reduce amendment rates. Integrate with EDC and protocol authoring systems (Veeva, Medidata, Word). Add advanced analytics modules for cost estimation, site selection, and feasibility scoring. Extend global regulatory coverage and introduce multi-language support. Long-Term (12+ Months) Implement generative protocol authoring from simple drug or mechanism prompts. Simulate trial outcomes pre-enrollment using predictive models trained on 1M+ trials. Automate regulatory submission workflows, including IND draft generation and FDA response preparation. Release open-source datasets, model weights, and benchmark frameworks for academic collaboration. Achieve <3-minute full analysis time, >0.90 ML accuracy, and scalable performance for 10,000+ concurrent users with HIPAA-compliant security.

## README (from the GitHub repository)

# TrialScope AI - Clinical Trial Intelligence Platform

**Cal Hacks 12.0 - Regeneron Tech Prize**  
**License**: MIT (Open Source)

---

## 🎤 Elevator Pitch

**Your AI-driven clinical trial intelligence platform that reviews, benchmarks, and regenerates protocol drafts into USDM-ready, FDA-aligned docs: Reducing amendments, delays, and cost overruns.**

---

## 💡 Inspiration

The members of the Trialscope AI team came into CalHacks with all but a singular question: 

"_Why do SO MANY promising discoveries at the bench fail to reach the patients at the bedside?_"

In fact, roughly **9 in 10** clinical developments **fail** between starting Phase I trials and receiving regulatory approval. While many of these failures stem from biological uncertainty, a surprisingly large proportion are lost not in the lab, but in clinical trial design and operations.

While wet-lab innovation races ahead, trial design still lives in sprawling word documents/PDFs - **Even at leading biopharma companies**. These protocols span hundreds of pages, presenting scattered trial design information. 
When foundational design choices are made inside such unstructured, manual systems, **trials become vulnerable to avoidable operational risks**: misaligned endpoints, impractical timelines, or regulatory gaps that can compromise even the most promising science.

_The result?_ **Delayed trials, avoidable amendments, and millions of dollars in wasted effort.**

**Enter TrialScope AI.**
_Our mission?_ Controlling the controllables, by **making clinical trial design as intelligent as the science it tests**, narrowing the chasm between therapeutic discovery and approval.

---

## 🎯 What it does

TrialScope AI transforms messy, unstructured trial drafts into structured and regulator-aligned designs, followed by regenerating improved versions using AI.

### Core Workflow

1. **Upload** any Phase II–III trial draft PDF doc.

2. **Convert** it into a machine-readable USDM structure (Schedule of Activities, endpoints, arms, eligibility, etc.)

3. **Generate insights** on factors that may slow down trial progress using data from **1M+ historical clinical studies**, benchmarking performance metrics such as duration, procedural burden, and amendment likelihood.

4. **Identify missing regulatory elements** by cross-referencing FDA guidance documents, while highlighting compliance gaps and potential design inefficiencies.

5. **Benchmark trial performance** against studies of similar drugs, mechanisms, and phases, providing justification on how design choices (e.g., endpoints, visit frequency, population scope) align with successful precedents.

6. **Regenerate** an improved, citation-linked draft and export it as USDM-ready JSON/XML for CRO or CTMS integration.

### Key Features

#### 📄 Protocol Intelligence System
- **PDF Processing**: Automatic PDF→Markdown→USDM conversion using Claude 4.5 Sonnet
- **Similar Trials Discovery**: Find up to 50 similar trials using natural language matching from 556K+ completed studies
- **Similarity Scoring**: Multi-factor semantic analysis (condition 35%, phase 20%, endpoints 25%, design 20%)
- **Baseline Metrics**: Weighted aggregation from top-K most similar trials for realistic benchmarking
- **Burden Analysis**: Rule-based complexity, recruitment difficulty, and patient burden scoring
- **ML Predictions**: XGBoost models with SHAP explainability for duration overrun risk prediction
- **FDA Compliance**: AI-powered regulatory guidance analysis using actual FDA PDF documents
- **Protocol Optimization**: AI-powered regeneration with citations and regulatory alignment
- **USDM Export**: Industry-standard CDISC format export for seamless CRO integration

#### 🔍 Natural Language Trial Search
- Query 556,743+ clinical trials using natural language powered by Claude AI with MCP tools
- Intelligent fallback between PostgreSQL database and live ClinicalTrials.gov API

**Processing Time**: 5-10 minutes for complete analysis

---

## 🏗️ How we built it

### Architecture Overview

```
┌────────────────────────────────────────────────────────────┐
│                   Frontend (Next.js 14)                     │
│   Trial Search | Protocol Upload | Analysis Dashboard       │
│        Real-time Progress Tracking via WebSockets           │
└──────────────────────────┬─────────────────────────────────┘
                           │ HTTP/REST + WebSockets
                           ▼
┌────────────────────────────────────────────────────────────┐
│                  Backend API (FastAPI)                      │
│  Claude 4.5 | PostgreSQL | MCP Server | ML Models | FDA    │
│  Async Processing | Session Management | WebSocket Updates │
└────────────────────────────────────────────────────────────┘
                           │
                           ▼
┌────────────────────────────────────────────────────────────┐
│              Data Layer & External Services                 │
│  556K Trials DB | FDA Guidance PDFs | ClinicalTrials.gov   │
└────────────────────────────────────────────────────────────┘
```

### Technology Stack

#### Backend (Python/FastAPI)
- **FastAPI** - High-performance async web framework with automatic API documentation
- **Claude 4.5 Sonnet** - AI processing for USDM conversion, FDA analysis, and protocol optimization
- **PostgreSQL 14+** - 556K+ completed trials from ClinicalTrials.gov + session storage
- **sentence-transformers** - Semantic similarity using all-MiniLM-L6-v2 (384-dim embeddings)
- **XGBoost + SHAP** - ML predictions with SHAP TreeExplainer for explainability
- **PyMuPDF + pdfplumber** - Hybrid PDF text extraction (NO OCR required)
- **WebSockets** - Real-time progress updates during long-running analysis
- **psycopg2** - PostgreSQL adapter for efficient database operations

#### Frontend (Next.js/TypeScript)
- **Next.js 14** - React framework with App Router for optimal performance
- **TypeScript** - Type safety across the entire frontend
- **Tailwind CSS** - Utility-first styling for rapid UI development
- **Recharts** - Interactive data visualizations (burden charts, risk gauges, SHAP plots)
- **Lucide React** - Consistent icon system
- **Shadcn UI** - High-quality, accessible component library

#### AI & ML Infrastructure
- **Anthropic Claude API** - USDM conversion (16K token output), FDA analysis, protocol optimization
- **Model Context Protocol (MCP)** - TypeScript/Bun MCP server for intelligent trial discovery
- **CDISC USDM v3.0** - Industry-standard clinical study data model
- **FDA Guidance Library** - 10+ regulatory PDF documents (oncology, general, genetics categories)

#### Data Processing Pipeline
1. **PDF Ingestion**: Hybrid extraction using PyMuPDF + pdfplumber
2. **Markdown Conversion**: Structured text with page markers and tables
3. **USDM Transformation**: Claude AI converts unstructured text to CDISC USDM v3.0 JSON
4. **Parallel Annotation**: 20 concurrent Claude API calls for similar trial annotation
5. **Multi-Factor Scoring**: Semantic embeddings + lexical matching for 4 similarity dimensions
6. **FDA Analysis**: AI-powered document selection + compliance gap identification
7. **ML Prediction**: XGBoost ensemble with SHAP feature attribution
8. **Protocol Regeneration**: Claude extended thinking mode for optimized draft generation

### Key Technical Innovations

#### 1. Two-Stage PDF Processing Pipeline
- **Stage 1**: Python libraries (PyMuPDF + pdfplumber) for text extraction - NO expensive OCR
- **Stage 2**: Claude AI for intelligent structure recognition and USDM conversion
- **Result**: Cost-effective processing with high accuracy on complex medical documents

#### 2. Semantic Similarity Engine
- **Condition Matching** (35%): Sentence-BERT embeddings with cosine similarity
- **Phase Alignment** (20%): Exact match + adjacent phase scoring (e.g., Phase 2 vs Phase 2/3)
- **Endpoint Overlap** (25%): Hybrid semantic + lexical (Jaccard index) matching
- **Design Similarity** (20%): Structural elements (randomiz

[README truncated for size]

## Detected evidence (automated analysis)

Indexed codebase: 127 recognized source files, 1042 KB.
- Anthropic (technology) — detected in the code
- CSS (language) — detected in the code
- FastAPI (technology) — detected in the code
- HTML (language) — detected in the code
- Next.js (technology) — detected in the code
- Python (language) — detected in the code
- React (technology) — detected in the code
- SQL (language) — detected in the code
- Tailwind CSS (technology) — detected in the code
- TypeScript (language) — detected in the code
- PostgreSQL (technology) — claimed on Devpost, not found in the code

## Codebase structure (from repository index)

### Files (120 of 159)

```
.gitignore
api_documentation/claude-api-docs/advanced/01-extended-thinking.md
api_documentation/claude-api-docs/advanced/02-prompt-caching.md
api_documentation/claude-api-docs/advanced/03-vision-pdf.md
api_documentation/claude-api-docs/advanced/04-batch-processing.md
api_documentation/claude-api-docs/core-features/01-tool-use.md
api_documentation/claude-api-docs/core-features/02-mcp.md
api_documentation/claude-api-docs/core-features/03-streaming.md
api_documentation/claude-api-docs/core-features/04-json-mode.md
api_documentation/claude-api-docs/examples/01-tool-use-examples.md
api_documentation/claude-api-docs/examples/02-streaming-examples.md
api_documentation/claude-api-docs/examples/03-mcp-examples.md
api_documentation/claude-api-docs/examples/04-json-examples.md
api_documentation/claude-api-docs/examples/05-complete-apps.md
api_documentation/claude-api-docs/getting-started/01-introduction.md
api_documentation/claude-api-docs/getting-started/02-quickstart.md
api_documentation/claude-api-docs/getting-started/03-messages-api.md
api_documentation/claude-api-docs/README.md
api_documentation/claude-api-docs/reference/01-api-reference.md
api_documentation/claude-api-docs/reference/02-error-handling.md
api_documentation/claude-api-docs/reference/03-best-practices.md
api_documentation/claude-api-docs/reference/04-models-pricing.md
api_documentation/claude-api-docs/STRUCTURE.md
api_documentation/ClinicalTrials.gov Data API - Complete Documentation.md
backend/.gitignore
backend/app/__init__.py
backend/app/config.py
backend/app/database/__init__.py
backend/app/database/session_manager.py
backend/app/main.py
backend/app/ml/__init__.py
backend/app/ml/baseline_ensemble_predictor.py
backend/app/ml/conformal_utils.py
backend/app/ml/create_demo_models.py
backend/app/ml/data_extraction.py
backend/app/ml/feature_engineering.py
backend/app/ml/models/duration_feature_names.txt
backend/app/ml/models/duration_model_metrics.json
backend/app/ml/models/duration_quantile_p10.pkl
backend/app/ml/models/duration_quantile_p25.pkl
backend/app/ml/models/duration_quantile_p50.pkl
backend/app/ml/models/duration_quantile_p75.pkl
backend/app/ml/models/duration_quantile_p90.pkl
backend/app/ml/models/duration_shap_p10.pkl
backend/app/ml/models/duration_shap_p25.pkl
backend/app/ml/models/duration_shap_p50.pkl
backend/app/ml/models/duration_shap_p75.pkl
backend/app/ml/models/duration_shap_p90.pkl
backend/app/ml/models/feature_names.txt
backend/app/ml/models/gradient_boosting.pkl
backend/app/ml/models/model_metrics.json
backend/app/ml/models/shap_explainer.pkl
backend/app/ml/train_duration_quantile_models.py
backend/app/ml/train_model.py
backend/app/models/__init__.py
backend/app/models/query.py
backend/app/routes/__init__.py
backend/app/routes/analysis.py
backend/app/routes/protocol.py
backend/app/routes/query.py
backend/app/routes/sessions.py
backend/app/services/__init__.py
backend/app/services/annotation_service.py
backend/app/services/baseline_generator.py
backend/app/services/burden_calculator.py
backend/app/services/claude.py
backend/app/services/ctgov_api.py
backend/app/services/fda_report_analyzer.py
backend/app/services/mcp_client.py
backend/app/services/mcp_tools.py
backend/app/services/ml_predictor.py
backend/app/services/postgres.py
backend/app/services/protocol_optimizer.py
backend/app/services/protocol_parser.py
backend/app/services/similarity_engine.py
backend/app/websocket_manager.py
backend/inspect_db.py
backend/README.md
backend/setup_env.sh
backend/tests/fixtures/sample_analysis_report.json
backend/tests/fixtures/sample_protocol_usdm.json
backend/tests/run_comprehensive_tests.py
backend/tests/test_annotation_service.py
backend/tests/test_baseline_generator.py
backend/tests/test_burden_calculator_comprehensive.py
backend/tests/test_burden_calculator.py
backend/tests/test_curly_bracket_parsing.py
backend/tests/test_fda_analyzer.py
backend/tests/test_fda_report_analyzer.py
backend/tests/test_feature_engineer.py
backend/tests/test_indexed_selection_demo.py
backend/tests/test_mcp_client_api.py
backend/tests/test_ml_predictor_comprehensive.py
backend/tests/test_ml_predictor.py
backend/tests/test_pdf_parser.py
backend/tests/test_protocol_parser.py
backend/tests/test_results.json
backend/tests/test_services.py
backend/tests/test_similarity_engine.py
backend/tests/test_status_normalization.py
database/.env
database/check_database.sh
database/data_dictionary.csv
database/install_postgres.sh
database/migration_add_optimizing_status.sql
database/migration_add_reasoning_column.sql
database/migration_add_updated_at_column.sql
database/nlm_protocol_definitions.html
database/nlm_results_definitions.html
database/README.md
database/schema_protocol_intelligence.sql
database/setup_database.sh
docs/backend/backend.md
docs/backend/QUICKSTART.md
docs/backend/README.md
fda/.DS_Store
front-end/.gitignore
front-end/app/globals.css
front-end/app/layout.tsx
front-end/app/page.tsx
[39 more files omitted for size]
```

### Dependencies

- front-end/package.json: @radix-ui/react-slot@^1.2.3, @tailwindcss/postcss@^4, @tanstack/react-query@^5.17.0, @types/node@^20, @types/react@^19, @types/react-dom@^19, class-variance-authority@^0.7.1, clsx@^2.1.1, eslint@^9, eslint-config-next@16.0.0, lucide-react@^0.548.0, next@16.0.0, react@19.2.0, react-dom@19.2.0, recharts@^2.10.3, tailwind-merge@^3.3.1, tailwindcss@^4, tw-animate-css@^1.4.0, typescript@^5
- requirements.txt: anthropic@==0.39.0, fastapi@==0.115.5, httpx@==0.27.0, joblib@==1.5.2, markdownify@==0.12.1, numpy@==2.3.4, pandas@==2.3.3, pdfplumber@==0.11.7, psycopg2-binary@==2.9.10, pydantic@==2.10.1, PyMuPDF@==1.26.5, python-dotenv@==1.0.1, python-multipart@==0.0.20, scikit-learn@==1.7.2, sentence-transformers@==5.1.2, shap@==0.49.1, SQLAlchemy@==2.0.44, usdm@==0.64.0, uvicorn[standard]@==0.32.1, websockets@==15.0.1, xgboost@==3.1.1

### Recent commits (newest first)

- Remove CLAUDE.md from repository and add to .gitignore
- Remove .cursor and hidden files from git tracking and update .gitignore
- Clean up Regeneron track files and finalize README
- Add elevator pitch to README
- Update README to match pitch with comprehensive project details
- Clean up temporary files and documentation
- Fix FDA analyzer: correct Claude model name and resolve SDK compatibility issues
- fix fda
- Revert context window upgrade changes to fix timeout errors
- Remove caching from analysis.py - use remote version
- Merge remote changes: Resolve conflicts in session_manager and analysis routes
- Critical fixes: context window upgrade and comprehensive improvements
- add more pdfs, get rid of caching
- feat: enhance protocol diff viewer with row-aligned layout and alphabetical sorting
- Implement progressive search retry strategy and phase normalization
- delete pdfs, make fda better
- Merge branch 'main' of https://github.com/Hilo-Hilo/cal-hacks-new
- Add protocol optimizer feature with AI-powered recommendations
- Merge branch 'main' of https://github.com/Hilo-Hilo/cal-hacks-new
- fix some ui bugs

## Key source files (fetched from GitHub, selected and truncated for size)

### PRD_COMPLETE.md

```markdown
# Clinical Trials MVP - Product Requirements Document

**Version**: 1.0 (MVP)  
**Date**: October 25, 2025  
**Target**: Cal Hacks 12.0 - Regeneron Tech Prize  
**License**: MIT (Open Source)

---

## Executive Summary

A simple, functional system that enables users to query clinical trials data using natural language. The system has three core components:
1. **Frontend** (Next.js) - User interface for queries
2. **MCP Server** (TypeScript/Bun) - Enables Claude AI to query PostgreSQL and ClinicalTrials.gov API
3. **Data Sources** - Local PostgreSQL database + ClinicalTrials.gov API

**Key Value Proposition**: Query 556,743+ clinical trials using natural language through Claude AI, with results from both local database and live API.

---

## System Architecture

### Simple 3-Tier Architecture

```
┌─────────────────────────────────────────────────────────────┐
│                    Frontend (Next.js)                        │
│              User enters natural language query              │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            │ HTTP/REST
                            │
┌───────────────────────────▼─────────────────────────────────┐
│                   Backend API (FastAPI)                      │
│                                                               │
│  ┌────────────────────────────────────────────────────────┐ │
│  │              Claude AI Integration                     │ │
│  │         (Anthropic API with MCP tools)                 │ │
│  └────────────────────────────────────────────────────────┘ │
│                            │                                 │
│          ┌─────────────────┴──────────────────┐            │
│          ▼                                     ▼            │
│  ┌──────────────┐                    ┌────────────────┐   │
│  │  MCP Tools   │                    │  MCP Tools     │   │
│  │  (Postgres)  │                    │  (API Client)  │   │
│  └──────────────┘                    └────────────────┘   │
└─────────────────────────────────────────────────────────────┘
          │                                     │
          ▼                                     ▼
┌──────────────────┐            ┌──────────────────────┐
│   PostgreSQL     │            │ ClinicalTrials.gov   │
│   (556K trials)  │            │        API           │
└──────────────────┘            └──────────────────────┘
```

### Component Responsibilities

| Component | Technology | Purpose |
|-----------|-----------|----------|
| **Frontend** | Next.js 14, TypeScript, Tailwind | User interface, query input |
| **Backend API** | FastAPI (Python) | Route requests, integrate Claude |
| **Claude AI** | Anthropic API (Sonnet 4) | Process natural language, use tools |
| **MCP Tools** | Python functions | Search database, call external API |
| **PostgreSQL** | PostgreSQL 14 | Local clinical trials data |
| **External API** | ClinicalTrials.gov API v2 | Live data acce
[truncated — 26643 more characters]
```

### docs/backend/QUICKSTART.md

```markdown
# Quick Start Guide

## 🚀 Get Started in 5 Minutes

### Prerequisites
- Python 3.11+
- PostgreSQL 14+
- Anthropic API key ([Get one here](https://console.anthropic.com/))

### Step 1: Setup Backend

```bash
cd backend
./setup.sh
```

This will:
- Create a virtual environment
- Install all dependencies
- Create a `.env` file template

### Step 2: Configure Environment

Edit `.env` and add your credentials:

```bash
# Required - Get from https://console.anthropic.com/
ANTHROPIC_API_KEY=sk-ant-your-actual-key-here

# Required - Your PostgreSQL password
POSTGRES_PASSWORD=your_postgres_password

# Optional - Use defaults or customize
POSTGRES_HOST=localhost
POSTGRES_PORT=5432
POSTGRES_DATABASE=clinical_trials
POSTGRES_USER=postgres
```

### Step 3: Setup Database

```bash
# Create database
createdb clinical_trials

# Restore data (choose one based on your file type)
pg_restore -d clinical_trials -v ../database/postgres.dmp
# OR if it's a SQL file:
# psql -d clinical_trials -f ../database/postgres.dmp

# Verify
psql clinical_trials -c "SELECT COUNT(*) FROM studies;"
```

Expected output: ~556,743 studies

### Step 4: Run the Server

```bash
./run.sh
```

Or manually:
```bash
source venv/bin/activate
uvicorn app.main:app --reload --port 8000
```

The API will be available at: http://localhost:8000

### Step 5: Test It!

#### Health Check
```bash
curl http://localhost:8000/health
```

Expected response:
```json
{
  "status": "healthy",
  "database": "connected",
  "api_key": "configured"
}
```

#### Query Clinical Trials
```bash
curl -X POST http://localhost:8000/api/query \
  -H "Content-Type: application/json" \
  -d '{"query": "Find Phase 3 lung cancer trials in California"}'
```

#### API Documentation
Visit http://localhost:8000/docs for interactive API documentation.

## 🎯 Example Queries

Try these natural language queries:

1. "Find all recruiting Phase 3 lung cancer trials"
2. "Show me diabetes studies in California"
3. "What are the most recent cancer immunotherapy trials?"
4. "Find studies sponsored by NIH"
5. "Show trials for COVID-19 treatments"
6. "What Phase 2 trials are recruiting in New York?"

## 🐛 Troubleshooting

### Database Connection Failed
```bash
# Check PostgreSQL is running
pg_isready

# Test connection
psql -h localhost -U postgres -d clinical_trials -c "SELECT 1"
```

### Anthropic API Key Error
```bash
# Verify key is set
cat .env | grep ANTHROPIC_API_KEY

# Test the key
curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: YOUR_KEY_HERE" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-4-20250514","max_tokens":10,"messages":[{"role":"user","content":"Hi"}]}'
```

### Import Errors
```bash
# Ensure virtual environment is activated
source venv/bin/activate

# Reinstall dependencies
pip install -r requirements.txt
```

## 📚 Next Steps

- Read the [full README](README.md) for detailed documentation
- Check the [API documentation](http://localhost:8000/docs) for endpoint details
- Integrate wit
[truncated — 506 more characters]
```

### requirements.txt

```
fastapi==0.115.5
uvicorn[standard]==0.32.1
python-multipart==0.0.20
anthropic==0.39.0
psycopg2-binary==2.9.10
httpx==0.27.0
python-dotenv==1.0.1
pydantic==2.10.1

# PDF Processing (NO OCR - Python only)
PyMuPDF==1.26.5
pdfplumber==0.11.7
markdownify==0.12.1

# CDISC USDM Standard (Official Library)
usdm==0.64.0

# ML & Similarity
sentence-transformers==5.1.2
scikit-learn==1.7.2
xgboost==3.1.1
shap==0.49.1
numpy==2.3.4
pandas==2.3.3

# Database
SQLAlchemy==2.0.44

# Async & WebSockets
websockets==15.0.1

# Additional utilities
joblib==1.5.2


```

### front-end/package.json

```
{
  "name": "front-end",
  "version": "0.1.0",
  "private": true,
  "scripts": {
    "dev": "next dev",
    "build": "next build",
    "start": "next start",
    "lint": "eslint"
  },
  "dependencies": {
    "@radix-ui/react-slot": "^1.2.3",
    "@tanstack/react-query": "^5.17.0",
    "class-variance-authority": "^0.7.1",
    "clsx": "^2.1.1",
    "lucide-react": "^0.548.0",
    "next": "16.0.0",
    "react": "19.2.0",
    "react-dom": "19.2.0",
    "recharts": "^2.10.3",
    "tailwind-merge": "^3.3.1"
  },
  "devDependencies": {
    "@tailwindcss/postcss": "^4",
    "@types/node": "^20",
    "@types/react": "^19",
    "@types/react-dom": "^19",
    "eslint": "^9",
    "eslint-config-next": "16.0.0",
    "tailwindcss": "^4",
    "tw-animate-css": "^1.4.0",
    "typescript": "^5"
  }
}

```

### front-end/app/layout.tsx

```typescript
import type { Metadata } from "next";
import { Geist, Geist_Mono } from "next/font/google";
import "./globals.css";

const geistSans = Geist({
  variable: "--font-geist-sans",
  subsets: ["latin"],
});

const geistMono = Geist_Mono({
  variable: "--font-geist-mono",
  subsets: ["latin"],
});

export const metadata: Metadata = {
  title: "TrialScope AI - Clinical Trial Intelligence",
  description: "AI-powered clinical trial protocol analysis and intelligence platform",
};

export default function RootLayout({
  children,
}: Readonly<{
  children: React.ReactNode;
}>) {
  return (
    <html lang="en">
      <body
        className={`${geistSans.variable} ${geistMono.variable} antialiased`}
      >
        {children}
      </body>
    </html>
  );
}

```

### backend/app/main.py

```python
"""
Main FastAPI application for Clinical Trials search.
Integrates Claude AI with MCP tools for natural language querying.
"""
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from contextlib import asynccontextmanager
import logging
from .config import config
from .routes import query, protocol, analysis, sessions
from .services.postgres import postgres
from .models.query import HealthResponse

# Configure logging
logging.basicConfig(
    level=getattr(logging, config.LOG_LEVEL.upper()),
    format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)

# Validate configuration
try:
    config.validate()
    logger.info("Configuration validated successfully")
except ValueError as e:
    logger.error(f"Configuration error: {e}")
    raise


@asynccontextmanager
async def lifespan(app: FastAPI):
    """Application lifespan events"""
    # Startup
    logger.info("Starting Clinical Trials API")
    logger.info(f"CORS origins: {config.CORS_ORIGINS}")
    yield
    # Shutdown
    logger.info("Shutting down Clinical Trials API")
    postgres.close()


# Create FastAPI app
app = FastAPI(
    title="Clinical Trials API",
    description="Natural language search for clinical trials using Claude AI with MCP tools",
    version="1.0.0",
    lifespan=lifespan
)

# CORS middleware for frontend
app.add_middleware(
    CORSMiddleware,
    allow_origins=config.CORS_ORIGINS,
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)

# Include routers
app.include_router(query.router, prefix="/api", tags=["query"])
app.include_router(protocol.router, prefix="/api", tags=["protocol"])
app.include_router(analysis.router, prefix="/api", tags=["analysis"])
app.include_router(sessions.router, prefix="/api", tags=["sessions"])


@app.get("/health", response_model=HealthResponse)
async def health():
    """
    Health check endpoint.
    Verifies database connection and API key configuration.
    """
    db_status = "connected" if postgres.test_connection() else "disconnected"
    api_key_status = "configured" if config.ANTHROPIC_API_KEY else "missing"

    return {
        "status": "healthy" if db_status == "connected" and api_key_status == "configured" else "unhealthy",
        "database": db_status,
        "api_key": api_key_status
    }


if __name__ == "__main__":
    import uvicorn
    uvicorn.run(
        "backend.app.main:app",
        host="0.0.0.0",
        port=config.PORT,
        reload=True
    )


```

### front-end/app/page.tsx

```typescript
'use client';

import { useState } from 'react';
import { SearchBox } from '@/components/SearchBox';
import { ResultCard } from '@/components/ResultCard';
import { LoadingSpinner } from '@/components/LoadingSpinner';
import { searchTrials, APIError } from '@/lib/api';
import { QueryResponse } from '@/lib/types';
import { AlertCircle, Database, Sparkles, ChevronDown, ChevronUp } from 'lucide-react';
import { Alert, AlertDescription, AlertTitle } from '@/components/ui/alert';

export default function Home() {
  const [isLoading, setIsLoading] = useState(false);
  const [results, setResults] = useState<QueryResponse | null>(null);
  const [error, setError] = useState<string | null>(null);
  const [isSummaryExpanded, setIsSummaryExpanded] = useState(true);

  const handleSearch = async (query: string) => {
    setIsLoading(true);
    setError(null);
    setResults(null);

    try {
      const data = await searchTrials(query);
      setResults(data);
    } catch (err) {
      if (err instanceof APIError) {
        setError(err.message);
      } else {
        setError('An unexpected error occurred. Please try again.');
      }
      console.error('Search error:', err);
    } finally {
      setIsLoading(false);
    }
  };

  return (
    <div className="min-h-screen bg-background">
      <div className="container mx-auto px-4 py-12 space-y-8">
        {/* Header Navigation */}
        <div className="flex justify-between items-center mb-8">
          <h1 className="text-3xl font-bold">TrialScope AI</h1>
          <div className="flex gap-3">
            <a
              href="/protocol/upload"
              className="px-4 py-2 bg-primary text-primary-foreground rounded-md hover:bg-primary/90 transition-colors"
            >
              Protocol Analysis
            </a>
            <a
              href="/sessions"
              className="px-4 py-2 border rounded-md hover:bg-accent transition-colors"
            >
              View Sessions
            </a>
          </div>
        </div>

        <SearchBox onSearch={handleSearch} isLoading={isLoading} />

        {isLoading && <LoadingSpinner message="Searching clinical trials..." />}

        {error && (
          <div className="w-full max-w-4xl mx-auto">
            <Alert variant="destructive">
              <AlertCircle className="h-4 w-4" />
              <AlertTitle>Error</AlertTitle>
              <AlertDescription>{error}</AlertDescription>
            </Alert>
          </div>
        )}

        {results && !isLoading && (
          <div className="w-full max-w-4xl mx-auto space-y-6">
            {/* Claude's Response */}
            {results.claude_response && (
              <div className="rounded-lg border bg-card">
                <button
                  onClick={() => setIsSummaryExpanded(!isSummaryExpanded)}
                  className="w-full p-6 flex items-center justify-between hover:bg-accent/50 transition-colors rounded-lg"
                >
                  <div className="flex items-center gap-2 text-sm font-medium">
                    <Sparkles className="h-4 w-4 text-primary" />
                    <span>AI Summary</span>
                  </div>
                  {isSummaryExpanded ? (
                    <ChevronUp className="h-4 w-4 text-muted-foreground" />
                  ) : (
                    <ChevronDown className="h-4 w-4 text-muted-foreground" />
                  )}
                </button>
                {isSummaryExpanded && (
                  <div className="px-6 pb-6">
                    <p className="text-sm text-muted-foreground leading-relaxed">
                      {results.claude_response}
                    </p>
                  </div>
                )}
              </div>
            )}

            {/* Results Summary */}
            {results.results && (
              <div className="flex items-center justify-between">
                <div className="flex items-center gap-2">
                  <h2 className="text-2xl font-semibold">
                    {results.results.totalCount === 0
                      ? 'No studies found'
                      : `Found ${results.results.totalCount} ${results.results.totalCount === 1 ? 'study' : 'studies'}`
                    }
                  </h2>
                  {results.tool_used && (
                    <div className="flex items-center gap-1 text-xs text-muted-foreground">
                      <Database className="h-3 w-3" />
                      <span>via {results.tool_used}</span>
                    </div>
                  )}
                </div>
              </div>
            )}

            {/* Results List */}
            {results.results && results.results.studies.length > 0 && (
              <div className="space-y-4">
                {results.results.studies.map((study) => (
                  <ResultCard key={study.nctId} study={study} />
                ))}
              </div>
            )}

            {/* No Results */}
            {results.results && results.results.totalCount === 0 && (
              <div className="text-center py-12 space-y-2">
                <p className="text-muted-foreground">
                  No clinical trials found matching your query.
                </p>
                <p className="text-sm text-muted-foreground">
                  Try adjusting your search terms or using different keywords.
                </p>
              </div>
            )}
          </div>
        )}

        {/* Footer */}
        {!isLoading && !results && (
          <div className="w-full max-w-4xl mx-auto pt-12">
            <div className="text-center space-y-4">
              <div className="inline-flex items-center gap-2 px-4 py-2 rounded-full bg-secondary text-secondary-foreground text-sm">
                <Database className="h-4 w-4" />
                <span>Powered by Claude AI + PostgreSQL + ClinicalTrials.gov API</span>
              </div>
              <p className="text-sm t
[truncated — 184 more characters]
```

### front-end/app/sessions/page.tsx

```typescript
'use client';

import { useEffect, useState } from 'react';
import { useRouter } from 'next/navigation';
import { listSessions, APIError } from '@/lib/api';
import { LoadingSpinner } from '@/components/LoadingSpinner';
import { Alert, AlertDescription, AlertTitle } from '@/components/ui/alert';
import { Badge } from '@/components/ui/badge';
import { Card } from '@/components/ui/card';
import { AlertCircle, Eye, Upload, Trash2, Home, Activity, Loader2 } from 'lucide-react';

export default function SessionsPage() {
  const router = useRouter();
  const [sessions, setSessions] = useState<any[]>([]);
  const [isLoading, setIsLoading] = useState(true);
  const [error, setError] = useState<string | null>(null);
  const [deletingId, setDeletingId] = useState<string | null>(null);

  useEffect(() => {
    console.log('🚀 Sessions page mounted');
    loadSessions(true); // Initial load with loading state
    
    // Poll for updates every 5 seconds (without loading state)
    const pollInterval = setInterval(() => {
      loadSessions(false); // Background refresh, no loading spinner
    }, 5000);
    
    return () => {
      clearInterval(pollInterval);
    };
  }, []);

  const loadSessions = async (showLoading: boolean = true) => {
    console.log('🔄 Loading sessions...');
    
    if (showLoading) {
      setIsLoading(true);
    }
    setError(null);

    try {
      const data = await listSessions();
      console.log('✅ Sessions loaded:', data.length, 'sessions');
      setSessions(data);
    } catch (err) {
      console.error('❌ Sessions load error:', err);
      if (err instanceof APIError) {
        setError(err.message);
      } else {
        setError('Failed to load sessions');
      }
    } finally {
      if (showLoading) {
        console.log('✓ Loading complete, setting isLoading = false');
        setIsLoading(false);
      }
    }
  };

  const handleDelete = async (sessionId: string, filename: string) => {
    if (!confirm(`Delete session "${filename}"? This will remove the uploaded file and all analysis data.`)) {
      return;
    }

    setDeletingId(sessionId);
    
    try {
      const response = await fetch(`http://localhost:8000/api/sessions/${sessionId}`, {
        method: 'DELETE',
      });

      if (!response.ok) {
        throw new Error('Failed to delete session');
      }

      console.log(`✅ Session deleted: ${sessionId}`);
      
      // Reload sessions list
      await loadSessions(true);
      
    } catch (err) {
      console.error('Delete error:', err);
      setError(err instanceof Error ? err.message : 'Failed to delete session');
    } finally {
      setDeletingId(null);
    }
  };

  const getStatusColor = (status: string) => {
    const colors: Record<string, string> = {
      'created': 'bg-slate-500',
      'parsing': 'bg-blue-500',
      'querying': 'bg-indigo-500',
      'annotating': 'bg-purple-500',
      'analyzing': 'bg-violet-500',
      'complete': 'bg-green-500',
      'error': 'bg-red-500',
    };
    return colors[status] || 'bg-gray-500';
  };

  if (isLoading && sessions.length === 0) {
    console.log('🔄 Rendering loading state...');
    return (
      <div className="min-h-screen bg-background">
        <div className="container mx-auto px-4 py-12">
          <div className="max-w-2xl mx-auto space-y-6">
            <div className="text-center space-y-6">
              <h1 className="text-3xl font-bold">Protocol Analysis Sessions</h1>
              
              <Card className="p-12">
                <div className="flex flex-col items-center gap-4">
                  <div className="relative">
                    <Loader2 className="h-16 w-16 animate-spin text-primary opacity-20" />
                    <div className="absolute inset-0 flex items-center justify-center">
                      <Activity className="h-8 w-8 text-primary animate-pulse" />
                    </div>
                  </div>
                  <div className="space-y-2">
                    <p className="text-lg font-medium">Loading your sessions...</p>
                    <p className="text-sm text-muted-foreground">This should only take a moment</p>
                  </div>
                </div>
              </Card>

              <div className="flex gap-3 justify-center">
                <button
                  onClick={() => router.push('/')}
                  className="px-4 py-2 border rounded-md hover:bg-accent transition-colors text-sm"
                >
                  <Home className="h-4 w-4 inline mr-2" />
                  Home
                </button>
                <button
                  onClick={() => loadSessions(true)}
                  className="px-4 py-2 bg-primary text-primary-foreground rounded-md hover:bg-primary/90 text-sm"
                >
                  Retry
                </button>
              </div>
            </div>
          </div>
        </div>
      </div>
    );
  }
  
  console.log('✓ Rendering sessions:', sessions.length, 'sessions found');

  return (
    <div className="min-h-screen bg-background">
      <div className="container mx-auto px-4 py-12 space-y-8">
        {/* Header */}
        <div className="flex items-center justify-between">
          <div>
            <h1 className="text-3xl font-bold">Protocol Analysis Sessions</h1>
            <p className="text-muted-foreground mt-2">
              View and manage your protocol analysis sessions
            </p>
          </div>
          <div className="flex gap-3">
            <button
              onClick={() => router.push('/')}
              className="px-4 py-2 border rounded-md hover:bg-accent transition-colors flex items-center gap-2"
            >
              <Home className="h-4 w-4" />
              Home
            </button>
            <button
              onClick={() => router.push('/protocol/upload')}
              className="px-4 py-2 bg-primary text-primary-foreground rounded-md hover:bg-primary/90 flex items-center
[truncated — 4175 more characters]
```

### front-end/app/protocol/upload/page.tsx

```typescript
'use client';

import { useState } from 'react';
import { useRouter } from 'next/navigation';
import { uploadProtocol, parseProtocol, queryTrials, annotateTrials, getSession, APIError } from '@/lib/api';
import { Upload, FileText, Loader2, CheckCircle2 } from 'lucide-react';
import { Alert, AlertDescription, AlertTitle } from '@/components/ui/alert';
import { Badge } from '@/components/ui/badge';

export default function ProtocolUpload() {
  const router = useRouter();
  const [file, setFile] = useState<File | null>(null);
  const [instructions, setInstructions] = useState('');
  const [isProcessing, setIsProcessing] = useState(false);
  const [currentStep, setCurrentStep] = useState('');
  const [error, setError] = useState<string | null>(null);
  const [progress, setProgress] = useState(0);
  const [isDragging, setIsDragging] = useState(false);

  const handleFileChange = (e: React.ChangeEvent<HTMLInputElement>) => {
    if (e.target.files && e.target.files[0]) {
      setFile(e.target.files[0]);
      setError(null);
    }
  };

  const handleDragOver = (e: React.DragEvent<HTMLDivElement>) => {
    e.preventDefault();
    e.stopPropagation();
    setIsDragging(true);
  };

  const handleDragLeave = (e: React.DragEvent<HTMLDivElement>) => {
    e.preventDefault();
    e.stopPropagation();
    setIsDragging(false);
  };

  const handleDrop = (e: React.DragEvent<HTMLDivElement>) => {
    e.preventDefault();
    e.stopPropagation();
    setIsDragging(false);

    if (isProcessing) return;

    const files = e.dataTransfer.files;
    if (files && files[0]) {
      const droppedFile = files[0];
      
      // Validate PDF file
      if (droppedFile.type === 'application/pdf' || droppedFile.name.endsWith('.pdf')) {
        setFile(droppedFile);
        setError(null);
      } else {
        setError('Please upload a PDF file');
      }
    }
  };

  const handleUpload = async () => {
    if (!file) {
      setError('Please select a PDF file');
      return;
    }

    setIsProcessing(true);
    setError(null);
    setProgress(0);

    try {
      // Step 1: Upload
      setCurrentStep('Uploading protocol...');
      setProgress(20);
      const uploadResult = await uploadProtocol(file, instructions);
      const sessionId = uploadResult.session_id;

      // Step 2: Parse
      setCurrentStep('Parsing PDF and converting to USDM...');
      setProgress(40);
      const parseResult = await parseProtocol(sessionId);

      // Step 3: Query similar trials
      setCurrentStep('Finding similar trials...');
      setProgress(60);
      const queryResult = await queryTrials(sessionId);

      // Step 4: Annotate trials and trigger automatic report generation
      setCurrentStep('Annotating trials (this may take a few minutes)...');
      setProgress(80);
      
      // Start annotation but don't wait for completion
      // Backend will automatically trigger report generation when annotation finishes
      annotateTrials(sessionId).catch((err) => {
        // Annotation continuing in background
      });
      
      // Poll for status instead of waiting
      setCurrentStep('Analysis running in background...');
      setProgress(85);
      
      // Redirect immediately to analysis page where it will poll for status
      setTimeout(() => {
        router.push(`/protocol/${sessionId}/analysis`);
      }, 1000);

    } catch (err) {
      if (err instanceof APIError) {
        setError(err.message);
      } else {
        setError('An unexpected error occurred. Please try again.');
      }
      
      setIsProcessing(false);
    }
  };

  // Get step emoji and title
  const getStepInfo = () => {
    if (progress >= 80) return { emoji: '⚡', title: 'Converting Trials', desc: 'Converting similar trials to USDM format' };
    if (progress >= 60) return { emoji: '🔍', title: 'Finding Similar Trials', desc: 'Searching ClinicalTrials.gov database' };
    if (progress >= 40) return { emoji: '📄', title: 'Parsing Protocol', desc: 'Extracting and structuring protocol data' };
    if (progress >= 20) return { emoji: '📤', title: 'Uploading', desc: 'Uploading protocol to secure server' };
    return { emoji: '🚀', title: 'Initializing', desc: 'Preparing analysis' };
  };

  const stepInfo = getStepInfo();

  return (
    <div className={`min-h-screen ${isProcessing ? 'bg-gradient-to-br from-background via-background to-primary/5' : 'bg-background'}`}>
      <div className="container mx-auto px-4 py-12">
        <div className="max-w-2xl mx-auto space-y-8">
          {!isProcessing && (
            <>
              {/* Navigation */}
              <div className="flex justify-between items-center">
                <a href="/" className="text-sm text-muted-foreground hover:text-primary">
                  ← Back to Trial Search
                </a>
                <a href="/sessions" className="text-sm text-muted-foreground hover:text-primary">
                  View Sessions →
                </a>
              </div>

              {/* Header */}
              <div className="text-center space-y-2">
                <h1 className="text-4xl font-bold">Protocol Intelligence</h1>
                <p className="text-muted-foreground">
                  Upload a Phase II/III clinical trial protocol for AI-powered analysis
                </p>
              </div>
            </>
          )}

          {isProcessing && (
            /* Beautiful Processing View */
            <div className="text-center space-y-6">
              <div className="text-6xl mb-4 animate-bounce">{stepInfo.emoji}</div>
              <h1 className="text-4xl font-bold bg-clip-text text-transparent bg-gradient-to-r from-primary to-primary/60">
                {stepInfo.title}
              </h1>
              <p className="text-lg text-muted-foreground">
                {stepInfo.desc}
              </p>
            </div>
          )}

          {/* Upload Form - Only show when not processing */}
          {!isProcessing && (
           
[truncated — 9097 more characters]
```

### front-end/app/protocol/[sessionId]/analysis/page.tsx

```typescript
'use client';

import { useEffect, useState } from 'react';
import { useParams, useRouter } from 'next/navigation';
import { getSession, getAnalysisReport, exportSession, APIError } from '@/lib/api';
import { BurdenAnalysis } from '@/components/analysis/BurdenAnalysis';
import { DurationDistribution } from '@/components/analysis/DurationDistribution';
import { SimilarTrialsPanel } from '@/components/analysis/SimilarTrialsPanel';
import { ProtocolOptimizer } from '@/components/analysis/ProtocolOptimizer';
import { LoadingSpinner } from '@/components/LoadingSpinner';
import { InfoTooltip } from '@/components/InfoTooltip';
import { DeveloperPanel } from '@/components/DeveloperPanel';
import { Alert, AlertDescription, AlertTitle } from '@/components/ui/alert';
import { Badge } from '@/components/ui/badge';
import { Card } from '@/components/ui/card';
import { 
  AlertCircle, 
  Download, 
  FileText, 
  TrendingUp, 
  Users, 
  Calendar,
  Target,
  Activity,
  Loader2
} from 'lucide-react';

export default function AnalysisPage() {
  const params = useParams();
  const router = useRouter();
  const sessionId = params.sessionId as string;

  const [session, setSession] = useState<any>(null);
  const [report, setReport] = useState<any>(null);
  const [isLoading, setIsLoading] = useState(true);
  const [error, setError] = useState<string | null>(null);
  const [exportLoading, setExportLoading] = useState(false);
  const [isPolling, setIsPolling] = useState(false);

  useEffect(() => {
    loadAnalysis();
  }, [sessionId]);

  // WebSocket for real-time updates (with polling fallback)
  useEffect(() => {
    if (!session || ['complete', 'error'].includes(session.status)) {
      return; // Don't connect if complete or error
    }

    const API_BASE_URL = process.env.NEXT_PUBLIC_API_URL || 'http://localhost:8000';
    const wsUrl = API_BASE_URL.replace('http', 'ws') + `/api/ws/${sessionId}`;

    console.log('🔌 Connecting to WebSocket:', wsUrl);
    const ws = new WebSocket(wsUrl);

    ws.onopen = () => {
      console.log('✅ WebSocket connected - will receive instant updates');
    };

    ws.onmessage = (event) => {
      const update = JSON.parse(event.data);
      console.log('📡 WebSocket update:', update);

      setSession((prev: any) => ({
        ...prev,
        status: update.status,
        metadata: { ...(prev?.metadata || {}), ...update.metadata },
        error_message: update.error_message
      }));

      // Handle completion
      if (update.status === 'complete') {
        console.log('✅ Analysis complete (WebSocket), loading report...');
        getAnalysisReport(sessionId)
          .then(reportData => {
            setReport(reportData);
            console.log('✅ Report loaded successfully!');
          })
          .catch(err => {
            console.error('Report loading error:', err);
            setError('Failed to load report.');
          });
      } else if (update.status === 'error') {
        setError(update.error_message || 'Analysis failed');
      }
    };

    ws.onerror = (error) => {
      console.warn('⚠️ WebSocket error, will fallback to polling:', error);
    };

    ws.onclose = () => {
      console.log('🔌 WebSocket disconnected');
    };

    return () => {
      console.log('🔌 Closing WebSocket connection');
      ws.close();
    };
  }, [session?.status, sessionId]);

  useEffect(() => {
    // Poll for status updates if analysis is in progress (FALLBACK for WebSocket)
    if (session && session.status && !['complete', 'error'].includes(session.status)) {
      setIsPolling(true);
      const pollInterval = setInterval(async () => {
        try {
          const sessionData = await getSession(sessionId);
          console.log('📊 Poll update:', {
            status: sessionData.status,
            progress: sessionData.metadata?.progress,
            step: sessionData.metadata?.step,
            trials_completed: sessionData.metadata?.trials_completed,
            trials_total: sessionData.metadata?.trials_total
          });
          setSession(sessionData);
          
          // Backend automatically generates report after annotation
          // Frontend just polls until status is 'complete'
          if (sessionData.status === 'complete') {
            // If complete, load report and stop polling
            console.log('✅ Analysis complete, loading report...');
            try {
              const reportData = await getAnalysisReport(sessionId);
              setReport(reportData);
              setIsPolling(false);
              clearInterval(pollInterval);
              console.log('✅ Report loaded successfully!');
            } catch (reportErr) {
              console.error('Report loading error:', reportErr);
              setError('Failed to load report. The analysis completed but report retrieval failed.');
              setIsPolling(false);
              clearInterval(pollInterval);
            }
          } else if (sessionData.status === 'error') {
            setError(sessionData.error_message || 'Analysis failed');
            setIsPolling(false);
            clearInterval(pollInterval);
          }
          // For all other statuses (created, parsing, querying, annotating, analyzing):
          // Just keep polling and showing progress
        } catch (err) {
          console.error('Polling error:', err);
        }
      }, 5000); // Poll every 5 seconds

      return () => clearInterval(pollInterval);
    }
  }, [session?.status, sessionId]);

  const loadAnalysis = async () => {
    setIsLoading(true);
    setError(null);

    try {
      // Get session details with retry logic
      let sessionData;
      let retries = 0;
      const maxRetries = 3;
      
      while (retries < maxRetries) {
        try {
          sessionData = await getSession(sessionId);
          break;
        } catch (err) {
          retries++;
          if (retries === maxRetries) throw err;
          console.log(`⚠️ Session load attempt ${re
[truncated — 26501 more characters]
```

[111 more indexed source files omitted to keep this export small. The full file list is in the Codebase structure section above.]