Project Info
Author: Vocal Coach AI Team Repository: Berkeley Hack 2025 Overview This project is a comprehensive AI-powered vocal coaching system built for the Berkeley Hackathon 2025. The application provides a full-stack solution for personalized vocal training, combining real-time voice analysis, intelligent coaching, and conversational AI agents to create a holistic vocal development experience. VocalAIAgent solves the problem of fragmented vocal training by combining multiple functionalities into one intuitive and multimodal assistant. This AI-powered system simplifies vocal coaching by: Understanding natural voice patterns and interpreting audio recordings to gather vocal characteristics Providing personalized recommendations based on vocal analysis and user preferences Fetching real-time data for vocal metrics, progress tracking, and performance insights Generating dynamic lesson plans tailored to the user's vocal type, skill level, and practice goals Offering intelligent coaching conversations grounded in real-time voice analysis data VocalAIAgent is not just an LLM chatbot. The agent is enhanced with numerous AI capabilities, including: Voice Understanding: Real-time pitch detection, vocal analysis with metrics like jitter, shimmer, vibrato rate Retrieval-Augmented Generation (RAG): Providing personalized coaching tips by retrieving relevant vocal techniques from a knowledge base Few-Shot Prompting: Generating dynamic lesson plans and exercises based on minimal user input Function Calling: Executing specific functions based on user commands, such as starting voice sessions, analyzing recordings, or generating progress reports Long Context Window: Managing and retaining user vocal profiles and practice history across multiple sessions Context Caching: Storing relevant vocal data temporarily to improve response speed and reduce redundant analysis AI Evaluation: Using LLM-based evaluation to assess vocal progress and provide "Vocal Scores" based on improvement and consistency Grounding: Ensuring that coaching recommendations are grounded in real-time vocal analysis data Embeddings: Utilizing embeddings for effective vocal pattern matching and personalized exercise recommendations Multimodal Integration: Understanding both voice inputs and conversational text for comprehensive coaching Problem Statement Vocal training can be an isolated and inconsistent process. Singers and speakers often struggle with: Lack of real-time feedback during practice sessions Limited access to personalized coaching based on their specific vocal characteristics Difficulty tracking progress and identifying improvement areas Fragmented resources across multiple platforms and tools Inconsistent practice routines without proper guidance VocalAIAgent addresses these challenges by providing a unified, intelligent coaching platform that combines voice analysis, personalized AI coaching, and comprehensive progress tracking in one seamless experience. π Key Features Core Vocal Analysis π΅ Real-Time Pitch Detection: Instant feedback during practice sessions with live pitch visualization π Deep Vocal Analysis: Advanced metrics including jitter, shimmer, vibrato rate, vocal range analysis π― Voice Type Classification: Automatic classification of voice types (soprano, alto, tenor, bass) π Progress Tracking: Comprehensive tracking of vocal improvements over time AI-Powered Coaching System π€ Dual-AI Architecture: Proactive Fetch.ai Agent + Reactive Letta Conversational Agent π¬ Stateful Conversations: AI coach that remembers context and discusses specific progress π Personalized Lesson Plans: Dynamic lesson generation based on vocal analysis and user goals π Exercise Recommendations: Tailored vocal exercises based on analysis results Advanced Features π£οΈ VAPI Voice Integration: Real-time voice conversations with AI coach π± Multimodal Interface: Support for voice input, text chat, and visual feedback π Lesson Feedback Loop: Comprehensive storage and analysis of lesson completion data π AI-Generated Reports: Daily summaries of performance trends and insights πͺ Community Features: Progress sharing and vocal challenges Data & Memory Management πΎ Persistent Memory: User preferences, vocal characteristics, and practice history retention π€ Export Capabilities: Save vocal analyses, lesson plans, and progress reports π Secure Data Storage: Supabase integration with proper authentication and RLS π Session Management: Comprehensive tracking of practice sessions and improvements Why This Matters Vocal training today lacks the personalized, data-driven approach that modern AI can provide. VocalAIAgent brings together voice science, conversational AI, and personalized coaching into one intelligent system, offering a more effective, engaging, and accessible vocal training experience. By combining real-time voice analysis, stateful AI conversations, and comprehensive progress tracking, this tool showcases the potential of Generative AI in revolutionizing music education and vocal development. π Tech Stack Frontend React: Modern component-based UI framework TypeScript: Type-safe development with enhanced IDE support Vite: Fast build tool and development server Tailwind CSS: Utility-first CSS framework for responsive design Framer Motion: Smooth animations and transitions Backend FastAPI: High-performance Python web framework with automatic API docs Python: Core backend language with extensive AI/ML libraries Uvicorn: ASGI server for production deployment AI Services Letta: Stateful conversational AI with long-term memory capabilities Fetch.ai: Autonomous agent system for proactive analysis and reporting VAPI: Voice AI platform for real-time voice conversations Database & Authentication Supabase: PostgreSQL database with built-in authentication and real-time features Row Level Security (RLS): Secure user data isolation Voice Processing Web Audio API: Real-time audio processing and pitch detection Custom Voice Analyzer: Advanced vocal metrics calculation Hosting & Deployment Netlify: Frontend hosting with automatic deployments Google Cloud Run: Scalable backend container hosting Docker: Containerized backend for consistent deployments How It Works Intent Recognition and Routing System VocalAIAgent uses a sophisticated routing system that directs user requests to appropriate handlers based on vocal coaching context: Vocal Analysis Pipeline The voice analysis system combines multiple AI techniques: Real-time Processing: Web Audio API captures and processes audio in real-time Feature Extraction: Advanced algorithms extract vocal characteristics (pitch, formants, etc.) AI Classification: Machine learning models classify voice type and detect patterns Contextual Analysis: Results are interpreted within the user's vocal development context Dual-AI Architecture Proactive Fetch.ai Agent Reactive Letta Conversational Agent Session Management and Memory VocalAIAgent maintains comprehensive session state and user memory: Current Capabilities Demonstrated β Voice Analysis & Processing Real-time pitch detection with Web Audio API Advanced vocal metrics (jitter, shimmer, vibrato) Voice type classification and range analysis Session recording and playback capabilities β AI-Powered Coaching Fetch.ai autonomous agents for progress analysis Letta conversational AI with stateful memory VAPI real-time voice conversations Personalized lesson and exercise generation β Data Management & Persistence Comprehensive user vocal profiles Session history and progress tracking Lesson feedback storage and retrieval Secure multi-user data isolation β User Experience Features Modern, responsive React interface Real-time visual feedback during voice sessions Progress dashboards and analytics Community features and challenges Current Errors and Solutions Issues Identified: URL Construction Error: Double slash in API endpoints causing malformed URLs Database Connection Issues: Lesson feedback storage failing due to Supabase credential problems Error Handling: Generic error messages making debugging difficult Solutions Implemented: Fixed URL Construction: Added trailing slash removal in frontend API calls Enhanced Error Logging: Improved backend error reporting with detailed messages Database Health Checks: Added endpoints to verify service connectivity Limitations & Future Work Current Limitations: Voice Processing Accuracy: Browser-based analysis has limitations compared to specialized hardware AI Model Training: Limited training data for vocal coaching specific AI models Scalability: Current architecture needs optimization for large-scale deployment Future Enhancements: High Priority: Enhanced Voice Processing: Integrate professional-grade voice analysis libraries Advanced AI Models: Fine-tune models specifically for vocal coaching contexts Mobile Applications: Native iOS/Android apps with enhanced voice processing Medium Priority: Social Features: Enhanced community aspects with vocal challenges and peer learning Integration Ecosystem: Connect with music learning platforms and DAWs Offline Capabilities: Voice analysis and basic coaching without internet connection Built With Core Technologies React 18 with TypeScript for modern, type-safe frontend development FastAPI for high-performance Python backend with automatic API documentation Supabase for PostgreSQL database, authentication, and real-time features Tailwind CSS for responsive, utility-first styling AI & Voice Technologies Letta for stateful conversational AI with long-term memory Fetch.ai for autonomous agent systems and proactive analysis VAPI for real-time voice AI conversations Web Audio API for browser-based voice processing DevOps & Deployment Docker for containerized backend deployment Google Cloud Run for scalable, serverless backend hosting Netlify for frontend hosting with automatic deployments VocalAIAgent demonstrates the transformative potential of AI in music education, combining cutting-edge voice processing, conversational AI, and personalized coaching to create a comprehensive vocal training platform.
VocalAIAgent - AI-Powered Vocal Coaching System π€
Winner of "Most Ambitious Vapi Project" at UC Berkeley Hackathon 2025
VocalAIAgent is a comprehensive AI-powered vocal coaching system that combines real-time voice analysis, intelligent coaching, and conversational AI agents to create a holistic vocal development experience. Built during the world's biggest AI in-person hackathon where over 1,200 developers competed.
YouTube Demo:
π― Problem Statement
Vocal training can be an isolated and inconsistent process. Singers and speakers often struggle with:
- Lack of real-time feedback during practice sessions
- Limited access to personalized coaching based on their specific vocal characteristics
- Difficulty tracking progress and identifying improvement areas
- Fragmented resources across multiple platforms and tools
- Inconsistent practice routines without proper guidance
Solution
VocalAIAgent addresses these challenges by providing a unified, intelligent coaching platform that combines voice analysis, personalized AI coaching, and comprehensive progress tracking in one seamless experience.
π Key Features
Core Vocal Analysis
- π΅ Real-Time Pitch Detection: Instant feedback during practice sessions with live pitch visualization
- π Deep Vocal Analysis: Advanced metrics including jitter, shimmer, vibrato rate, vocal range analysis
- π― Voice Type Classification: Automatic classification of voice types (soprano, alto, tenor, bass)
- π Progress Tracking: Comprehensive tracking of vocal improvements over time
AI-Powered Coaching System
- π€ Dual-AI Architecture: Proactive Fetch.ai Agent + Reactive Letta Conversational Agent
- π¬ Stateful Conversations: AI coach that remembers context and discusses specific progress
- π Personalized Lesson Plans: Dynamic lesson generation based on vocal analysis and user goals
- π Exercise Recommendations: Tailored vocal exercises based on analysis results
Advanced Features
- π£οΈ VAPI Voice Integration: Real-time voice conversations with AI coach
- π± Multimodal Interface: Support for voice input, text chat, and visual feedback
- π Lesson Feedback Loop: Comprehensive storage and analysis of lesson completion data
- π AI-Generated Reports: Daily summaries of performance trends and insights
- οΏ½οΏ½ Community Features: Progress sharing and vocal challenges
Data & Memory Management
- πΎ Persistent Memory: User preferences, vocal characteristics, and practice history retention
- π€ Export Capabilities: Save vocal analyses, lesson plans, and progress reports
- π Secure Data Storage: Supabase integration with proper authentication and RLS
- οΏ½οΏ½ Session Management: Comprehensive tracking of practice sessions and improvements
π Tech Stack
Frontend
- React: Modern component-based UI framework
- TypeScript: Type-safe development with enhanced IDE support
- Vite: Fast build tool and development server
- Tailwind CSS: Utility-first CSS framework for responsive design
- Framer Motion: Smooth animations and transitions
Backend
- FastAPI: High-performance Python web framework with automatic API docs
- Python: Core backend language with extensive AI/ML libraries
- Uvicorn: ASGI server for production deployment
AI Services
- Letta: Stateful conversational AI with long-term memory capabilities
- Fetch.ai: Autonomous agent system for proactive analysis and reporting
- VAPI: Voice AI platform for real-time voice conversations
Database & Authentication
- Supabase: PostgreSQL database with built-in authentication and real-time features
- Row Level Security (RLS): Secure user data isolation
Voice Processing
- Web Audio API: Real-time audio processing and pitch detection
- Custom Voice Analyzer: Advanced vocal metrics calculation
Hosting & Deployment
- Netlify: Frontend hosting with automatic deployments
- Google Cloud Run: Scalable backend container hosting
- Docker: Containerized backend for consistent deployments
π Getting Started
Prerequisites
- Node.js (v18 or higher)
- Python (v3.8 or higher)
- Docker (for backend deployment)
Frontend Setup
cd src
npm install
npm run dev
Backend Setup
cd backend
pip install -r requirements.txt
uvicorn main:app --reload
Environment Variables
Create .env files in both frontend and backend directories with the necessary API keys and configuration.
Why This Matters
Vocal training today lacks the personalized, data-driven approach that modern AI can provide. VocalAIAgent brings together voice science, conversational AI, and personalized coaching into one intelligent system, offering a more effective, engaging, and accessible vocal training experience.
By combining real-time voice analysis, stateful AI conversations, and comprehensive progress tracking, this tool showcases the potential of Generative AI in revolutionizing music education and vocal development.
Live Demo
https://prismatic-buttercream-5f0d5a.netlify.app/
Built with β€οΈ at UC Berkeley Hackathon 2025
Analysis
View
Metric
- 72
- 35
- 5
- 2
Figures cover GitHub contributors during the hackathon window. A co-authored commit counts in full for each author, so per-member totals add up to more than the whole-team figures.
Technology
- CSSIn code
- FastAPIIn code
- HTMLIn code
- JavaScriptIn code
- PythonIn code
- ReactIn code
- SupabaseIn code
- Tailwind CSSIn code
- TypeScriptIn code
- DockerClaimed
9 of 10 appear in the indexed code. 1 claimed on Devpost could not be matched to code, which may simply mean the tool leaves no trace in the repository.
AI coding agents
No AI coding agent signals were found in this repository.
Detected from committed agent config files and commit authorship. Absence of a signal is not proof an agent was unused.
Codebase size
Source size
613 KB
Source files
65
Counts recognized source files only; vendored directories, binaries and lockfiles are excluded, so this is smaller than the repository on disk.
Repository
ramizik/vocal-ai
81 files Β· 799 KB Β· @ aea5097
Structure
Interface
30 files Β· 37%Screens, components and styles rendered to the user.
Application logic
32 files Β· 40%Domain rules, services and shared utilities.
+2 more
Supporting
Layers are inferred from where files sit in the tree, not from reading the code. A project that names its directories unconventionally will read oddly here β open the file browser to check anything the diagram implies.
Languages
- TypeScript55%
- Python41%
- Markdown3%
- CSS1%
- JavaScript0%
- YAML0%
- Other (2)0%
Share of indexed source by file size. Binary and vendored files are excluded.
Dependencies
backend/requirements.txt
pypi Β· 22- fastapi
- groq
- gunicorn
- httpx
- letta-client
- librosa
- matplotlib
- numpy
- pandas
- pydantic
- pydub
- python-dateutil
- python-dotenv
- python-json-logger
- python-multipart
- pytz
- requests
- scikit-learn
- +4 more
package.json
npm Β· 20- @supabase/supabase-js
- @vapi-ai/web
- @vitejs/plugin-react
- autoprefixer
- framer-motion
- lucide-react
- pitchy
- postcss
- react
- react-dom
- react-particles
- react-router-dom
- tailwindcss
- tsparticles
- tsparticles-engine
- tsparticles-slim
- typescript
- vite
- +2 more
Declared in the repositoryβs manifests at the indexed commit. A declared package is not proof it is used, and runtime dependencies are listed first.
This projectβs features have not been analysed yet.
Export this project's context (description, README, evidence, key source files) to chat with an AI agent elsewhere.
