# Project export: ConvoCoach

This document was generated by HackStack to give an AI agent context about a hackathon project. Sections are labeled with their provenance; content marked as truncated was cut to keep this document small.

## Project metadata

- Hackathon: TreeHacks 2025
- Tagline: Convo Coach is an AI-powered speech therapist that helps individuals with speaking disabilities improve fluency, articulation, and confidence through real-time feedback and personalized coaching.
- Devpost: https://devpost.com/software/convocoach-ipvah3
- GitHub: https://github.com/lag-gam/ConvoCoach
- Video: https://www.youtube.com/embed/bdrICaryIzs?enablejsapi=1&hl=en_US&rel=0&start=&version=3&wmode=transparent
- Team: 2 GitHub contributor(s) — Agrantlv (3 commits), lag-gam (2 commits)

## Devpost submission (written by the team)

### Inspiration

For over 70 million people globally, speech disorders turn communication—a fundamental human connection—into a source of anxiety. Stuttering, articulation challenges, and social anxiety can erode confidence, limit opportunities, and create isolation. While traditional speech therapy plays a crucial role, it remains out of reach for many due to cost, geographic barriers, and social stigma, leaving a gap that technology has the potential to bridge. The idea for Convo Coach emerged from a personal experience. One of our team members was once criticized in school for overusing filler words like “um”—a small comment that carried lasting emotional weight. This moment revealed a universal truth: even minor speech differences can fuel self-doubt. Many people hold back from speaking, not because they lack the words, but because they fear judgment. Over time, this fear discourages practice, leading to a cycle of avoidance that reinforces insecurities. Speech challenges are more than technical obstacles—they are deeply human struggles. A person might remain silent in a meeting to avoid stuttering or rehearse conversations endlessly due to anxiety. Convo Coach was created to break this cycle. By providing real-time analysis of speech patterns and filler words, it empowers users to practice and improve in a private, judgment-free space. More than just a tool, it’s a step toward making speech a source of confidence, not anxiety, ensuring that everyone has the opportunity to be heard.

### What it does

Convo Coach helps users reduce filler words in their speech by providing real-time AI-powered feedback. When a user records themselves speaking, the app accurately transcribes their speech, highlighting unnecessary fillers like “um,” “like,” and “uh.” After analyzing the speech, an AI avatar responds dynamically, providing personalized feedback on areas for improvement, such as suggesting pauses before complex phrases or reducing filler words for clearer communication. By offering instant, interactive feedback in a judgment-free environment, Convo Coach helps users refine their speech and build confidence in their communication skills.

### How we built it

ConvoCoach is an AI-powered conversational trainer that helps users improve their speaking skills in real time. By leveraging advanced AI technologies, it listens to speech, transcribes it, analyzes patterns, and provides personalized coaching through a talking AI avatar. The system enhances fluency by identifying issues like filler words, pauses, stuttering, and pacing irregularities, offering actionable feedback in a natural and interactive way. To process speech, ConvoCoach uses AssemblyAI for transcription and OpenAI API for speech analysis, detecting patterns such as overuse of filler words, extended pauses, and inconsistent pacing. Based on this, it generates concise, constructive feedback with practical suggestions like “Try slowing down a bit” or “Reduce filler words like ‘umm’ and ‘like.’” The feedback is then converted into speech by D-ID, which syncs it with a talking avatar to create a more engaging and immersive coaching experience. By integrating AssemblyAI for transcription, GPT-4 for analysis, D-ID for avatar interaction, and Flask for backend processing, ConvoCoach delivers a seamless, real-time speech training tool. This AI-powered system makes conversational coaching more accessible and effective, transforming speech practice into an interactive, judgment-free experience.

### Challenges we ran into

Integrating a Flask API for speech processing while maintaining low latency was a major challenge. Speech transcription and real-time sentiment analysis required optimizing backend processing to ensure responses felt instant and natural. Balancing accuracy with speed proved difficult, as deeper speech analysis introduced delays that could disrupt user experience. Developing a real-time AI avatar that delivers personalized feedback added another layer of complexity. Generating dynamic, context-aware responses required fine-tuning sentiment analysis models and ensuring that the AI’s tone and pacing adapted to user input. Creating a seamless connection between the avatar and speech analysis while avoiding response lag was a significant hurdle. Synchronizing the full-stack architecture between Flask and React required careful API optimization. Managing user speech data efficiently, ensuring smooth frontend-to-backend communication, and refining feedback delivery were key technical challenges. Through rigorous testing and iteration, we built a responsive, interactive system that delivers instant, AI-driven speech feedback to help users refine their communication skills.

### Accomplishments we're proud of

As newcomers to the hackathon scene, we are incredibly proud of building a fully functional project within such a short timeframe. Stepping outside our comfort zones, we embraced new technologies and tackled the challenges of creating a full-stack AI-powered speech coach from scratch. Beyond just technical achievements, we’re most proud that Convo Coach has real-world impact, providing an accessible tool that can help people improve their communication skills. We also think our Kenna avatar is pretty cool.

### What we learned

Through the development of ConvoCoach, we gained valuable experience in integrating multiple AI technologies into a seamless, real-time system. We deepened our understanding of Machine Learning for speech analysis, learning how to process live audio input, extract meaningful insights, and generate feedback in a way that feels natural and intuitive. Working with AssemblyAI, GPT-4, and D-ID, we navigated the challenges of interfacing different AI models, ensuring smooth synchronization between transcription, analysis, and avatar-based feedback to create an engaging user experience. Building ConvoCoach also pushed us to refine our skills in real-time data processing, minimizing latency while maintaining accuracy. We optimized API calls, synchronized AI outputs, and streamlined backend operations to ensure smooth performance. Beyond the technical aspects, this project reinforced the social and educational value of AI-powered speech coaching, highlighting its potential to help students, professionals, and individuals with speech difficulties improve their communication skills. This experience has shown us how AI can be leveraged not just for automation, but for meaningful, human-centered applications that empower users to communicate more effectively.

### What's next

The development of ConvoCoach has opened up exciting possibilities for expanding AI-driven speech training. Moving forward, we aim to enhance the system by introducing gamification features, allowing users to earn experience points, level up, and complete speech challenges that encourage continuous improvement. By incorporating progress tracking, users will be able to monitor their speech development over time, fostering a sense of motivation and achievement as they refine their communication skills. Beyond individual speech coaching, we see potential for ConvoCoach to be applied in education and classroom engagement. A possible extension of the platform would be an AI-powered substitute teacher that allows educators to upload lesson plans, which are then delivered interactively through AI avatars. Students would join as avatars, engage in discussions through voice or chat, and interact with AI-driven prompts in real-time. By integrating with platforms like Pear Deck, this system could facilitate structured class discussions and quizzes, making learning more engaging and accessible. The broader vision for ConvoCoach is to expand its impact on accessibility and human-centric AI solutions. The platform has the potential to assist individuals with speech difficulties, improve public speaking skills, and provide an inclusive learning environment for those who may struggle with verbal communication. As AI-driven speech coaching continues to evolve, we are committed to refining ConvoCoach into a scalable, adaptive, and widely accessible tool that empowers users to communicate with confidence in any setting.

## README (from the GitHub repository)

# ConvoCoach
AI Conversation Coach

AI-powered conversational training tool designed to help users improve their speaking skills in real time. It listens to conversations via a microphone, transcribes speech, and analyzes key aspects such as clarity, pacing, filler words, confidence, and articulation. Using advanced AI, ConvoCoach provides instant feedback and personalized coaching to help users refine their communication skills.
To make the experience immersive and engaging, ConvoCoach features a talking AI avatar that provides real-time coaching, offering verbal and visual feedback as if the user were speaking to a real conversation partner. Additionally, the tool includes a gamified progression system, where users can earn XP, level up, and unlock challenges based on their conversational improvements.
By leveraging technologies like Whisper AI for speech-to-text, GPT-4 for conversation analysis, TTS (Text-to-Speech) for voice responses, and D-ID for a talking avatar, ConvoCoach creates an interactive and engaging learning experience. Whether for public speaking, professional interviews, or everyday conversations, ConvoCoach helps users build confidence and refine their speech patterns effectively.
💡 Perfect for students, professionals, and anyone looking to improve their conversational fluency in a fun and structured way! 🚀


## Detected evidence (automated analysis)

Indexed codebase: 18 recognized source files, 29 KB.
- CSS (language) — detected in the code
- HTML (language) — detected in the code
- JavaScript (language) — detected in the code
- Python (language) — detected in the code
- React (technology) — detected in the code
- Tailwind CSS (technology) — detected in the code
- TypeScript (language) — detected in the code
- Flask (technology) — claimed on Devpost, not found in the code

## Codebase structure (from repository index)

### Files (28 of 28)

```
.DS_Store
.gitignore
app.py
avatar.py
frontend/.gitignore
frontend/eslint.config.js
frontend/index.html
frontend/package-lock 2.json
frontend/package.json
frontend/postcss.config.js
frontend/README.md
frontend/src/App.css
frontend/src/App.tsx
frontend/src/index.css
frontend/src/main.tsx
frontend/src/vite-env.d.ts
frontend/tailwind.config.js
frontend/tsconfig.app.json
frontend/tsconfig.json
frontend/tsconfig.node.json
frontend/vite.config.ts
index.html
package.json
README.md
tes.m4a
test_speech_analysis.py
user_speech_to_text.py
vite.config.js
```

### Dependencies

- frontend/package.json: @eslint/js@^9.19.0, @tailwindcss/postcss@^4.0.6, @types/react@^19.0.8, @types/react-dom@^19.0.3, @vitejs/plugin-react@^4.3.4, autoprefixer@^10.4.20, eslint@^9.19.0, eslint-plugin-react-hooks@^5.0.0, eslint-plugin-react-refresh@^0.4.18, globals@^15.14.0, postcss@^8.5.2, react@^19.0.0, react-dom@^19.0.0, tailwindcss@^4.0.6, typescript@~5.7.2, typescript-eslint@^8.22.0, vite@^6.1.0
- package.json: @headlessui/react@^1.7.7, @heroicons/react@^2.0.18, @tailwindcss/vite@^4.0.6, @vitejs/plugin-react@^4.3.4, autoprefixer@^10.4.20, lucide-react@^0.475.0, postcss@^8.5.2, react@^18.2.0, react-dom@^18.2.0, recoil@^0.7.6, shadcn-ui@^0.9.4, socket.io-client@^4.8.1, tailwindcss@^3.4.17, vite@^6.1.0

### Recent commits (newest first)

- added icon to AI agent
- enhanced UI
- added frontend more
- Made basic frontend with user-ai interaction
- implement avatar to get ai text response
- Added speech to chat response pipeline
- Resolved merge conflict in .gitignore
- fixed merge conflict
- resolved
- added gitignore for assemblyai
- Adds the main function to convert the transcript to a response (only prints)
- updated speech to text for pipeline
- Merge branch 'main' of https://github.com/lag-gam/ConvoCoach
- User Speech and sentiment and confidence analysis
- Updated commit with simple UI
- Your commit message here
- Update README.md
- Initial commit

## Key source files (fetched from GitHub, selected and truncated for size)

### package.json

```
{
  "name": "convo-coach-ui",
  "version": "0.0.1",
  "private": true,
  "scripts": {
    "dev": "vite",
    "build": "vite build",
    "preview": "vite preview"
  },
  "dependencies": {
    "@headlessui/react": "^1.7.7",
    "@heroicons/react": "^2.0.18",
    "lucide-react": "^0.475.0",
    "react": "^18.2.0",
    "react-dom": "^18.2.0",
    "recoil": "^0.7.6",
    "socket.io-client": "^4.8.1"
  },
  "devDependencies": {
    "@tailwindcss/vite": "^4.0.6",
    "@vitejs/plugin-react": "^4.3.4",
    "autoprefixer": "^10.4.20",
    "postcss": "^8.5.2",
    "shadcn-ui": "^0.9.4",
    "tailwindcss": "^3.4.17",
    "vite": "^6.1.0"
  }
}

```

### frontend/package.json

```
{
  "name": "convocoach",
  "private": true,
  "version": "0.0.0",
  "type": "module",
  "scripts": {
    "dev": "vite",
    "build": "tsc -b && vite build",
    "lint": "eslint .",
    "preview": "vite preview"
  },
  "dependencies": {
    "react": "^19.0.0",
    "react-dom": "^19.0.0"
  },
  "devDependencies": {
    "@eslint/js": "^9.19.0",
    "@tailwindcss/postcss": "^4.0.6",
    "@types/react": "^19.0.8",
    "@types/react-dom": "^19.0.3",
    "@vitejs/plugin-react": "^4.3.4",
    "autoprefixer": "^10.4.20",
    "eslint": "^9.19.0",
    "eslint-plugin-react-hooks": "^5.0.0",
    "eslint-plugin-react-refresh": "^0.4.18",
    "globals": "^15.14.0",
    "postcss": "^8.5.2",
    "tailwindcss": "^4.0.6",
    "typescript": "~5.7.2",
    "typescript-eslint": "^8.22.0",
    "vite": "^6.1.0"
  }
}

```

### app.py

```python
import os
import time
import requests
from flask import Flask, request, jsonify
from flask_cors import CORS
from dotenv import load_dotenv

import user_speech_to_text as STT #speech to text
import test_speech_analysis as chat_response

# Load environment variables
load_dotenv()

ASSEMBLYAI_API_KEY = os.getenv("ASSEMBLYAI_API_KEY")
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
AVATAR_API_KEY = os.getenv("AVATAR_API_KEY")

if not ASSEMBLYAI_API_KEY:
    raise ValueError("Missing AssemblyAI API key in .env!")
if not OPENAI_API_KEY:
    raise ValueError("Missing OpenAI API key in .env!")
if not AVATAR_API_KEY:
    raise ValueError("Missing D-ID API key in .env!")

app = Flask(__name__)
CORS(app, resources={r"/generate-video": {"origins": "http://localhost:5173"}})




@app.route("/generate-video", methods=["POST"])
def generate_video():
    """
    Endpoint that:
    1. Records + transcribes audio.
    2. Generates GPT-based response.
    3. Creates a D-ID talk (video) with a chosen avatar reading the AI response.
    4. Polls until the video is ready and returns the video URL.
    """
    try:
        if "audio" not in request.files:
            return jsonify({"error": "No audio file uploaded"}), 400

        audio_file = request.files["audio"]
        filename = f"user_audios/uploaded_{int(time.time())}.wav"
        os.makedirs("user_audios", exist_ok=True)  # Ensure directory exists
        audio_file.save(filename)
       
        transcription = STT.transcribe_audio(filename)
        transcript_text = transcription.get("text", "").strip()

        if not transcript_text:
            return jsonify({"error": "No valid transcription received."}), 400


        ai_response = chat_response.generate_avatar_response(transcript_text)
        print("Received AI response:", ai_response)

       
        # Choose your avatar's image URL
        source_url = 'https://media-hosting.imagekit.io//4078ec62a6824c90/screenshot_1739688772692.png?Expires=1834296773&Key-Pair-Id=K2ZIVPTIP2VGHC&Signature=LwWvBbL2shd0AqMt~XyVwUwxey2AGKgPPrMEAql3lULNTvG8mF~Urcxhpy7KvW2Cy2LL1VWU69UBsOXc5SWU9FgsvvVz03ggSbtVtApbitZ2eh7uKwiv0rzsvhTFCwmJBxPpiU9D-CdY~ynDhlHodH~xE8wpd2WYAzdChBhLW2QfeEPDs3zP12LXeZFnO0-05Xco0LTnwQOUK12ZDUsIZLtkWIln6Ooqkb9R-BjBDKT9zGnw5IY5D2JHWrTiJOLX3H0W0Ti65BcXYrbYiILndV2xVM-aIfVPoCtz2bYrSPswkG5yPYDErl79jApaQCvIc3mjZBDvCpEjC4u2U-7xGw__'

        # Create talk endpoint
        create_talk_url = "https://api.d-id.com/talks"
        
        # D-ID API uses Basic Auth with your D-ID API key as the "username" and no password
        headers = {
            "Authorization": f"Basic {AVATAR_API_KEY}",
            "Content-Type": "application/json"
        }

        # Payload for creating the talk
        data = {
            "source_url": source_url,
            "script": {
                # Optionally specify voice provider settings here:
                # "provider": {
                #     "type": "amazon",
                #     "voice_id": "Joey"
                # },
                "type": "text",
                "input": ai_response
            }
        }

        # Create the talk
        response = requests.post(create_talk_url, headers=headers, json=data)

        if response.status_code != 201:
            return jsonify({
                "error": f"Failed to create talk: {response.status_code} - {response.text}"
            }), 500

        talk_id = response.json().get("id")
        print(f"Talk created successfully with ID: {talk_id}")

        # ---------------------------
        # Step 4: Poll until video is ready
        # ---------------------------
        get_talk_url = f"https://api.d-id.com/talks/{talk_id}"

        result_url = None
        while True:
            poll_response = requests.get(get_talk_url, headers=headers)
            if poll_response.status_code == 200:
                talk_data = poll_response.json()
                status = talk_data.get("status")
                if status == "done":
                    result_url = talk_data.get("result_url")
                    print(f"Video generated successfully: {result_url}")
                    break
                elif status == "error":
                    return jsonify({"error": "Error in video generation."}), 500
                else:
                    print("Video is still processing...")
            else:
                return jsonify({"error": f"Failed to retrieve talk: {poll_response.status_code} - {poll_response.text}"}), 500

            time.sleep(5)  # Wait 5 seconds before polling again

        # Return the final video link
        return jsonify({
            "transcript": transcript_text,
            "ai_response": ai_response,
            "video_url": result_url
        })

    except Exception as e:
        return jsonify({"error": str(e)}), 500


if __name__ == "__main__":
    # Run the Flask server
    app.run(port=5000, debug=True)
```

### frontend/src/main.tsx

```typescript
import { StrictMode } from 'react'
import { createRoot } from 'react-dom/client'
import './index.css'
import App from './App.tsx'
import './index.css'; // Import Tailwind CSS

createRoot(document.getElementById('root')!).render(
  <StrictMode>
    <App />
  </StrictMode>,
)

```

### frontend/src/App.tsx

```typescript
"use client"

import { useState, useRef, useEffect } from "react"
import { Mic, MicOff, Video, VideoOff, Loader } from "lucide-react"

const avatars = [
  {
    name: "Default Avatar",
    url: "https://media.licdn.com/dms/image/v2/D5603AQFQemGbPTJW5w/profile-displayphoto-shrink_200_200/profile-displayphoto-shrink_200_200/0/1723056712090?e=2147483647&v=beta&t=a6e5a0TMt4R3_FQgyjjbz1eayWz3WvE9NWSRjnpChWI",
  },
  { name: "Avatar 2", url: "https://static.planetminecraft.com/files/resource_media/screenshot/1428/stevepmc7855107.jpg" },
  { name: "Avatar 3", url: "https://example.com/avatar3.png" },
]

const prompts = [
  "Introduce yourself and your background",
  "Describe your ideal work environment",
  "Explain a challenging situation you've overcome",
  "Discuss your long-term career goals",
  "Share a recent accomplishment you're proud of",
]

export default function ConvoCoach() {
  const [recording, setRecording] = useState(false)
  const [videoUrl, setVideoUrl] = useState("")
  const [aiResponse, setAiResponse] = useState("")
  const [loading, setLoading] = useState(false)
  const [selectedAvatar, setSelectedAvatar] = useState(avatars[0].url)
  const [selectedPrompt, setSelectedPrompt] = useState("")
  const [webcamActive, setWebcamActive] = useState(false)
  const mediaRecorder = useRef(null)
  const audioChunks = useRef([])
  const videoRef = useRef(null)

  useEffect(() => {
    if (webcamActive) {
      startWebcam()
    } else {
      stopWebcam()
    }
  }, [webcamActive])

  const startWebcam = async () => {
    try {
      const stream = await navigator.mediaDevices.getUserMedia({ video: true })
      if (videoRef.current) {
        videoRef.current.srcObject = stream
      }
    } catch (err) {
      console.error("Error accessing webcam:", err)
    }
  }

  const stopWebcam = () => {
    if (videoRef.current?.srcObject) {
      const tracks = videoRef.current.srcObject.getTracks()
      tracks.forEach((track) => track.stop())
      videoRef.current.srcObject = null
    }
  }

  const startRecording = async () => {
    try {
      const stream = await navigator.mediaDevices.getUserMedia({ audio: true })
      mediaRecorder.current = new MediaRecorder(stream)
      mediaRecorder.current.ondataavailable = (event) => {
        audioChunks.current.push(event.data)
      }
      mediaRecorder.current.onstop = processRecording
      mediaRecorder.current.start()
      setRecording(true)
    } catch (err) {
      console.error("Error accessing microphone:", err)
    }
  }

  const stopRecording = () => {
    if (mediaRecorder.current) {
      mediaRecorder.current.stop()
      setRecording(false)
    }
  }

  const processRecording = async () => {
    setLoading(true)
    try {
      const audioBlob = new Blob(audioChunks.current, { type: "audio/wav" })
      const formData = new FormData()
      formData.append("audio", audioBlob, "recording.wav")
      formData.append("source_url", selectedAvatar)
      formData.append("prompt", selectedPrompt)

      const response = await fetch("http://127.0.0.1:5000/generate-video", {
        method: "POST",
        body: formData,
      })

      if (!response.ok) {
        throw new Error(`HTTP error! status: ${response.status}`)
      }

      const data = await response.json()

      if (data.error) {
        console.error("Error from server:", data.error)
        setAiResponse("Error generating response. Try again.")
        return
      }

      setAiResponse(data.ai_response)
      setVideoUrl(data.video_url)
    } catch (error) {
      console.error("Fetch error:", error)
      setAiResponse("Connection failed. Check server and try again.")
    } finally {
      setLoading(false)
      audioChunks.current = []
    }
  }

  return (
    <div className="min-h-screen bg-gradient-to-br from-purple-500 to-pink-500 p-8">
      <div className="container space-y-8">
        <h1 className="text-4xl font-bold text-white text-center">ConvoCoach</h1>

        {/* Practice Settings */}
        <div className="card">
          <h2 className="text-2xl font-semibold mb-4">Practice Settings</h2>
          <div className="grid grid-cols-1 md:grid-cols-2 gap-6">
            <div>
              <label className="block text-sm font-medium text-gray-700 mb-2">
                Choose Your Speaker:
              </label>
              <select
                className="w-full p-2 border rounded-lg"
                value={selectedAvatar}
                onChange={(e) => setSelectedAvatar(e.target.value)}
              >
                {avatars.map((avatar, index) => (
                  <option key={index} value={avatar.url}>
                    {avatar.name}
                  </option>
                ))}
              </select>
            </div>
            <div>
              <label className="block text-sm font-medium text-gray-700 mb-2">
                Select a Prompt:
              </label>
              <select
                className="w-full p-2 border rounded-lg"
                value={selectedPrompt}
                onChange={(e) => setSelectedPrompt(e.target.value)}
              >
                <option value="">Choose a prompt...</option>
                {prompts.map((prompt, index) => (
                  <option key={index} value={prompt}>
                    {prompt}
                  </option>
                ))}
              </select>
            </div>
          </div>
        </div>

        {/* Flexbox container for Webcam and AI Assistant */}
        <div className="flex-section">
          {/* Webcam Section */}
          <div className="flex-1 card">
            <h3 className="text-xl font-semibold mb-2 text-[#1e3a8a]">Your Webcam</h3>
            <div className="video-container">
              {webcamActive ? (
                <video
                  ref={videoRef}
                  autoPlay
                  playsInline
                  muted
                  className="w-full h-auto object-contain"
                />
         
[truncated — 2641 more characters]
```

### vite.config.js

```javascript
import { defineConfig } from "vite";
import react from "@vitejs/plugin-react";

export default defineConfig({
  plugins: [react()],
});


```

### index.html

```html
<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <meta name="viewport" content="width=device-width, initial-scale=1.0">
    <title>ConvoCoach</title>
    <script type="module" src="/src/main.jsx"></script>
</head>
<body>
    <div id="root"></div>
</body>
</html>


```

### test_speech_analysis.py

```python
import openai
import os
import re
import numpy as np
from dotenv import load_dotenv
import user_speech_to_text

# Load .env for API Key
load_dotenv()
api_key = os.getenv("OPENAI_API_KEY")

if not api_key:
    raise ValueError("API key not found. Make sure '.env' exists and contains 'OPENAI_API_KEY'.")

# Initialize OpenAI client
client = openai.OpenAI(api_key=api_key)

#TODO: Add text speech feedback (num words in response/num_words)
def generate_avatar_response(transcribed_text):
    """
    Uses OpenAI API to analyze speech and provide conversational feedback while keeping the topic in mind.
    """

    prompt = f"""
    Imagine you are a **vocal coach assistant** designed to help a wide variety of users improve their speech clarity while keeping the conversation flowing.
    Your task is to:
    1️⃣ **Analyze the user's speech transcript** and look for:
       - **Unfinished thoughts** (sentences that end abruptly).
       - **Repeated words** (signs of stuttering or hesitation).
       - **Filler words** (uh, um, like, you know, actually, basically, so).
       - **Long pauses** (indicated by "...").
    2️⃣ **Continue the conversation naturally** by responding in a way that keeps the user's topic in mind.
    3️⃣ **Provide subtle, encouraging feedback** in a conversational manner—avoid sounding too robotic or overly critical.

    **Rules for Your Response:**
    - Be **supportive and friendly**.
    - Your response should be **2-3 short sentences**.
    - First, **engage with what the user was trying to say**.
    - Then, **gently provide feedback** to help them improve.
    - Most importantly, only give feedback **when necessary**. If speech was good and coherent, compliment/commend them.
    
    **User's Speech Transcript:**
    "{transcribed_text}"

    Based on this transcript, respond as if you were talking to them in real-time. Keep the conversation flowing while subtly helping them improve.
    """

    response = client.chat.completions.create(
        model="gpt-4-turbo",
        messages=[{"role": "user", "content": prompt}],
        temperature=0.6,  # Balances structure and variation
        max_tokens=200  # Ensures concise responses
    )

    return response.choices[0].message.content.strip()


# Transcript (Test Case)
def get_response():
    audio_file = user_speech_to_text.record_audio()
    transcript = user_speech_to_text.transcribe_audio(audio_file)
    transcript_text = transcript.get("text", "").strip()

    if not transcript_text:
        print("⚠️ No valid transcription received. Exiting...")
        exit()

    # Get AI avatar response
    avatar_response = generate_avatar_response(transcript_text)
    print(avatar_response)
    return avatar_response

```

### user_speech_to_text.py

```python
import os
import time
import wave
import assemblyai as aai
from dotenv import load_dotenv

# Load API Key
load_dotenv(".env")
aai.settings.api_key = os.getenv("ASSEMBLYAI_API_KEY")

if not aai.settings.api_key:
    raise ValueError("AssemblyAI API key is missing. Make sure it's set in .env1!")

# Audio Configuration
RATE = 16000  # Sample rate
CHANNELS = 1  # Mono
RECORD_SECONDS = 15  # Duration to record


def record_audio():
    """Records audio from the microphone and saves it as a WAV file."""
    import pyaudio

    filename = f"user_audios/recorded_audio_{int(time.time())}.wav"

    p = pyaudio.PyAudio()
    stream = p.open(format=pyaudio.paInt16,
                    channels=CHANNELS,
                    rate=RATE,
                    input=True,
                    frames_per_buffer=1024)

    print(f"Recording for {RECORD_SECONDS} seconds...")
    frames = []

    for _ in range(0, int(RATE / 1024 * RECORD_SECONDS)):
        data = stream.read(1024)
        frames.append(data)

    print("Recording complete.")

    stream.stop_stream()
    stream.close()
    p.terminate()

    os.makedirs("user_audios", exist_ok=True)
    with wave.open(filename, 'wb') as wf:
        wf.setnchannels(CHANNELS)
        wf.setsampwidth(p.get_sample_size(pyaudio.paInt16))
        wf.setframerate(RATE)
        wf.writeframes(b''.join(frames))

    return filename  # Return the filename for processing


def transcribe_audio(audio_file):
    """
    Transcribes the given audio file and returns the transcript + sentiment analysis.
    
    :param audio_file: Path to the audio file to transcribe.
    :return: A dictionary containing transcription text and sentiment analysis results.
    """
    transcriber = aai.Transcriber()
    config = aai.TranscriptionConfig(sentiment_analysis=True)
    
    transcript = transcriber.transcribe(audio_file, config)
    try:
        os.remove(audio_file)
    except Exception as e:
        print(f"Error deleting file: {e}")

    return {
        "text": transcript.text,
        "sentiment_analysis": [
            {
                "text": result.text,
                "sentiment": result.sentiment,
                "confidence": result.confidence,
                "start": result.start,
                "end": result.end
            }
            for result in transcript.sentiment_analysis
        ]
    }


# Allow this script to be run standalone for testing
if __name__ == "__main__":
    audio_file = record_audio()
    print("Transcribing and analyzing sentiment...")
    result = transcribe_audio(audio_file)

    print("\nTranscription:")
    print(result["text"])

    print("\nSentiment Analysis Results:")
    for entry in result["sentiment_analysis"]:
        print(f"Text: {entry['text']}")
        print(f"Sentiment: {entry['sentiment']}")
        print(f"Confidence: {entry['confidence']}")
        print(f"Timestamp: {entry['start']} - {entry['end']}\n")

```

### avatar.py

```python
import requests
import time
from dotenv import load_dotenv
import os
import test_speech_analysis as chat_response

load_dotenv()  # Load variables from .env file

api_key = os.getenv('AVATAR_API_KEY')
if api_key is None:
    raise ValueError("AVATAR_API_KEY environment variable is not set")




# URL of the avatar image
# source_url = 'https://media-hosting.imagekit.io//4078ec62a6824c90/screenshot_1739688772692.png?Expires=1834296773&Key-Pair-Id=K2ZIVPTIP2VGHC&Signature=LwWvBbL2shd0AqMt~XyVwUwxey2AGKgPPrMEAql3lULNTvG8mF~Urcxhpy7KvW2Cy2LL1VWU69UBsOXc5SWU9FgsvvVz03ggSbtVtApbitZ2eh7uKwiv0rzsvhTFCwmJBxPpiU9D-CdY~ynDhlHodH~xE8wpd2WYAzdChBhLW2QfeEPDs3zP12LXeZFnO0-05Xco0LTnwQOUK12ZDUsIZLtkWIln6Ooqkb9R-BjBDKT9zGnw5IY5D2JHWrTiJOLX3H0W0Ti65BcXYrbYiILndV2xVM-aIfVPoCtz2bYrSPswkG5yPYDErl79jApaQCvIc3mjZBDvCpEjC4u2U-7xGw__'
# Henok
source_url = 'https://media-hosting.imagekit.io//b9a94f2580c647be/screenshot_1739696990595.png?Expires=1834304991&Key-Pair-Id=K2ZIVPTIP2VGHC&Signature=QoHsh~Mtks6TznmRGXzCjBoPPts-fRoXzjB~3arufkhEKO5eRo2GTIV8teM8PxseAVlYJJcV3am5~397h0K0-CPSzmD6UnkDcvJfpCPoYPzovU3hhWqXbGKD4Q~8YY5GcLwCuYbxiGAThLVqUgonT8tY2EuuE07YQTyM9XUHqJ5OmLauk9ikO~k6bkk1H75DecK0rMMI0CO3oHLs72fvE4LLNzvgIXiTmyzuDC0JxmvdZGsA6oZUghuLxtCGZr8DcvKNGlM~rlCVeGVTHWgl6RluUxrfdxfKXg9jLj7x0GlPia3M~VSkS9MW8C8jnCZJzNXN4msVV5iJh~bT5Tyfdg__'
# Text you want the avatar to speak
text_input = chat_response.get_response()
print("received response from text_speech_analysis")
# Endpoint to create a talk
create_talk_url = 'https://api.d-id.com/talks'

# Headers for the request
headers = {
    'Authorization': f'Basic {api_key}',
    'Content-Type': 'application/json'
}

# Data payload for the request
data = {
    'source_url': source_url,
    'script': {
        # "provider": {
        #     "type": "amazon",
        #     "voice_id": "Joey"
        #     },
        'type': 'text',
        'input': text_input
    }
}

response = requests.post(create_talk_url, headers=headers, json=data)

if response.status_code == 201:
    talk_id = response.json().get('id')
    print(f'Talk created successfully with ID: {talk_id}')
else:
    print(f'Failed to create talk: {response.status_code} - {response.text}')
    exit()

# Endpoint to retrieve the generated video
get_talk_url = f'https://api.d-id.com/talks/{talk_id}'


# Polling to check the status of the video generation
while True:
    response = requests.get(get_talk_url, headers=headers)
    if response.status_code == 200:
        talk_data = response.json()
        if talk_data.get('status') == 'done':
            result_url = talk_data.get('result_url')
            print(f'Video generated successfully: {result_url}')
            break
        elif talk_data.get('status') == 'error':
            print('Error in video generation.')
            break
        else:
            print('Video is being processed...')
    else:
        print(f'Failed to retrieve talk: {response.status_code} - {response.text}')
        break
    time.sleep(5)  # Wait for 5 seconds before polling again
```

[8 more indexed source files omitted to keep this export small. The full file list is in the Codebase structure section above.]