# Project export: Pet Talks

This document was generated by HackStack to give an AI agent context about a hackathon project. Sections are labeled with their provenance; content marked as truncated was cut to keep this document small.

## Project metadata

- Hackathon: TreeHacks 2025
- Tagline: Pet Talks: a pet camera that allows you to converse with your pet by detecting gestures, movements, and sounds. Scans for unusual behavior leading to medical problems.
- Devpost: https://devpost.com/software/pet-talks
- GitHub: https://github.com/joshua-linsanity/pet-talk
- Video: https://www.youtube.com/embed/bEU4Y0NPwXI?enablejsapi=1&hl=en_US&rel=0&start=&version=3&wmode=transparent
- Team: 2 GitHub contributor(s) — Joshua Lin (1 commits), Jenny Huynh (1 commits)

## Devpost submission (written by the team)

### Overview

Introducing Pet-Talks! 🐰🥕💬

### Inspiration

Instead of having a one-sided conversation, imagine having a conversation with your pet, where they speak in human words. Your Disney princess dreams can finally come true! Introducing Pet Talks: have conversations with your pet by detecting gestures, movements, and even sounds. Pet Talks can also detect unusual behavior/gestures that can quickly detect medical problems. As bunny owners, we find ourselves talking to our bunny every day as if she were a human. Our pets are emotional and sophisticated creatures, and we are always striving to connect with our pets in any way possible. Have you ever looked at your pet and asked yourself: what is going on inside their head? What if… we can vocalize those emotions? Instead of having a one-sided conversation, imagine having a conversation with your pet, where they speak in human words.

### What it does

Introducing Pet Talks, a pet cam that allows you to have conversations with your pet by detecting gestures, movements, and even sounds. Pet Talks can also detect unusual behavior or gestures that can quickly detect medical problems.GI stasis is a common and potentially fatal condition affecting thousands of rabbits daily. Detecting its early signs can be challenging, as changes in their behavior and eating habits often go unnoticed when they're out of sight. "Pet Talks – Because every hop, twitch, and nibble has a story to tell!"

### How we built it

We started implementing a desktop version, which uses our phone as the webcam and our laptop as the chat tool. The webcam takes short videos or bursts of photos every time we start typing, and hit enter. Those photos are then fed through the LLM model configured, which generates a medical response or simply a live status update for the bunny.

### Challenges we ran into

We initially attempted to use the Meta Oculus to implement a live video feed from the webcam. However, the Oculus had many issues starting up the wireless connection. Also, at first, we asked Gemini to interact with the user as if it were a rabbit, which posed a challenge as the responses seemed too mindless and lacked reliable information. We would prompt, “Are you healthy?” and the bunny would respond, “I am just a cute fluffy bunny. Give me carrots!” This provided no valid information on her health or actions.

### Accomplishments we're proud of

We began tailoring Gemini to give a more accurate response through better instructions and some anti-jail break methods. However, we were also able to utilize multiple LLM solutions such as Gemini, PerplexityAI, and OpenAI’s models, hinting at a future where the user experience is fully customizable. The perplexity API call that we set up to be tailored is really nice since we can gather a detailed report with minimal effort.

### What we learned

We learned that LLMs are difficult to effectively finetune, which is why we need to make use of the existing sponsor resources provided to us. We need a balance between cool technology and what is already familiar; since we didn’t want to be focusing too much on making use of technology, but rather let the passion in our idea make itself known. The presentation is also equally as crucial in our project since that is our user’s first impression, and not all group members have experience DEMOing so we are very thrilled to be able to have fun during this technical challenge. What's Next for Pet Talks With more time, we plan to create an iOS version in which users can have better accessibility when viewing real-time footage from a remote location. The IOS app will be paired with a 360 camera system with AI detection, geofencing, and of course, the chat and pet translation feature, and the health check as well. The app will alert users wherever unusual activity occurs, such as when the bunny hasn't been eating for the past 3 hours, the user will be alerted and if this goes on for more than 12, the user will get a GI status alert and description, the user can choose to automatically send the message to their local vet, and book a vet appointment if severe. We also plan to do text-to-speech, and users can hear their pet speak in unique voices.

## README (from the GitHub repository)

## Introducing Pet-Talks! 🐰🥕💬
### Inspiration
Instead of having a one-sided conversation, imagine having a conversation with your pet, where they speak in human words. Your Disney princess dreams can finally come true! Introducing Pet Talks: have conversations with your pet by detecting gestures, movements, and even sounds. Pet Talks can also detect unusual behavior/gestures that can quickly detect medical problems.
As bunny owners, we find ourselves talking to our bunny every day as if she were a human. Our pets are emotional and sophisticated creatures, and we are always striving to connect with our pets in any way possible. Have you ever looked at your pet and asked yourself: what is going on inside their head? What if… we can vocalize those emotions? Instead of having a one-sided conversation, imagine having a conversation with your pet, where they speak in human words. 

### What it does
Introducing Pet Talks, a pet cam that allows you to have conversations with your pet by detecting gestures, movements, and even sounds. Pet Talks can also detect unusual behavior or gestures that can quickly detect medical problems.GI stasis is a common and potentially fatal condition affecting thousands of rabbits daily. Detecting its early signs can be challenging, as changes in their behavior and eating habits often go unnoticed when they're out of sight.
"Pet Talks – Because every hop, twitch, and nibble has a story to tell!"
### How we built it
We started implementing a desktop version, which uses our phone as the webcam and our laptop as the chat tool.
The webcam takes short videos or bursts of photos every time we start typing, and hit enter. Those photos are then fed through the LLM model configured, which generates a medical response or simply a live status update for the bunny. 
### Challenges we ran into
We initially attempted to use the Meta Oculus to implement a live video feed from the webcam. However, the Oculus had many issues starting up the wireless connection.
Also, at first, we asked Gemini to interact with the user as if it were a rabbit, which posed a challenge as the responses seemed too mindless and lacked reliable information. We would prompt, “Are you healthy?” and the bunny would respond, “I am just a cute fluffy bunny. Give me carrots!” This provided no valid information on her health or actions. 
### Accomplishments that we're proud of
We began tailoring Gemini to give a more accurate response through better instructions and some anti-jail break methods. However, we were also able to utilize multiple LLM solutions such as Gemini, PerplexityAI, and OpenAI’s models, hinting at a future where the user experience is fully customizable. 
The perplexity API call that we set up to be tailored is really nice since we can gather a detailed report with minimal effort.
###  What we learned
We learned that LLMs are difficult to effectively finetune, which is why we need to make use of the existing sponsor resources provided to us. We need a balance between cool technology and what is already familiar; since we didn’t want to be focusing too much on making use of technology, but rather let the passion in our idea make itself known.
The presentation is also equally as crucial in our project since that is our user’s first impression, and not all group members have experience DEMOing so we are very thrilled to be able to have fun during this technical challenge.
### What's Next for Pet Talks
With more time, we plan to create an iOS version in which users can have better accessibility when viewing real-time footage from a remote location. The IOS app will be paired with a 360 camera system with AI detection,  geofencing, and of course, the chat and pet translation feature, and the health check as well. The app will alert users wherever unusual activity occurs, such as when the bunny hasn't been eating for the past 3 hours, the user will be alerted and if this goes on for more than 12, the user will get a GI status alert and description, the user can choose to automatically send the message to their local vet, and book a vet appointment if severe. We also plan to do text-to-speech, and users can hear their pet speak in unique voices.


## Detected evidence (automated analysis)

Indexed codebase: 4 recognized source files, 35 KB.
- Python (language) — detected in the code
- OpenAI (technology) — claimed on Devpost, not found in the code

## Codebase structure (from repository index)

### Files (7 of 7)

```
.gitignore
capture/.DS_Store
capture/diff
capture/helpers.py
capture/main.py
capture/video.py
README.md
```

### Dependencies

No dependency index available.

### Recent commits (newest first)

- Update README.md
- perplexity!
- complete
- UI FIXED
- updated ui
- pre-background attempt
- the bubble readjusts!
- switched to openai
- updating loading, need scroll fix
- loading symbol + shorter responses
- working with gemini
- calm before the storm
- pre-implementation of vision query
- better ui
- windows semi-working
- combined panel
- changed ui design to blue for ios
- initial push
- Initial commit

## Key source files (fetched from GitHub, selected and truncated for size)

### capture/main.py

```python
import sys
import cv2
import PIL
import random
import base64
from PyQt5.QtWidgets import *
from PyQt5.QtCore import *
from PyQt5.QtGui import *
from openai import OpenAI

worker = OpenAI()

##############################
# Worker Thread for Open AI  #
##############################

def get_load_message():
    messages = [
        "Hopping right into action, hold on a sec!",
        "Nibbling on some carrots of data, please wait!",
        "Hang tight, I'm leaping into your request!",
        "Carrots in hand, I’m busy crunching details!",
        "Hold your ears, I'm bounding over to it!",
        "Just a bunny hop away, almost there!",
        "Gathering carrots and cuddles, please wait!",
        "Ears up, I'm whisking up your request!",
        "In a hoppy hurry, your info is coming!",
        "Binkying through data, please hold on!",
        "Hopping into action, one carrot at a time!",
        "Pause for a moment, I'm on a bunny trail!",
        "Hare-ing on to the details, please wait!",
        "Bunny breath in, bunny breath out, almost ready!",
        "Tail wags and bunny hops, your info is near!",
        "Just a hop and a skip away from completion!",
        "Crunching carrots and data, hang tight!",
        "Hold tight, I'm just nibbling through the details!",
        "Hop, skip, and wait a bit longer, almost there!",
        "Bunny-speed loading in progress, hold your whiskers!"
    ]
    return random.choice(messages)

def create_circular_pixmap(image_path, size):
    # Load the image
    pixmap = QPixmap(image_path)
    if pixmap.isNull():
        # Create a placeholder pixmap if loading fails
        pixmap = QPixmap(size, size)
        pixmap.fill(Qt.gray)

    # Crop the pixmap to a square based on the smallest dimension
    s = min(pixmap.width(), pixmap.height())
    rect = QRect((pixmap.width() - s) // 2, (pixmap.height() - s) // 2, s, s)
    pixmap = pixmap.copy(rect)

    # Create a square pixmap with transparency to hold the circular image
    circular = QPixmap(s, s)
    circular.fill(Qt.transparent)

    # Draw the circular image
    painter = QPainter(circular)
    painter.setRenderHint(QPainter.Antialiasing)
    path = QPainterPath()
    path.addEllipse(0, 0, s, s)
    painter.setClipPath(path)
    painter.drawPixmap(0, 0, pixmap)
    painter.end()

    # Scale the resulting pixmap to the desired size
    return circular.scaled(size, size, Qt.KeepAspectRatio, Qt.SmoothTransformation)

def encode_image(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

##############################
# Worker Thread for Open AI  #
##############################

class SweatshopWorker(QObject):
    finished = pyqtSignal(str)
    error = pyqtSignal(str)

    def __init__(self, question, image_path, name, species):
        super().__init__()
        self.question = question
        self.image_path = image_path
        self.name = name
        self.species = species

    def run(self):
        try:
            base64_image = encode_image(self.image_path)
            prompt = (
                    "Determine if the following query is health-related or conversational. "
                    "Examples of health-related queries: 'Are you healthy?' or 'Are you okay?' etc. "
                    "If the query is health-related, reply HEALTH. Else, reply CONVO. "
                    "User query: "
                    )
            response = worker.chat.completions.create(
                model="gpt-4o-mini",
                messages=[
                    {
                        "role": "user",
                        "content": [
                            {
                                "type": "text",
                                "text": prompt + self.question,
                            },
                        ],
                    }
                ],
                max_tokens=300
            )
            response = response.choices[0].message.content
            health = (response == "HEALTH")

            if health: 
                prompt = (
                    f"You are a professional veterinarian specializing in dogs, cats, and bunnies. "
                    "Carefully *analyze the attached image* along with the user query and offer your clinical diagnosis. "
                    "(Note the user query may be addressed to the pet, but respond as a veterinarian. "
                    "If the pet appears healthy, respond as such. "
                    "Otherwise, report the health concerns that may be present in the pet. "
                    "Keep all responses under 1000 characters. "
                    "User query: "
                )
                response = worker.chat.completions.create(
                    model="sonar-pro",
                    messages=[
                        {
                            "role": "user",
                            "content": [
                                {
                                    "type": "text",
                                    "text": prompt + self.question,
                                },
                                {
                                    "type": "image_url",
                                    "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"},
                                },
                            ],
                        }
                    ],
                    max_tokens=300
                )
                response = response.choices[0].message.content

                prompt = ("Consider the following diagnosis by a veterinarian. "
                    f"You are named {self.name}. "
                    "Replace all third person references with first person references. "
                    "Make sure your tone is warm and cute! "
                    "**Make sure to preserve the medical/clinical information in the diagnosis.** "
                )
                response = 
[truncated — 10812 more characters]
```

### capture/helpers.py

```python
##############################
# Round Out the Profile Pic  #
##############################
from PyQt5.QtWidgets import *
from PyQt5.QtCore import *
from PyQt5.QtGui import *
import base64


def create_circular_pixmap(image_path, size):
    # Load the image
    pixmap = QPixmap(image_path)
    if pixmap.isNull():
        # Create a placeholder pixmap if loading fails
        pixmap = QPixmap(size, size)
        pixmap.fill(Qt.gray)

    # Crop the pixmap to a square based on the smallest dimension
    s = min(pixmap.width(), pixmap.height())
    rect = QRect((pixmap.width() - s) // 2, (pixmap.height() - s) // 2, s, s)
    pixmap = pixmap.copy(rect)

    # Create a square pixmap with transparency to hold the circular image
    circular = QPixmap(s, s)
    circular.fill(Qt.transparent)

    # Draw the circular image
    painter = QPainter(circular)
    painter.setRenderHint(QPainter.Antialiasing)
    path = QPainterPath()
    path.addEllipse(0, 0, s, s)
    painter.setClipPath(path)
    painter.drawPixmap(0, 0, pixmap)
    painter.end()

    # Scale the resulting pixmap to the desired size
    return circular.scaled(size, size, Qt.KeepAspectRatio, Qt.SmoothTransformation)

def encode_image(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

```

### capture/video.py

```python
import sys
import cv2
import PIL
import time
from collections import deque
from PyQt5.QtWidgets import *
from PyQt5.QtCore import *
from PyQt5.QtGui import *
from helpers import *
from google import genai
from google.genai import types

client = genai.Client(api_key="AIzaSyARoDWGfPnfxYkzyMhU5I3ULLluH7o6qaM")

##############################
# Worker Thread for Gemini   #
##############################
class GeminiWorker(QObject):
    finished = pyqtSignal(str)
    error = pyqtSignal(str)
    
    def __init__(self, question, video_path, name, species):
        super().__init__()
        self.question = question
        self.video_path = video_path
        self.name = name
        self.species = species

    def run(self):
        try:
            # Determine if the user query is conversational or medical
            message = (
                "Determine if the following query is health-related or conversational. "
                "Examples of health-related queries: 'Are you healthy?' or 'Are you okay?' etc. "
                "If the query is health-related, reply HEALTH. Else, reply CONVO. "
                "User query: "
            )
            code = client.models.generate_content(
                model="gemini-2.0-flash",
                contents=message + self.question
            )
            health = (code.text.strip() == "HEALTH")

            # This is the long-running Gemini call using the video file
            video_file = client.files.upload(path=self.video_path)
            while videO_file.state.name == "PROCESSING":
                time.sleep(1)
                video_file = client.files.get(name=video_file.name)
            if video_file.state.name == "FAILED":
                raise ValueError(video_file.state.name)

            if health: 
                prompt = (
                    f"You are a professional veterinarian specializing in dogs, cats, and bunnies. "
                    "Carefully *analyze the attached video* along with the user query and offer your clinical diagnosis. "
                    "(Note the user query may be addressed to the pet, but respond as a veterinarian. "
                    "If the pet appears healthy, respond as such. "
                    "Otherwise, report the health concerns that may be present in the pet. "
                    "Keep all responses under 1000 characters. "
                    "User query: "
                )
                diagnosis = client.models.generate_content(
                    model="gemini-2.0-flash",
                    contents=[video_file, prompt]
                )

                message = ("Consider the following diagnosis by a veterinarian. "
                    f"You are named {self.name}. "
                    "Replace all third person references with first person references. "
                    "Make sure your tone is warm and friendly. "
                    "**Make sure to preserve the medical/clinical information in the diagnosis.** "
                )
                response = client.models.generate_content(
                    model="gemini-2.0-flash",
                    contents=message + diagnosis.text
                )
                self.finished.emit(response.text)
            else:
                prompt = (
                    f"You are a {self.species} named {self.name}. Consider the video of yourself "
                    "attached. Respond to the human's message in a conversational and cute tone. "
                    "Message: "
                )
                response = client.models.generate_content(
                    model="gemini-2.0-flash",
                    contents=[video_file, prompt + self.question]
                )
                self.finished.emit(response.text)

        except Exception as e:
            self.error.emit(str(e))

##############################
# Video Widget (OpenCV feed) #
##############################
class VideoWidget(QLabel):
    def __init__(self, parent=None):
        super(VideoWidget, self).__init__(parent)
        self.cap = cv2.VideoCapture(0)
        if not self.cap.isOpened():
            print("Cannot open camera")
            sys.exit(1)
        self.timer = QTimer(self)
        self.timer.timeout.connect(self.update_frame)
        self.timer.start(30)
        self.setSizePolicy(QSizePolicy.Expanding, QSizePolicy.Expanding)
        # Buffer to store (timestamp, frame) tuples for the past 5 seconds
        self.frame_buffer = deque()

    def update_frame(self):
        ret, frame = self.cap.read()
        if ret:
            now = time.time()
            # Store a copy of the original BGR frame in the buffer
            self.frame_buffer.append((now, frame.copy()))
            # Remove frames older than 5 seconds
            while self.frame_buffer and (now - self.frame_buffer[0][0] > 5):
                self.frame_buffer.popleft()
            # Convert frame to RGB for display
            rgb_frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
            h, w, ch = rgb_frame.shape
            bytes_per_line = ch * w
            qimg = QImage(rgb_frame.data, w, h, bytes_per_line, QImage.Format_RGB888)
            scaled_pix = QPixmap.fromImage(qimg).scaledToHeight(self.height(), Qt.SmoothTransformation)
            self.setPixmap(scaled_pix)

    def get_recent_video_clip(self, filename="recent_clip.mp4"):
        if not self.frame_buffer:
            return None
        # Get frame size from the first buffered frame
        first_frame = self.frame_buffer[0][1]
        height, width, channels = first_frame.shape
        fps = self.cap.get(cv2.CAP_PROP_FPS)
        if fps == 0 or fps is None:
            fps = 30
        fourcc = cv2.VideoWriter_fourcc(*'mp4v')
        out = cv2.VideoWriter(filename, fourcc, fps, (width, height))
        for ts, frame in self.frame_buffer:
            out.write(frame)
        out.release()
        return filename

    def get_current_frame(self):
        ret, frame = self.cap.read()
        retu
[truncated — 7351 more characters]
```