Project Info
This project did not submit a demo video on Devpost.
Inspiration
Current Vision Language Models (VLMs), despite their ability to learn "common sense" from internet data and handle long-tail cases, fall short in autonomous driving applications due to two critical limitations: they can't process high frame-rate video inputs and can't meet real-time latency requirements. This gap presents significant implications for real-world applications: Autonomous vehicle development needs better analysis of edge cases Driving schools require objective assessment tools Insurance companies seek efficient claim verification methods Law enforcement needs automated traffic violation analysis
What it does
TeslaQA automatically analyzes traffic videos and answers questions about driver behavior, traffic rules, and road safety. By first converting video content into descriptive text using BLIP, and then leveraging ChatGPT for question-answering, our system provides detailed explanations of traffic scenarios.
How we built it
Used BLIP to convert video frames into detailed text descriptions Employed ChatGPT API to analyze these descriptions and answer specific questions Created a pipeline that processes 5-second traffic videos Developed custom prompt engineering to ensure accurate and relevant responses
Challenges we ran into
Initially struggled with TimeSformer model and BLIP model for video understanding Faced issues with different video frame rates and lengths Had to optimize the text descriptions to be both concise and informative Needed to carefully design prompts to get consistent, accurate answers
Accomplishments we're proud of
Successfully combined computer vision and language models Created an interpretable system that can explain its reasoning Achieved accurate answers for complex traffic scenarios Built a scalable solution that can handle various traffic situations
What we learned
The importance of model selection for specific tasks How to effectively combine multiple AI models The value of converting visual data to text for better interpretability Techniques for prompt engineering with GPT models
What's next
Fine-tune BLIP model for specialized traffic understanding: Generate more precise descriptions of vehicle movements, signals, and road conditions Better capture critical driving behaviors and safety moments Develop traffic-specific visual attention mechanisms Fine-tune BLIP model for specialized traffic understanding: Generate more precise descriptions of vehicle movements, signals, and road conditions Better capture critical driving behaviors and safety moments Develop traffic-specific visual attention mechanisms Enhance text summary quality through: Adding detailed spatial relationships between vehicles Incorporating multi-frame temporal understanding Building specialized traffic vocabulary and context Enhance text summary quality through: Adding detailed spatial relationships between vehicles Incorporating multi-frame temporal understanding Building specialized traffic vocabulary and context Technical Improvements: Integrate with existing traffic monitoring systems Develop APIs for easy integration Create efficient frame selection algorithms Optimize for real-time processing Technical Improvements: Integrate with existing traffic monitoring systems Develop APIs for easy integration Create efficient frame selection algorithms Optimize for real-time processing Model Enhancement: Train on larger, more diverse traffic datasets Implement more sophisticated prompting strategies Add multi-camera support Develop hierarchical summarization capabilities Model Enhancement: Train on larger, more diverse traffic datasets Implement more sophisticated prompting strategies Add multi-camera support Develop hierarchical summarization capabilities Our goal is to make roads safer by providing better tools for understanding and analyzing traffic scenarios, ultimately contributing to the development of more reliable autonomous driving systems and better driver education. The kaggle link is listed below. And our team name is 'passionfruit.'
2025treehacks_tesla
https://www.kaggle.com/competitions/tesla-real-world-video-q-a/overview
Analysis
View
Metric
- 4
Figures cover GitHub contributors during the hackathon window. A co-authored commit counts in full for each author, so per-member totals add up to more than the whole-team figures.
Technology
- OpenAIClaimed
- PythonClaimed
0 of 2 appear in the indexed code. 2 claimed on Devpost could not be matched to code, which may simply mean the tool leaves no trace in the repository.
AI coding agents
No AI coding agent signals were found in this repository.
Detected from committed agent config files and commit authorship. Absence of a signal is not proof an agent was unused.
Codebase size
Source size
95 B
Source files
1
Counts recognized source files only; vendored directories, binaries and lockfiles are excluded, so this is smaller than the repository on disk.
Repository
HangLiu01/2025treehacks_tesla
3 files · 53 KB · @ 9a383a7
Structure
Application logic
2 files · 67%Domain rules, services and shared utilities.
Supporting
Layers are inferred from where files sit in the tree, not from reading the code. A project that names its directories unconventionally will read oddly here — open the file browser to check anything the diagram implies.
Languages
- Markdown100%
Share of indexed source by file size. Binary and vendored files are excluded.
This project’s features have not been analysed yet.
Export this project's context (description, README, evidence, key source files) to chat with an AI agent elsewhere.