Project Info
The Team Taha Abdullah, Computer Science @ UC Davis '24 [in] Yahya Qteishat, Electrical Engineering & Computer Sciences @ UC Berkeley '24 [in] Youssef Qteishat, Computer Science @ UC Davis '24 [in] Nour Zahzah, Mechanical Engineering @ UC Berkeley '27 [in]
Inspiration
Struggles of immigrants and refugees navigating language barriers Hands-off user experiences offered by commercial smart glasses (Meta glasses, Google Glass, AR, etc.)
What it does
Glasses capture a live stream of the user's POV via CAM module attached to glasses frame Snapshot triggered on dev GUI ESP32 sends POV image to Flask backend hosted on Render (mocked in demo video bc of SSL cert issues) Gemini interprets the language and translates into English, and vice versa Translated text is displayed in server console and LCD display (WIP)
How we built it
Hardware: Nour custom modeled and printed the following components: glasses frame with integrated camera housing, LCD display mount Middleware: Youssef and Taha wrote the Python (Flask) middleware server, hosted on Render, that integrates with Gemini for multimodal translation of uploaded image Electronics: Yahya integrated the ESP32-CAM micro-controller with a 2 MP camera and the FTDI USB-to-TTL adapter Programmed the ESP32 to capture the image through the camera and pass that as the body of a POST request to the api endpoint. (mocked in the demo bc running into cert issues on ESP32)
Challenges we ran into
Getting a snap fit when attaching the camera module to its holder on the frame Handling latency when making API calls from ESP32 to Flask Server ESP32-CAM doesn't have a usb chip. First tried using the usb chip from a second USB-32 but we didn't have a micro-usb data transfer cable. so we switched to the FTDI approach. HTML on ESP32 is a zipped archive thats cut up into a C byte array, very cumbersome to change Displaying languages with non-Latin and especially cursive scripts
Accomplishments we're proud of
The quality of the hardware! Nour did an amazing amazing job. Establishing a wifi connection between the ESP32 and Flask backend Getting live video feed from CAM module Hosting Flask servers on Render
What we learned
Prompt Engineering Integrating an ESP32-CAM micro-controller with an FTDI USB-TTL Adapter Designing CAD models in Fusion 360 3D printing models Refinements Triggering image capture and translation from button on the side instead of from a dev GUI on laptop Finishing the integration of our LCD screen so we can display the output there isntead of in server console
What's next
Integrating a GPS module to get the location to account for differences between standard and colloquial language, updating the user prompt accordingly Utilizing a microphone, speaker, and STT/TTS provider to translate a conversation in real-time Bigger display and additional buttons for custom user configuration
SnapLens
Frictionless Translation for social good
The Team
Taha Abdullah, Computer Science @ UC Davis '24 [in] Yahya Qteishat, Electrical Engineering & Computer Sciences @ UC Berkeley '24 [in] Youssef Qteishat, Computer Science @ UC Davis '24 [in] Nour Zahzah, Mechanical Engineering @ UC Berkeley '27 [in]
Inspiration
- Struggles of immigrants and refugees navigating language barriers
- Hands-off user experiences offered by commercial smart glasses (Meta glasses, Google Glass, AR, etc.)
What it does
- Glasses capture a live stream of the user's POV via CAM module attached to glasses frame
- Snapshot triggered on dev GUI
- ESP32 sends POV image to Flask backend hosted on Render (mocked in demo video bc of SSL cert issues)
- Gemini interprets the language and translates into English, and vice versa
- Translated text is displayed in server console and LCD display (WIP)
How we built it
Hardware:
- Nour custom modeled and printed the following components: glasses frame with integrated camera housing, LCD display mount
Middleware:
- Youssef and Taha wrote the Python (Flask) middleware server, hosted on Render, that integrates with Gemini for multimodal translation of uploaded image
Electronics:
- Yahya integrated the ESP32-CAM micro-controller with a 2 MP camera and the FTDI USB-to-TTL adapter
- Programmed the ESP32 to capture the image through the camera and pass that as the body of a POST request to the api endpoint. (mocked in the demo bc running into cert issues on ESP32)
Challenges we ran into
- Getting a snap fit when attaching the camera module to its holder on the frame
- Handling latency when making API calls from ESP32 to Flask Server
- ESP32-CAM doesn't have a usb chip. First tried using the usb chip from a second USB-32 but we didn't have a micro-usb data transfer cable. so we switched to the FTDI approach.
- HTML on ESP32 is a zipped archive thats cut up into a C byte array, very cumbersome to change
- Displaying languages with non-Latin and especially cursive scripts
Accomplishments that we're proud of
- The quality of the hardware! Nour did an amazing amazing job.
- Establishing a wifi connection between the ESP32 and Flask backend
- Getting live video feed from CAM module
- Hosting Flask servers on Render
What we learned
- Prompt Engineering
- Integrating an ESP32-CAM micro-controller with an FTDI USB-TTL Adapter
- Designing CAD models in Fusion 360
- 3D printing models
Refinements
- Triggering image capture and translation from button on the side instead of from a dev GUI on laptop
- Finishing the integration of our LCD screen so we can display the output there isntead of in server console
What's next for SnapLens
- Integrating a GPS module to get the location to account for differences between standard and colloquial language, updating the user prompt accordingly
- Utilizing a microphone, speaker, and STT/TTS provider to translate a conversation in real-time
- Bigger display and additional buttons for custom user configuration
Analysis
View
Metric
- 3
- 1
Figures cover GitHub contributors during the hackathon window. A co-authored commit counts in full for each author, so per-member totals add up to more than the whole-team figures.
Technology
- CIn code
- C++In code
- FlaskIn code
- PythonIn code
- Google GeminiClaimed
4 of 5 appear in the indexed code. 1 claimed on Devpost could not be matched to code, which may simply mean the tool leaves no trace in the repository.
AI coding agents
No AI coding agent signals were found in this repository.
Detected from committed agent config files and commit authorship. Absence of a signal is not proof an agent was unused.
Codebase size
Source size
184 KB
Source files
5
Counts recognized source files only; vendored directories, binaries and lockfiles are excluded, so this is smaller than the repository on disk.
Repository
tmabdull/snap-lens
10 files · 190 KB · @ b0a27ab
Structure
Application logic
7 files · 70%Domain rules, services and shared utilities.
Supporting
Layers are inferred from where files sit in the tree, not from reading the code. A project that names its directories unconventionally will read oddly here — open the file browser to check anything the diagram implies.
Languages
- C82%
- C++14%
- Markdown2%
- Python2%
Share of indexed source by file size. Binary and vendored files are excluded.
Dependencies
requirements.txt
pypi · 6- arabic-reshaper
- Flask
- google-genai
- gunicorn
- Pillow
- python-dotenv
Declared in the repository’s manifests at the indexed commit. A declared package is not proof it is used, and runtime dependencies are listed first.
This project’s features have not been analysed yet.
Export this project's context (description, README, evidence, key source files) to chat with an AI agent elsewhere.