BeaverHacksJudge
Back to Projects

SPOOT: Sound Point Of Origin Tracker

A real-time sound awareness system that identifies where sounds come from and uses Gemini to evaluate their importance.

1 / 2
About This Project

Inspiration

When searching for a project idea, Carson's roommate Ian shared a unique challenge related to his disability. Ian is deaf in one ear, which makes it difficult for him to tell where sounds are coming from. For example, if an ambulance passes by, someone yells, or a loud noise happens nearby, he can hear the sound but cannot easily identify its direction.

Ian also mentioned that it can be difficult to tell when someone is calling his name, especially in noisy environments. We wanted to build something that could help people with single-sided deafness quickly understand where important sounds are coming from and whether those sounds are worth paying attention to.

What It Does

SPOOT detects the direction of nearby sounds in real time and displays that direction through a visual interface. When a sound is detected, the system immediately shows where it is coming from, helping the user know where to look.

The system also listens for speech and uses Gemini to understand the context of what was said. Gemini helps SPOOT decide whether the sound is likely directed at the user, whether it is urgent, and how it should be shown on the HUD. For example, if someone says the user's name or gives a warning like "watch out," Gemini can mark that sound as high priority and generate a short, clear message for the user.

In short, SPOOT helps answer two important questions:

  1. Where did that sound come from?
  2. Does Gemini think that sound matters to me?

How We Built It

We built SPOOT using a Raspberry Pi, a ReSpeaker microphone array, a Python backend, Gemini, and a web-based HUD. The microphone array captures audio and provides direction-of-arrival data, which tells us the angle that a sound is coming from.

Our Python backend processes the microphone data, measures volume levels, and broadcasts sound events over WebSockets. The frontend receives those updates and displays directional alerts in real time.

Gemini acts as the intelligence layer of the project. After the system detects speech, it sends the transcript, sound direction, and volume information to Gemini. Gemini then evaluates whether the sound is likely directed at the user, how important it is, what type of event it represents, and what short message should be shown on the HUD.

We designed the system so that the direction appears immediately, while Gemini runs afterward as a smarter reasoning layer. This keeps the HUD fast while still allowing the system to provide meaningful context instead of only showing raw sound data.

Challenges We Ran Into

The biggest challenge we ran into was at the very start. When we first started using the microphone array, we tried logging the sound direction to the console, but the direction would never change. At first, we thought we had broken the microphone. Eventually, we discovered that only 2 of the 4 microphones were enabled.

Building the physical device was also challenging, especially the glasses. We 3D-printed most of the holders for our parts, which took several hours and required us to carefully think through how everything would fit together.

Another challenge was balancing speed and intelligence. We wanted the direction indicator to update almost instantly, but Gemini reasoning takes longer than raw direction detection. To solve this, we split the system into two stages: an immediate directional alert and a slower Gemini-powered importance update.

Accomplishments That We Are Proud Of

We are proud that SPOOT combines hardware, real-time audio processing, Gemini reasoning, and a visual interface into one working assistive prototype.

We are especially excited that the project is based on a real problem shared by someone close to us. Instead of building a random demo, we created something that could genuinely help people with single-sided deafness feel more aware of their surroundings.

We are also proud that we were able to get the microphone array, WebSocket communication, speech transcription, Gemini integration, and HUD working together in a short amount of time. Gemini helped us move the project beyond simply pointing toward sound by allowing SPOOT to reason about whether the sound is actually important.

What We Learned

We learned the importance of planning hardware designs before building. We had an idea and the parts we needed, but we struggled at first to physically connect everything in a clean and wearable way.

We also learned more about audio direction detection, microphone arrays, speech-to-text, WebSockets, and how to use Gemini as a reasoning layer. One important design lesson was that Gemini should not block the real-time HUD. Instead, it works best as a second layer that adds meaning after the direction is shown.

Most importantly, we learned that assistive technology is not just about detecting information. It is about presenting the right information at the right time in a way that actually helps the user.

What Is Next

After the hackathon, we would like to improve SPOOT by adding an external sound classification model that can identify non-speech sounds, such as cars, sirens, alarms, knocks, or explosions. Gemini could then use those labels, along with direction and volume, to provide richer context about the user's environment.

We would also like to make the device smaller, more comfortable, and easier to wear. In the future, SPOOT could support more personalized Gemini-powered alerts, better noise filtering, and stronger integration with smart glasses or other wearable devices.

Google — Best use of Gemini
Team: Sneaky Squad
GitHub
Team Members
  • Carson Secrest

    Carson Secrest

  • Jakhangir  Tynshimov

    Jakhangir Tynshimov