BeaverHacksJudge
Back to Projects

Punch Harder

Punch Harder. Literally. The AI will judge your form.

1 / 9
About This Project

Inspiration

Self defense is a really important skill, but starting out can be hard. We all wanted to get better at punching things and decided that NVIDIA nemotron-3-nano-omni (Nomni, as we call him) could be our coach.

What It Does

You shadowbox in front of a camera for 15 seconds. We count your punches in real time, classify them, and send the annotated data and video to NVIDIA Nomni. Once the 15 seconds are done, you get a full punch by punch breakdown with form analysis and advice. We also added a small arcade aspect by having a leaderboard.

How We Built It

Lots of caffeine, low fi music, and plenty of snacks.

But on the serious side, we had quite a few moving components. First, we have the React Vite Frontend. We use mediapipe to track your body. We designed a custom algorithm that tracks both your hands and counts the punches. We send this video stream with the punch annotations to our Python Flask backend. For maximum precision, we trained our own YOLOv8 image model on datasets of punches that we sourced from the web. Our model had an average accuracy of 90%. The model analyzes the frames of each punch and classifies the type of punch (hook, uppercut, jab, cross). It also provides a confidence score (better form = higher confidence).

Once we've assembled the punch annotations, we send this data as well as the original video to NVIDIA nemotron-3-nano-omni-30b-a3b-reasoning (Nomni). This is one of the hottest new models out there and we picked it because it is specifically designed to support multimodal input (meaning video in our case).

Once Nomni grades your form, we're nearly done. We return all the feedback to the Frontend which splits the replay into separate bits. Each bit is generally one combo or a punch and transition. Nomni gives specific pieces of advice for each bit. Finally, the icing on the cake. We call NVIDIA magpie_tts to read out the feedback live as you watch the slo mo replay of your punches.

Building all this was not easy and involved a combination of hand coding, vibe coding, and directed agent management. Sometimes this meant politely prompting chat-gpt-5 in cursor while other times it meant impolitely screaming at claude code at 3am to fix the damn bug.

And that's how it works.

Challenges We Ran Into

  • The project has a lot of moving pieces so combining all of them was tricky. We were careful to test components individually before combining them.
  • Our first training attempt of YOLOv8 was run on a CPU. After 3 hours we realized that's a no go. Luckily one of us has an RTX 4070 laptop. The model trained in 30 minutes.
  • The algorithm to detect punches was really tricky. We tried lots of different approaches (elbow angles, velocity, etc.). Eventually we settled on measuring the distance from the wrist of each hand to the body. If the distance is increasing, this marks the start of a punch. We had to play around with the algorithms parameters to get it to count punches correctly.
  • There were no tables at the Toyota club on Saturday :(( so we ended up going to the library :))
  • Occasionally NVIDIA would rate limit us. This would cause random bugs. At 3 am. :L

Accomplishments That We Are Proud Of

We're super proud of finishing a working demo. We trained a model and learned to use different NVIDIA Nemotron technologies. We're also proud of sticking together as a team. And... we all got a little better at punching while working on this ;)

What We Learned

Multimodal models like NVIDIA Nemotron are powerful, but they are not drop-in replacements for a human coach. The model does not “watch” our clips the way we do: video is represented and processed as structured inputs (effectively tokenized multimodal context), so we cannot hand the raw file to the LLM and expect reliable end-to-end analysis. That pushed us to split the problem: local perception (e.g., YOLO for punch windows and types) feeds structured labels and short clips into the model for coaching text, rather than asking one model to do everything from pixels alone.

Also, never train YOLOv8 on a CPU when you have a GPU.

What Is Next

Punch-only analysis is a deliberate first slice; expanding to kicks would mean extending the detector and label schema, retraining or fine-tuning YOLOv8, and revisiting how we align coaching to new strike types—we did not have time for that in this iteration. We also want punch velocity (or proxy metrics from pose/keyframe deltas) so feedback can pair qualitative coaching with numbers users can track over rounds.

We would love to deploy Punch Harder and make it accessible for anyone to use. We can expand to have more detailed coaching, step by step lessons, and even a versus mode. Also, we could integrate with sensors or smartwatches to provide even more data. Perhaps Punch Harder could even become the go-to training app for professional boxers. One day... One day... ;)

NVIDIA — Best use of Nemotron
Team: Team Pacific Boxers
GitHub
Team Members
  • Fedya Semenov

    Fedya Semenov

  • Jason Tran

    Jason Tran

jadelopezraymundo

jadelopezraymundo

  • Anna Tymoshenko

    Anna Tymoshenko