Session

Off the Edge: Grading Google AI Edge and Gemini Against a Practitioner's Eye

𝗔𝗜 𝗮𝗹𝗿𝗲𝗮𝗱𝘆 𝗵𝗲𝗹𝗽𝘀 𝘂𝘀 𝗹𝗲𝗮𝗿𝗻, 𝘄𝗿𝗶𝘁𝗲 𝗮𝗻𝗱 𝗱𝗲𝗰𝗶𝗱𝗲. 𝗖𝗮𝗻 𝗶𝘁 𝗵𝗲𝗹𝗽 𝘂𝘀 𝗺𝗼𝘃𝗲 𝗯𝗲𝘁𝘁𝗲𝗿?

Getting better at a physical skill takes someone who can see what you cannot see: your own body while it is moving. Where your arms are. How high you rise. Whether this second looks like the one before it.

That is a harder thing to ask of a machine than it sounds. It has to understand space rather than words (position, speed, repetition) and then say something about it that is worth acting on.

This talk sets out to find how far that has come, using an application built for the purpose and a sport where the answer can actually be checked: 𝗷𝘂𝗺𝗽𝗶𝗻𝗴 𝗿𝗼𝗽𝗲.

That sport was chosen for exactly that reason. Someone who practises it can grade the answer. A few seconds of footage are enough to say what is wrong and why, so a model offering coaching gets no benefit of the doubt.

The application runs on Google Cloud, on two technologies that read movement in completely different ways. 𝗚𝗼𝗼𝗴𝗹𝗲 𝗔𝗜 𝗘𝗱𝗴𝗲'𝘀 𝗣𝗼𝘀𝗲 𝗟𝗮𝗻𝗱𝗺𝗮𝗿𝗸𝗲𝗿 measures the body. 𝗚𝗲𝗺𝗶𝗻𝗶 𝘃𝗶𝗱𝗲𝗼
𝘂𝗻𝗱𝗲𝗿𝘀𝘁𝗮𝗻𝗱𝗶𝗻𝗴 is asked for the harder thing: not a description of the video, but advice that changes the next session.

Three questions drove the whole build, and this talk is how each one turned out.

- 𝗛𝗼𝘄 𝗳𝗮𝗿 𝗱𝗼𝗲𝘀 𝗣𝗼𝘀𝗲 𝗟𝗮𝗻𝗱𝗺𝗮𝗿𝗸𝗲𝗿 𝗴𝗼 𝗼𝗻 𝗿𝗲𝗮𝗹 𝗳𝗼𝗼𝘁𝗮𝗴𝗲? The specification says thirty-three body points, thirty times a second.
- 𝗜𝘀 𝗚𝗲𝗺𝗶𝗻𝗶 𝘃𝗶𝗱𝗲𝗼 𝘂𝗻𝗱𝗲𝗿𝘀𝘁𝗮𝗻𝗱𝗶𝗻𝗴 𝗴𝗲𝗻𝘂𝗶𝗻𝗲𝗹𝘆 𝘂𝘀𝗲𝗳𝘂𝗹 𝗮𝘀 𝗮 𝗰𝗼𝗮𝗰𝗵? Not whether it sounds convincing, every model does that now. Does it say something specific enough to change how someone trains tomorrow?
- 𝗪𝗵𝗲𝗻 𝘁𝗵𝗲 𝗺𝗲𝗮𝘀𝘂𝗿𝗲𝗺𝗲𝗻𝘁 𝗮𝗻𝗱 𝘁𝗵𝗲 𝗮𝗱𝘃𝗶𝗰𝗲 𝗱𝗶𝘀𝗮𝗴𝗿𝗲𝗲, 𝘄𝗵𝗶𝗰𝗵 𝗼𝗻𝗲 𝗱𝗼 𝘆𝗼𝘂 𝗯𝗲𝗹𝗶𝗲𝘃𝗲?

𝗜𝗻 𝘁𝗵𝗶𝘀 𝘀𝗲𝘀𝘀𝗶𝗼𝗻 𝘄𝗲 𝗴𝗼 𝘁𝗵𝗿𝗼𝘂𝗴𝗵 𝘁𝗵𝗲 𝘄𝗵𝗼𝗹𝗲 𝗿𝗼𝘂𝘁𝗲:
  1. How the application is built on Google Cloud, the choices that survived contact with real video and the ones that did not, and the point where each technology reached its limit.
  2. We will run the application live, on real sessions, and read together what it has to say about them.
  3. Then comes the coaching it produced, what practitioners of the sport made of it, and a grade for both technologies.

You will leave knowing what Pose Landmarker really delivers and where it stops, what this stack costs and demands on Google Cloud, and whether Gemini video understanding is worth putting in front of your own users, whatever it is you are helping them get better at.

𝗧𝗵𝗲 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻 𝘄𝗮𝘀 𝗻𝗲𝘃𝗲𝗿 𝘄𝗵𝗲𝘁𝗵𝗲𝗿 𝘁𝗵𝗲 𝗺𝗼𝗱𝗲𝗹𝘀 𝗮𝗿𝗲 𝗶𝗺𝗽𝗿𝗲𝘀𝘀𝗶𝘃𝗲. 𝗧𝗵𝗲𝘆 𝗮𝗿𝗲. 𝗜𝘁 𝗶𝘀 𝘄𝗵𝗲𝘁𝗵𝗲𝗿 𝘁𝗵𝗲 𝗷𝘂𝗺𝗽𝗶𝗻𝗴 𝗶𝘀 𝗯𝗲𝘁𝘁𝗲𝗿 𝗶𝗻 𝗠𝗮𝗿𝗰𝗵 𝘁𝗵𝗮𝗻 𝗶𝘁 𝘄𝗮𝘀 𝗶𝗻 𝗝𝗮𝗻𝘂𝗮𝗿𝘆.


𝗣𝗿𝗲𝗳𝗲𝗿𝗿𝗲𝗱 𝗳𝗼𝗿𝗺𝗮𝘁: 30 to 45 minute technical session. Adapts to a 15 to 20 minute lightning talk by keeping the three questions and the verdict, and dropping the architecture walkthrough.

𝗔𝘂𝗱𝗶𝗲𝗻𝗰𝗲: developers, mobile and cloud builders, and anyone about to put a multimodal feature in front of users. 𝗡𝗼 𝗽𝗿𝗲𝗿𝗲𝗾𝘂𝗶𝘀𝗶𝘁𝗲 beyond general cloud familiarity, and no computer vision background needed. The lessons apply to any domain where a model is asked to judge something it can only partly see, though the examples are on Google Cloud.

Field report, not a product pitch. Nothing shown is for sale. The application is deployed and processes real video in production, it is not a mock up built for the stage. 𝗘𝘃𝗲𝗿𝘆 𝗻𝘂𝗺𝗯𝗲𝗿 𝗾𝘂𝗼𝘁𝗲𝗱 𝘄𝗮𝘀 𝗺𝗲𝗮𝘀𝘂𝗿𝗲𝗱 𝗶𝗻 𝘁𝗵𝗲 𝗰𝗼𝗻𝘁𝗮𝗶𝗻𝗲𝗿 𝘁𝗵𝗮𝘁 𝗿𝘂𝗻𝘀 𝗶𝘁, not on a laptop, and the talk explains why that distinction cost us a wrong figure.

Demo is recorded, with a live version if the room network allows. All footage is our own, no third party athlete is filmed, and no personal data is shown.

𝗙𝗶𝗿𝘀𝘁 𝗽𝘂𝗯𝗹𝗶𝗰 𝗱𝗲𝗹𝗶𝘃𝗲𝗿𝘆. 𝗖𝗮𝗻 𝗯𝗲 𝗽𝗿𝗲𝘀𝗲𝗻𝘁𝗲𝗱 𝗶𝗻 𝗘𝗻𝗴𝗹𝗶𝘀𝗵 𝗼𝗿 𝗙𝗿𝗲𝗻𝗰𝗵, 𝘀𝘂𝗯𝗺𝗶𝘁𝘁𝗲𝗱 𝗶𝗻 𝗘𝗻𝗴𝗹𝗶𝘀𝗵.

Boris-Wilfried Nyasse

GDE, Oloodi - Founder

Montréal, Canada

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top