The visitor's situation

Someone walking through an exhibition wears the glasses and stops in front of an object. The guide should work out which exhibit they're facing, give a short introduction by text and speech, and let them ask a question without losing track of the object. When they move to the next exhibit, the guide should move with them.

Camera to server

An Android camera pipeline captures frames and hands them to a Python service that recognizes the exhibit. I integrated the two. The main decision was to prioritize fresh frames: when the visitor turns their head, the frame that matters is the latest one, not an older frame still waiting to be processed. Giving newer frames priority made recognition follow the current view more closely.

Exhibit context

Once an exhibit is recognized, it becomes the active context. Its reviewed material supplies the introduction and is also the knowledge base for follow-up answers, so a question like "when was this made?" is answered about this object and not a neighbouring one. Recognizing a different exhibit replaces the context.

Reviewed introductions vs. AI answers

The introduction and the follow-up answers come from different places, and we kept them separate on purpose.

I also trimmed the on-glasses text and speech so they're short enough to follow while standing in a gallery.

What acceptance covered

Three flows passed user acceptance and regression checks during the internship: exhibit recognition, switching between exhibits, and introduction playback.

That is the extent of what I can claim. I don't have recognition accuracy, latency figures, or results across different Rokid models, so this page doesn't give any.

This was internship work, so there are no screenshots or source code here. I'm happy to walk through it in conversation.

Email me about this project