The visitor's situation
Someone walking through an exhibition wears the glasses and stops in front of an object. The guide should work out which exhibit they're facing, give a short introduction by text and speech, and let them ask a question without losing track of the object. When they move to the next exhibit, the guide should move with them.
Camera to server
An Android camera pipeline captures frames and hands them to a Python service that recognizes the exhibit. I integrated the two. The main decision was to prioritize fresh frames: when the visitor turns their head, the frame that matters is the latest one, not an older frame still waiting to be processed. Giving newer frames priority made recognition follow the current view more closely.
Exhibit context
Once an exhibit is recognized, it becomes the active context. Its reviewed material supplies the introduction and is also the knowledge base for follow-up answers, so a question like "when was this made?" is answered about this object and not a neighbouring one. Recognizing a different exhibit replaces the context.
Reviewed introductions vs. AI answers
The introduction and the follow-up answers come from different places, and we kept them separate on purpose.
- Introduction: prepared and reviewed ahead of time, then played as written.
- Follow-up answers: generated by a language model at the moment of the question and grounded in that exhibit's reviewed material. They can still be wrong, which is why they aren't mixed into the reviewed introduction.
I also trimmed the on-glasses text and speech so they're short enough to follow while standing in a gallery.
What acceptance covered
Three flows passed user acceptance and regression checks during the internship: exhibit recognition, switching between exhibits, and introduction playback.
That is the extent of what I can claim. I don't have recognition accuracy, latency figures, or results across different Rokid models, so this page doesn't give any.
This was internship work, so there are no screenshots or source code here. I'm happy to walk through it in conversation.