The problem
Course deadlines, requirements, and event details are spread across university pages, course sites, and portals that need a login. A student can get a plausible answer from a chatbot, but not a way to see where it came from or whether an important page was never checked.
What the demo does
- Course Desk
- The student asks a question, such as whether an assignment deadline changed, and gets an answer with numbered citations.
- Browser view
- A read-only Steel browser session sits beside the chat, so the student can watch pages open as they're read and page back through them afterwards.
- Citations
- Every quoted piece of evidence has to appear in the text captured during the same run. Clicking a citation shows the source it came from.
- Coverage receipt
- A per-source record of what happened: read, blocked by a login, or skipped. A source that couldn't be reached is reported that way, not counted as reviewed.
Live pages, fixtures, and logins
The app labels every run with the mode it actually used, because a demo can easily blur these:
- Live web browses current public U of T pages. Our acceptance question was about a carillon recital at Soldiers' Tower.
- Live fixture runs a real Steel browser session against public pages whose course content (a course called DEMO101) is explicitly fictional. This gave us a repeatable demo story.
- Local fixture is the fallback when those pages can't be reached. It uses bundled fictional data and is not live browsing, and the UI says so.
Sources like Quercus and Piazza need a login. Students can save portal credentials in a keychain; they're encrypted on the server, used only for the matching login host, and never sent to the language model. Multi-factor login (UTORMFA) can't be completed automatically, so those sources show up as blocked.
My part
I led the team and was responsible for bringing the React front end, the Fastify server, and the Steel/Playwright browser workflow together into one flow. I implemented source routing, evidence extraction, and the coverage reporting that surfaces unreachable or insufficient sources. I also added the preflight and regression checks we ran before demos.
How we checked it
The preflight script runs six checks: server health directly and through the front-end proxy, one live-web acceptance run, and three live-fixture acceptance runs. The regression checks cover source routing, browser execution, cancellation, cleanup of the browser session, and evidence integration. Both the fixture story and the live carillon question passed three runs in a row before the demo.
Limits
This is a hackathon demo, not a university information service. Acceptance covered the two scenarios above, not arbitrary courses. A cited answer can still miss a relevant rule, scanned PDFs without a text layer would need OCR, and live runs depend on OpenAI and Steel API access.