Scholia, a new AI study partner submitted for the Sanity Challenge: Path One, demonstrates how structured content models can solve the hallucination problem in educational agents. Built by developer Sumon Selim, the tool answers student queries using only specific course materialsβlecture recordings, slide decks, textbooks, and graded problem sets. Unlike generic LLM wrappers that struggle with precise citations, Scholia anchors every claim to a verifiable source, such as a specific timestamp in a YouTube lecture or a page number in an open textbook.
The Power of Deterministic Schemas
The core engineering insight is that the agentβs reliability stems from Sanityβs structured document schema rather than prompt engineering alone. The system utilizes eighteen distinct schema types, including lectures with embedded segment arrays containing startSec and endSec timestamps, and slide decks with specific slide numbers. This structure allows the backend to perform deterministic GROQ queries to find exact locators. For instance, when a student asks where they are losing marks, the system traverses a chain from submission feedback to rubric criteria to topics, and finally to specific lecture segments, without relying on the LLM to guess the connection.
Handling Knowledge Gaps and Hallucinations
Scholia distinguishes between course-specific knowledge and external facts through a strict allowlist fallback mechanism. If a query falls outside the courseβs scope, such as asking about the Python walrus operator in a pre-3.8 course, the agent searches trusted external sites like docs.python.org or MIT OpenCourseWare. These external results are visually segregated in an amber 'Outside course material' section, ensuring students never confuse external references with their actual syllabus. This approach mitigates the risk of the model inventing course content, a common failure mode in RAG systems.
Infrastructure and Deployment Nuances
The technical stack includes Next.js 16, Sanity Studio 6, and the Vercel AI SDK 7, deployed on AWS via a container Lambda behind the Lambda Web Adapter. The developer highlights specific infrastructure hurdles, such as the requirement for lambda:InvokeFunction permissions for CloudFront on Function URLs created after October 2025. Additionally, the system uses Amazon Bedrock with Nova 2 Lite for inference, chosen due to credit constraints, while handling citation verification in the UI. The frontend parses citation syntax like [lecture 6 @ 12:40] and validates them against tool outputs, flagging any unverified claims to maintain academic integrity.
Key Takeaways
- Scholia uses 18 Sanity schema types to enforce strict relationships between course content and student feedback.
- Deterministic GROQ queries handle citation locators, reducing the LLM's cognitive load and error rate.
- External knowledge is isolated in a separate UI section to prevent confusion with course-specific material.
- The project is live at scholia.mol.la, featuring three MIT OpenCourseWare courses with openly licensed textbooks.
The Bottom Line
Scholia proves that for domain-specific agents, the quality of the structured data schema matters more than the raw power of the underlying LLM. By offloading precise citation logic to deterministic queries, developers can build trustworthy educational tools even with smaller, cost-effective models.