Touchwood, a submission for the Hacktoberfest Open-Source AI Challenge, introduces a novel approach to digital wellness by coupling browser-based content locks with on-device computer vision. The system, developed by lillygrace-l, blocks access to social media and news feeds via a Chrome Manifest V3 extension but unlocks them only after the user completes a physical quest verified by the CLIP ViT-B/32 vision model running locally on their smartphone.

The Mechanics of Local Verification

The application operates as a Progressive Web App (PWA) that generates quests based on OpenStreetMap data, directing users to nearby features like lakes, parks, or public art. Upon arrival, the user takes a photo, which is analyzed by the on-device CLIP model to confirm two conditions: that the user is outdoors and that the subject matches the quest target. A successful verification generates a one-time 6-digit code, derived from an HMAC-SHA256 hash of a shared secret and a 5-minute time window, which unlocks the browser extension for 30 minutes.

Why Open-Weight Models Matter Here

The developer explicitly chose open-weight models to avoid the privacy and latency pitfalls of closed APIs. By running CLIP via transformers.js and using WebLLM for optional narrative generation with models like Gemma 2 2B or Qwen2.5 0.5B, the system ensures that geotagged photos never leave the device. This architecture allows for offline functionality after the initial 155 MB model download, a critical feature for users who may lose signal during their outdoor walks.

Debugging Vision Models Without Black Boxes

The project highlights the tangible benefits of inspectable models during development. The creator encountered a false positive where a beach towel was misclassified as water due to label overlap in the softmax distribution. By accessing the raw logits and implementing a margin-based check against decoy labels, the developer corrected the threshold without waiting for an API vendor’s update. This level of control is impossible with closed-source endpoints that return simple boolean results.

Key Takeaways

  • Touchwood uses CLIP ViT-B/32 for zero-shot image classification to verify outdoor quests.
  • The system requires no backend server, using HMAC-SHA256 for offline code validation.
  • Open-weight models enable on-device debugging of classification thresholds and false positives.
  • The app supports optional local LLMs like Gemma 2 2B for generating quest narratives via WebGPU.

The Bottom Line

Touchwood proves that on-device open-weight models are no longer just experimental toys; they are practical tools for building privacy-first applications with deterministic control over user experience.