In a compelling demonstration of the Hacktoberfest Weekend Challenge, developer luxury-dev built 'Super Stylist,' a local-first web app designed to help a friend manage 4C hair growth. Rather than relying on cloud-based APIs, the application runs Googleβs Gemma 4 E2B model locally via Ollama on a laptop, ensuring that sensitive personal data, including progress photos and routine notes, never leaves the userβs device ecosystem.
Optimizing Local Inference for Latency
The developer faced significant performance hurdles, noting that the initial setup with Gemma 4 default settings resulted in response times exceeding three minutes. By disabling the modelβs 'thinking' mode for chat interactions, the developer reduced short answer times to under five seconds, achieving a writing speed of 14.4 tokens per second. This optimization was critical for usability, as the developer realized that for a daily routine app, speed often trumps deep reasoning. To balance speed with decision quality, the app employs a hybrid approach: chat responses remain fast and non-thinking, while weekly reviews can optionally enable 'thinking' mode. This allows the model to reason through complex adherence data for up to 100 seconds, providing a more nuanced plan for the following week without burdening the user with delays during quick daily interactions.
Architecting for Privacy and Control
The technical stack combines Next.js 16 and TypeScript with Dexie for IndexedDB storage on the user's phone. An ngrok tunnel connects the phone to the laptopβs Ollama instance, creating a private network that bypasses public cloud providers entirely. The developer emphasized that this architecture allows for full control over model behavior, such as switching providers or tweaking output schemas, which is often impossible with closed-source APIs. A key engineering decision involved moving logical constraints from the prompt to the application code. Instead of asking the LLM to judge adherence consistency, the app calculates completion rates itself and passes explicit rules to the model. This prevents the AI from inventing arbitrary parameters, a common issue with smaller local models, and ensures that recommendations like 'continue,' 'modify,' or 'simplify' are grounded in deterministic data rather than probabilistic guesswork.
User Feedback and Iteration
Early user testing revealed friction points that code alone couldn't solve. The friendβs initial feedback, 'Itβs very dope,' was followed by complaints about slow chat responses and confusing onboarding. The developer responded by redesigning the interface to clarify the daily kit and review process, and by implementing a sentence-by-sentence safety filter that removes medical claims or brand names before they reach the user. This iterative process highlighted the importance of testing local AI apps in real-world, non-ideal network conditions.
Key Takeaways
- Local inference with Gemma 4 E2B can achieve sub-5-second response times if 'thinking' modes are disabled for simple interactions.
- Hybrid architectures that use fast models for chat and slower, reasoning-enabled models for periodic reviews optimize the user experience.
- Moving logical rules from prompts to application code prevents small models from hallucinating evaluation criteria.
- Privacy-first designs using local storage and private tunnels eliminate API costs and data sovereignty concerns for personal projects.
The Bottom Line
This project proves that local AI isn't just for hobbyists; it's a viable architecture for privacy-sensitive applications where latency can be engineered around. By treating the LLM as one component in a deterministic system rather than the sole decision-maker, developers can build robust, cost-effective tools that users actually trust.