Developer Rahul Kumar has released "Padh Ke Batao" (Read and Tell Me), an open-source tool designed to interpret dense, formal documents for elderly users who struggle with English or complex bureaucratic language. Built for the Hacktoberfest 2026 Weekend Challenge, the application runs entirely on a standard laptop without a dedicated GPU, processing photos of bank notices, hospital reports, and government circulars locally. The project addresses a critical accessibility gap: millions of households rely on family members to decode official mail, a process that often delays urgent actions like voter registration updates or medical follow-ups.
Local Vision Model Optimization
The core of the application is the Qwen3.5:4b open-weight vision model, running locally via Ollama. Kumar faced significant performance hurdles with the 3.4 GB model on an i7-13700H CPU. Initial tests showed a 147-second latency, primarily caused by the model processing 2,391 image tokens. By resizing input images to a 768-pixel width, Kumar reduced processing time to roughly 50 seconds. However, this optimization introduced a severe hallucination risk: at lower resolutions, the model invented penalties and phone numbers for harmless Hindi government announcements. The final implementation uses a pixel budget of approximately 900,000 pixels to balance speed with accuracy, ensuring dense Devanagari text is read correctly without fabricating threats.
Privacy-First Architecture
Kumar prioritized data sovereignty, ensuring that sensitive financial and medical information never leaves the user's machine. The backend uses a Fastify server listening only on localhost (127.0.0.1), sending photos directly to the local Ollama instance. The output is structured as newline-delimited JSON with strict fields for document type, summary, key details, actions, deadlines, and urgency. This local-first approach eliminates the need for cloud AI subscriptions or third-party data retention policies, making it a viable solution for users concerned about uploading personal documents to commercial APIs.
Voice Integration and UX Design
To accommodate users who prefer auditory information, the app integrates ElevenLabs for natural Hindi and English speech synthesis. The implementation is carefully scoped: only the summary, urgency, and action steps are sent to the cloud API, while the photo and detailed account information remain on the local device. The interface features high-contrast text, large buttons, and a "Listen" function that caches audio to prevent unnecessary API calls. For offline scenarios, the app falls back to the operating system's native text-to-speech engine, ensuring the tool remains functional without an internet connection.
Key Takeaways
- Local Inference is Viable on CPUs: Running a 4B vision model on consumer hardware is practical if you optimize image token counts and accept ~60-second latency.
- Resolution vs. Hallucination: Aggressive image downscaling can cause vision models to invent details in dense scripts like Devanagari; a pixel budget of ~900k is a safer compromise.
- Structured Output Prevents Drift: Using Ollama's JSON schema enforcement and strict language prompts (e.g., forcing Devanagari over Romanized Hindi) stabilizes output quality.
- Privacy by Design: Limiting cloud API usage to non-sensitive summary text allows developers to leverage high-quality TTS without compromising document privacy.
The Bottom Line
Padh Ke Batao proves that local AI can solve real-world accessibility problems without cloud dependency or expensive hardware. Itβs a masterclass in pragmatic engineering: accepting latency for privacy and tuning hyperparameters to prevent dangerous hallucinations in critical documents.