Ishan Naik has released PixaFind, a Windows application that solves the problem of remembering file contents but forgetting file names. Built on a Tauri 2 shell with a Rust backend, the tool leverages ONNX Runtime for local inference and SQLite for vector storage. This architecture allows users to perform semantic search without sending data to the cloud, keeping indexing and retrieval strictly local.

Dual-Encoder Architecture for Precision

The core of PixaFind’s capability lies in its use of two distinct vision-language models: SigLIP2 and DINOv3. SigLIP2 handles text-to-image search, allowing users to query with natural language descriptions like 'cat in a sunbeam.' DINOv3 is reserved for image-to-image similarity, enabling users to find visually similar photos by selecting an existing image. Naik emphasizes that these embedding spaces are kept separate; a DINOv3 vector and a SigLIP2 text vector do not share compatible meaning, so each search route requires its specific model and preprocessing pipeline.

Managing ONNX Runtime Constraints on Windows

Running these models locally on Windows requires careful management of the ONNX Runtime. The application supports DirectML and CUDA acceleration, with a CPU fallback for broader compatibility. However, Naik notes that Microsoft’s DirectML execution provider documentation requires sequential graph execution and disabled memory-pattern optimization. This constraint forces developers to serialize access to shared sessions or provision separate sessions, which increases memory overhead. The implementation uses the ort 2.0.0-rc.10 API, reusing sessions across images to avoid loading model weights for every forward pass.

Hybrid Indexing for Documents and Images

Beyond images, PixaFind indexes documents by combining dense vector embeddings with a BM25 keyword index. This hybrid approach addresses the limitation of semantic search, which often misses exact identifiers like error codes or invoice numbers. The system supports approximately 30 document extensions, including PDFs, DOCX, and source code files. For ranking, the tool merges results from both routes, using rank-based fusion to normalize scores from the different scales of vector similarity and lexical matching.

Infrastructure Costs and Offline Limitations

While the executable is lightweight at about 5 MB, the total footprint is significantly larger due to model weights and runtime dependencies. The release README indicates a one-time model download of roughly 500 MB. Naik warns that this initial provisioning requires network access, even though subsequent searches are fully offline. For air-gapped environments, operators must manually provision the installer and model assets before disconnecting the machine. Additionally, the local SQLite vector storage means disk usage grows with the library size; 100,000 items with 768 dimensions consume approximately 293 MiB for vector payloads alone.

Key Takeaways

  • PixaFind uses SigLIP2 for text-to-image queries and DINOv3 for visual similarity, keeping the embedding spaces strictly separate.
  • The application relies on ONNX Runtime with DirectML or CUDA support, requiring strict session management to comply with provider constraints.
  • Document search combines dense embeddings with BM25 to ensure exact matches for identifiers alongside semantic results.
  • Initial setup requires a ~500 MB model download, but all indexing and searching occur locally with no telemetry.

The Bottom Line

PixaFind demonstrates that local-first AI tools are viable on Windows if developers respect the hard constraints of ONNX Runtime and hardware acceleration. It’s a solid reference for building privacy-preserving search without relying on cloud APIs.