Liquid AI has officially released d1-3B and d1-omni-600M, two open-weight models designed for rapid decision-making rather than token generation. The flagship d1-3B scores 48.57 on the Decision Index v0.2.1, outperforming every model under 10B and matching the performance of Decider 35B-A3B, which is twelve times larger. This release marks a significant shift toward efficient, single-pass inference for edge deployments.
Architecture and Multimodal Capabilities
Unlike previous Liquid Foundation Models, the d1 family does not produce tokens. Instead, they generate answers in a single forward pass. d1-3B is trained from the decoder-only LFM2.5-VL-3B backbone, accepting text and image inputs. In contrast, d1-omni-600M is built on the bidirectional LFM2.5-Encoder-350M backbone. This experimental checkpoint integrates vision and audio encoders, allowing it to process text plus image or text plus audio combinations, scoring 15.95 on the Decision Index.
Latency Benchmarks on Edge Hardware
The models are optimized for the NVIDIA stack, delivering impressive latency metrics on constrained hardware. On a Jetson AGX Thor, d1-3B answers a single question in 16 ms, while the Jetson AGX Orin 64 GB handles the same task in 26 ms. Even the smaller Jetson Orin Nano achieves 50 ms response times, which Liquid AI claims is sufficient for real-time decisions. On desktop GPUs, the model responds in 8 ms on an NVIDIA RTX 4090 and 9 ms on an AMD MI325X.
Benchmark Performance and Training Details
In text benchmarks, d1-3B achieves a mean score of 82.9 across seven public datasets, surpassing Decider 4Bโs 81.1. The smaller d1-omni-600M reaches a mean of 78.4, outperforming Decider 2B (77.1) with only a quarter of the parameters. Liquid AI notes that simple techniques like shuffling answer options and fixing data shortcuts yielded better results than advanced methods. The models are available on Hugging Face with day-one support for llama.cpp, ensuring compatibility across Apple, AMD, Qualcomm, and NVIDIA platforms.
Key Takeaways
- d1-3B matches the decision quality of models 12x its size (Decider 35B-A3B).
- Sub-50ms inference is achievable on Jetson Orin Nano, enabling real-time edge AI.
- d1-omni-600M introduces multimodal decision-making for text, image, and audio.
- Models are open-weight and support llama.cpp for local deployment.
The Bottom Line
Liquid AI is proving that you don't need massive parameter counts to get high-quality decisions, just the right architecture and training data.