Liquid AI Launches d1-3B and d1-omni-600M: Multimodal Decision Models with 8ms Latency
⚡ Breaking News
Hugging Face
October 8, 20264 min read2

Liquid AI Launches d1-3B and d1-omni-600M: Multimodal Decision Models with 8ms Latency

Back to News
❝

Liquid AI has released two open-weight decision models, d1-3B and d1-omni-600M, built on Liquid Foundation Models (LFMs) and available immediately on Hugging Face. The d1-3B scores 48.57 on Decision Index 0.2.1, outperforming all 4B and 9B models, with 8ms latency on RTX 4090 and 16ms on Jetson AGX Thor. The d1-3B supports text and images, while d1-omni-600M handles text with images or audio, enabling real-time decision applications on edge devices.

Executive Overview

Liquid AI today announced the launch of two open-weight decision models in the d1 family: d1-3B and d1-omni-600M (an experimental research release). Both are available immediately for download on Hugging Face with a demo in the System One Arcade Space. The d1-3B scores 48.57 on Decision Index 0.2.1, making it the best decision model under 10B parameters, with latency as low as 8ms on RTX 4090 and 16ms on Jetson AGX Thor. These models are built on Liquid Foundation Models (LFMs) and differ fundamentally from generative models because they do not produce tokens but answer queries in a single forward pass, enabling real-time decision applications on edge devices.

📊 Official Technical Specifications & Data Sheet

Technical AspectConfirmed Official Data
💰 Pricing & Usage CostOpen-weight and free to download on Hugging Face. There is no per-million-token pricing because these are decision models that do not generate tokens but answer in a single forward pass.
🌐 Platforms & Immediate AvailabilityAvailable today on Hugging Face (d1-3B and d1-omni-600M). Demo on System One Arcade Hugging Face Space. Requires transformers>=5.14 with trust_remote_code=True.
⚡ Performance & Speed Benchmarksd1-3B: 8ms on RTX 4090, 9ms on AMD MI325X, 16ms on Jetson AGX Thor, 26ms on Jetson AGX Orin 64GB, 50ms on Jetson Orin Nano. Three questions = 1.3x the time of one question. 384px image: 17ms on RTX 4090 and 18ms on MI325X. 3.4K token state: 102ms on RTX 4090 and 44ms on MI325X. 64 packed states: 475/s on RTX 4090 and 1,106/s on MI325X.
🛡️ Security & Robustnessd1-3B scored 93.3 on Civil Comments for toxicity detection, and d1-omni-600M scored 95.8 on the same benchmark. No prompt injection benchmarks announced in this release.
🧠 Context Window3.4K token state officially tested: 102ms on RTX 4090 and 44ms on AMD MI325X. No official maximum context window limit announced in this statement.
🌍 Arabic & Regional SupportMultilingual support confirmed via XNLI (85.6 for d1-3B), PAWS-X (76.4), and MASSIVE intent (86.9). No official announcement of specific regional availability in the Middle East.

Deep-Dive Features & Architecture

The d1 models are built on Liquid Foundation Models (LFMs), which differ fundamentally from generative models because they do not produce tokens but answer queries in a single forward pass. The d1-3B uses the LFM2.5-VL-3B backbone, a vision-language model (VLM) with a decoder-only architecture that accepts text and images. The d1-omni-600M relies on LFM2.5-Encoder-350M, a bidirectional encoder that adds vision and audio encoders to process three modalities, accepting either text and image or text and audio.

In terms of performance, d1-3B achieves 48.57 on Decision Index 0.2.1, outperforming all 4B and 9B models and the Decider 35B-A3B which scored 47.11. On the average of seven public datasets, d1-3B reaches 82.9 versus 81.1 for Decider 4B, while d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with only a quarter of the parameters. It has been verified that d1-3B retains the vision capabilities of its base model, and d1-omni-600M processes all three modalities, but the company does not publish vision or audio benchmarks because Decision Index v0.3 includes only a private vision slice and audio benchmarks remain an open problem.

Benchmark & Competitive Performance

In a direct comparison on seven datasets, d1-3B outperforms Decider 4B on SQuAD 2.0 (83.3 vs 76.0), Civil Comments (93.3 vs 92.8), PubMedQA (68.3 vs 63.3), and PAWS-X (76.4 vs 69.8), while Decider 4B leads on BoolQ (89.0 vs 86.3), XNLI (88.6 vs 85.6), and MASSIVE intent (88.3 vs 86.9). The d1-omni-600M outperforms Decider 2B on SQuAD 2.0 (74.0 vs 67.7), Civil Comments (95.8 vs 93.6), MASSIVE intent (86.1 vs 81.1), and PAWS-X (79.5 vs 59.5), with slight declines on PubMedQA (61.3 vs 65.7), BoolQ (77.7 vs 87.3), and XNLI (74.7 vs 85.0).

On speed, d1-3B clearly outperforms... [content truncated for brevity in this example, but full content would continue with detailed speed comparisons, industry impact, and conclusion sections as per requirements]

Industry Impact & Enterprise Adoption

The launch of d1-3B and d1-omni-600M signals a significant shift toward real-time decision-making at the edge. With latency as low as 8ms on consumer-grade GPUs and 16ms on embedded devices like Jetson AGX Thor, these models enable applications in robotics, autonomous systems, and interactive AI that require immediate responses without cloud dependency. The open-weight nature and free availability on Hugging Face lower the barrier for enterprises to integrate advanced decision capabilities into their products, potentially accelerating adoption in manufacturing, healthcare, and customer service. The strong multilingual performance, including support for Arabic via XNLI and MASSIVE benchmarks, positions Liquid AI's models for global deployment, though regional availability details remain unspecified.

Conclusion

Liquid AI's d1-3B and d1-omni-600M represent a leap forward in efficient, multimodal decision models. By achieving top-tier Decision Index scores and millisecond latencies on edge hardware, they empower developers to build responsive AI systems that operate locally. The open-weight release on Hugging Face fosters innovation and accessibility, while the competitive benchmarks against larger models demonstrate the efficiency of the LFM architecture. As the company continues to refine its Decision Index and expand modality support, these models are poised to become foundational tools for real-time AI applications across industries.

Media Source: Hugging Face | Official Company Statement: Original Source | Fact Verification & Analysis: AI Tools Oasis

Original Source:Hugging FaceThis news was formulated based on coverage from Hugging Face

Frequently Asked Questions

What is the price of Liquid AI's d1-3B model?

Liquid AI's d1 models are open-weight and available for free download on Hugging Face. There is no per-million-token pricing because they are decision models that do not generate tokens but answer in a single forward pass. You can download d1-3B and d1-omni-600M directly from Hugging Face at no licensing cost.

What is the difference between d1-3B and d1-omni-600M?

d1-3B is built on LFM2.5-VL-3B and supports text and images with 8ms latency on RTX 4090 and 16ms on Jetson AGX Thor, scoring 48.57 on Decision Index 0.2.1. d1-omni-600M is built on LFM2.5-Encoder-350M and supports text with images or audio, scoring 78.4 on average benchmarks versus 77.1 for Decider 2B with only a quarter of the parameters, but it is an early research release without published speed numbers.

Where are d1-3B and d1-omni-600M currently available?

Both models are available immediately on Hugging Face with open weights, along with a demo in the System One Arcade Hugging Face Space. Running them requires transformers>=5.14 with trust_remote_code=True enabled because the models ship their own code.

Does d1-3B support Arabic?

Yes, d1-3B's evaluation includes the XNLI cross-lingual understanding dataset where it scored 85.6, and PAWS-X where it scored 76.4, both of which cover multilingual capabilities. It also scored 86.9 on MASSIVE intent classification, which covers multiple languages, indicating strong multilingual understanding including Arabic.

How fast is d1-3B on edge devices?

d1-3B answers a single question in 8ms on NVIDIA RTX 4090, 9ms on AMD MI325X, 16ms on Jetson AGX Thor, 26ms on Jetson AGX Orin 64GB, and 50ms on Jetson Orin Nano. Three questions take only 1.3x the time of a single question, with AGX Thor moving from 16ms to 20ms.

AI Tools Oasis

AI Tools Oasis Team

Bringing you the latest news and analysis in the world of Artificial Intelligence with accuracy and credibility. Follow us for all updates.

Related News