Liquid AI 喺 2026 年 8 月 4 號推出 LFM2.5 2.6B,一個得 2.69B 總參數嘅開源語言模型,專門設計嚟喺裝置上直接運行 AI agent 工作流程(好似 planning、tool calling、web searching 同 multi step tasks),唔使來回雲端。 呢個模型嘅記憶體佔用好細,唔過 2.5 GB,所以喺 laptop、edge server 甚至智能手機上都行得到。官方測試話,喺 Apple M5 Max 上可以跑到 220 tokens/s,喺 AMD Ryzen AI Max+ 395 上就有 113 tokens/s,而喺手機上大約 30 tokens/s。

Create a landscape editorial hero image for this Studio Global article: What is Liquid AI's LFM2.5-2.6B model — its release date, parameter count, key design focus on agentic workflows (including planning, tool c. Article summary: Here is a comprehensive breakdown of Liquid AI's **LFM2.5-2.6B** model.. Topic tags: general, academic, documentation, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails, icons, and tiny thumbnail layouts. Make it useful as an illustrative visual, not as factua
Liquid AI 嘅 LFM2.5-2.6B 係一個好細巧、開放權重嘅語言模型,專登設計嚟畀 AI agent 工作流程直接喺你部裝置上運行——完全唔使經雲端,數據唔會離開你部機,仲可以慳返每次 infer 嘅邊際成本。呢個模型喺 2026 年 8 月 4 號推出,總共有 2.69B 參數,記憶體佔用唔過 2.5 GB,所以喺 laptop、edge server 甚至智能手機上都行得到 。
LFM2.5-2.6B 係 2026 年 8 月 4 號 推出嘅 。佢係一個 dense model,有 2.69B 總參數同 30 層:包括 22 個 double-gated short convolution blocks 同 8 個 grouped-query attention (GQA) layers,用嘅係 LFM2 hybrid 架構
。
呢個模型係專為「喺裝置上運行 agentic 工作流程」而設計嘅——即係規劃、呼叫工具、上網搜尋、執行多步驟任務——數據唔會離開你部機,而且每次 infer 都唔使額外畀錢 。佢有原生嘅 tool calling 能力,而且喺訓練嗰陣已經喺真實嘅 agent harness 入面測試過,包括 Hermes Agent、OpenClaw 同 Pi
。
LFM2.5-2.6B 運行時嘅記憶體佔用 唔過 2.5 GB 。以下係 Liquid AI 官方公布嘅推理速度:
| 硬件 | 解碼速度 |
|---|---|
| Apple M5 Max | 220 tokens/s |
| AMD Ryzen AI Max+ 395 | 113 tokens/s |
| 智能手機 | 大約 30 tokens/s |
| 單張 Nvidia H100(高並發) | 大約 15,000 tokens/s |
呢個模型支援 llama.cpp (GGUF 格式), MLX (Apple Silicon), vLLM, SGLang, ONNX 同埋標準嘅 Transformers 。
你哋可以去 Hugging Face 下載開放權重,連結係 LiquidAI/LFM2.5-2.6B,入面仲有分開嘅 GGUF、MLX 同 ONNX 量化版本嘅倉庫 。使用條款係 LFM1.0 開放權重許可證
。
Liquid AI 係 2023 年由 MIT CSAIL 分拆出嚟 嘅初創公司,由 Ramin Hasani(CEO)、Mathias Lechner(CTO)、Alexander Amini(CSO)同 Daniela Rus(MIT 教授兼 CSAIL 主任)共同創立 。公司嘅技術係從 MIT 嘅 liquid neural networks 研究演變過嚟嘅
。
官方都講明,LFM2.5-2.6B 唔建議用喺 agentic coding 或者知識密集型嘅任務上。喺編程 benchmark 上面,較大嘅模型(例如 Qwen3.5-9B)仍然有優勢。官方 Hugging Face 嘅模型卡都建議,如果要做複雜嘅 agentic coding 或者深度知識檢索,最好揀大啲嘅模型 。
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
Liquid AI 喺 2026 年 8 月 4 號推出 LFM2.5 2.6B,一個得 2.69B 總參數嘅開源語言模型,專門設計嚟喺裝置上直接運行 AI agent 工作流程(好似 planning、tool calling、web searching 同 multi step tasks),唔使來回雲端。
Liquid AI 喺 2026 年 8 月 4 號推出 LFM2.5 2.6B,一個得 2.69B 總參數嘅開源語言模型,專門設計嚟喺裝置上直接運行 AI agent 工作流程(好似 planning、tool calling、web searching 同 multi step tasks),唔使來回雲端。 呢個模型嘅記憶體佔用好細,唔過 2.5 GB,所以喺 laptop、edge server 甚至智能手機上都行得到。官方測試話,喺 Apple M5 Max 上可以跑到 220 tokens/s,喺 AMD Ryzen AI Max+ 395 上就有 113 tokens/s,而喺手機上大約 30 tokens/s。
模型喺 Hugging Face 上以開放權重形式下載(LiquidAI/LFM2.5 2.6B),支援 llama.cpp、MLX、vLLM、SGLang、ONNX 等框架。佢係 MIT CSAIL 嘅 spinoff,但官方都講明呢個模型唔建議用喺 agentic coding 或者知識密集型嘅任務上。