DeepSeek Harness v0.1.0 rc.8 19 अगस्त 2026 को जारी हुआ। इसका सबसे अहम बदलाव इमेज ट्रांसपोर्ट है: compatible vision models को इमेज सीधे भेजी जा सकती है, जबकि text only models के लिए OCR और pixel analysis से structured... इस 14 बिंदु वाले अपडेट में /goal और /plan में mixed text image input, @ मेन्यू से files और sessio...
शोध उत्तर

Create a landscape editorial hero image for this Studio Global article: What changed in DeepSeek Harness v0.1.0-rc.8, released on August 21, 2026 as its first major post-beta update after the August 13 public bet. Article summary: The release was published on August 19 UTC, although coverage around August 20–21 described it as the first major update following the August 13 public beta/open-source launch. Its main change was to make images first-cl. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
DeepSeek Harness v0.1.0-rc.8 मूल रूप से orchestration और input का बड़ा अपग्रेड है—यह इस बात का प्रमाण नहीं कि हर DeepSeek मॉडल अब इमेज समझ सकता है। रिलीज़ 19 अगस्त 2026 को प्रकाशित हुई, यानी public beta और open-source launch के छह दिन बाद; कुछ रिपोर्टों में घोषणा या चर्चा की तारीख 20–21 अगस्त बताई गई। 9
16 इसके 14 बदलावों ने इमेज को agent workflows में इस्तेमाल करने योग्य बनाया और साथ ही subagents, terminal behavior, tool execution, storage तथा compatibility में सुधार किए।
4
RC.8 DeepSeek model adapter में configurable native-image request path जोड़ता है। यदि चुना गया मॉडल या backend image input को support करता है, तो Harness attachment को अलग चीज़ मानने के बजाय इमेज सीधे मॉडल तक भेज सकता है। 2
8
यह सुविधा planning से जुड़े मुख्य commands तक पहुंचती है:
/goal अब text और image के मिश्रित input को स्वीकार करता है।/plan में भी mixed text-image input दिया जा सकता है।@ menu से local files और पिछली sessions को reference किया जा सकता है। Coding-agent workflows के लिए इसका सीधा मतलब है कि screenshot, UI mock-up, diagram या error state को अब उपयोगकर्ता के वर्णन में बदलना जरूरी नहीं। वह सीधे goal और plan का हिस्सा बन सकता है।
लेकिन यहां एक जरूरी सावधानी है: RC.8 image transport और orchestration path जोड़ता है; यह अपने-आप किसी text-only model को स्वतंत्र visual understanding नहीं देता। Native interpretation इस बात पर निर्भर करती है कि configured model या compatible backend image input समझता है या नहीं। 8
9
Claude Code और Codex को subagent system के भीतर on-demand profile bundles के रूप में install किया जा सकता है। Codex को non-interactive permission mode और कई named instances का support भी मिला है। इससे एक ही workflow में अलग-अलग configuration वाले Codex subagents चलाना संभव होता है। 2
4
Windows PTY terminals को persistent PowerShell sessions मिली हैं। Minimal preset में यह सुविधा default रूप से enabled है, इसलिए लगातार चलने वाले commands के लिए हर बार shell state दोबारा बनाने की जरूरत नहीं पड़ती। 2
इस रिलीज़ में tool calling को बेहतर किया गया है। इसमें concurrent web_search, subagent की report मिलने के बाद parent task को जगाना, और dsh web के लिए local browser को अपने-आप खोलना शामिल है। बड़े session forks और SQLite storage के performance में भी सुधार किया गया है। 2
4
RC.8 ऐसे request failures को ठीक करता है जो बहुत बड़ी single image या conversation history में जमा हो गई बहुत-सी images के कारण आते थे। Streaming response cancel होने के बाद conversation को जारी रखने या fork करने पर reply prefixes बनाए रखने की समस्या भी सुधारी गई है। Custom OpenAI-compatible gateways के साथ request format में अंतर और inference-content return न मिलने जैसी interoperability समस्याओं को भी बेहतर संभाला गया है। 2
4
इन बदलावों से यह रिलीज़ केवल image-upload feature नहीं रह जाती। Images अब planning, session history और delegated work में भाग ले सकती हैं, जबकि लंबी या tool-heavy sessions के दौरान runtime भी अधिक resilient बनता है। 2
13
यदि चुना गया मॉडल image-input support घोषित नहीं करता, तो Harness tools की मदद से इमेज को structured textual evidence में बदल सकता है। रिपोर्ट किए गए pipeline में coordinates के साथ OCR text, color-ratio statistics, चुनी हुई pixel rows की scans, image dimensions और color-mode metadata शामिल हैं। इसके बाद text model इसी evidence के आधार पर reasoning करता है। 1
2
इस स्थिति में मॉडल को मूल visual field उसी तरह नहीं मिलता जैसे किसी native vision model को मिलता है। वह measurements और निकाले गए descriptions से अनुमान लगाता है—यह उपयोगी fallback है, लेकिन इसमें जानकारी का कुछ नुकसान होता है।
Tool-layer vision उन images के लिए अधिक उपयोगी है जिनमें संरचना साफ़ तौर पर निकाली जा सकती है, जैसे:
ऐसे कामों में readable text, alignment, boundaries, colors या दूसरे measurable signals अहम होते हैं, जिन्हें OCR और targeted pixel inspection कुछ हद तक पकड़ सकते हैं।
Natural photographs, fine-grained object recognition, visual style, subtle context और complex spatial या scene relationships के मामलों में यह तरीका काफी कमजोर पड़ता है। OCR और sampled pixel statistics हर उस detail को विश्वसनीय रूप से दोबारा नहीं बना सकते जिसकी इन tasks में जरूरत होती है। 1
2
व्यावहारिक नियम सीधा है: वास्तविक image understanding के लिए native vision model इस्तेमाल करें। वहीं, text-only main model होने पर या targeted evidence extraction के लिए tool-layer vision उपयोगी है। यह text निकालने, dominant colors देखने, सरल boundaries ढूंढने या UI में pixel-level differences validate करने में मदद कर सकता है, लेकिन इसके output को native visual reasoning के बराबर नहीं मानना चाहिए। 2
3
RC.8 से पहले ही community tools text-only Harness agents को visual evidence देने के कई तरीके आजमा रहे थे। रिपोर्ट किए गए विकल्पों में dsh-vision, dsh-vision-toolkit और modlens शामिल हैं; plugin directories में local OCR और pseudo-vision bridges भी सूचीबद्ध हैं। 10
23
दो plugins इस पूरे design space को समझने में मदद करते हैं:
dsh-vision-toolkit long-screenshot OCR, visual grounding, UI reconstruction और pixel comparison जैसे workflows पर केंद्रित है। ये plugins RC.8 के native-image path की जगह नहीं लेते, बल्कि उसके साथ काम कर सकते हैं। Compatible vision model इमेज को सीधे प्राप्त कर सकता है; plugin specialized OCR, layout, pixel या external-model tooling दे सकता है; और text-only model उस तैयार evidence के आधार पर reasoning कर सकता है। इस अलगाव से developers richer native perception और अधिक targeted, inspectable tool output के बीच अपनी जरूरत के अनुसार चुनाव कर सकते हैं।
Studio Global AI
इस पृष्ठ में एक स्रोत-समर्थित उत्तर शामिल है जिसे आप Studio Global के अंदर जारी रख सकते हैं।
DeepSeek Harness v0.1.0 rc.8 19 अगस्त 2026 को जारी हुआ। इसका सबसे अहम बदलाव इमेज ट्रांसपोर्ट है: compatible vision models को इमेज सीधे भेजी जा सकती है, जबकि text only models के लिए OCR और pixel analysis से structured...
DeepSeek Harness v0.1.0 rc.8 19 अगस्त 2026 को जारी हुआ। इसका सबसे अहम बदलाव इमेज ट्रांसपोर्ट है: compatible vision models को इमेज सीधे भेजी जा सकती है, जबकि text only models के लिए OCR और pixel analysis से structured... इस 14 बिंदु वाले अपडेट में /goal और /plan में mixed text image input, @ मेन्यू से files और sessions का संदर्भ, Claude Code और Codex subagents, Codex का non interactive permission mode, persistent Windows PowerShell se...
Tool layer vision screenshots, slides, diagrams और text heavy graphics के लिए उपयोगी है, लेकिन प्राकृतिक तस्वीरों, सूक्ष्म visual context और complex scene understanding के लिए native vision का विकल्प नहीं है। [1][2]