The evidence falls into three buckets:
exchange-core, made more than 1,000 tool calls and changed more than 4,000 lines of code. Kimi K2.6 is not being presented as just another chatbot. Microsoft Foundry places it in the category of agentic, multimodal models and says it is designed for long-horizon reasoning, coding and autonomous execution.
SiliconFlow describes Kimi K2.6 as an open-source multimodal model focused on long-horizon coding, autonomous agent orchestration and coding-driven design. It also lists benchmark figures such as 58.6 on SWE-Bench Pro and 86.3 on BrowseComp Agent Swarm. Ollama’s model page similarly describes Kimi K2.6 as an open-source, native multimodal agentic model for long-horizon coding, coding-driven design, proactive autonomous execution and swarm-based task orchestration.
That is enough to support a cautious statement: Kimi K2.6 is positioned as a long-horizon coding-agent model. But positioning and benchmark tables do not, by themselves, prove that it can reliably operate for hours in arbitrary real-world repositories without human oversight.
One of the clearest public references is the Kimi Forum announcement. In its long-horizon coding section, it refers to 4,000+ tool calls, more than 12 hours of continuous execution and generalization across languages such as Rust, Go and Python.
The more specific 13-hour version comes mainly from articles and posts summarizing Moonshot’s release material. The DEV Community article says Moonshot reported that Kimi K2.6 spent 13 hours rewriting parts of exchange-core, an open-source matching engine, made more than 1,000 tool calls, modified more than 4,000 lines of code and produced throughput gains without human intervention. The Neuron also describes a 13-hour exchange-core overhaul with more than 1,000 tool calls. A Kimi_Moonshot post on X mentions a 13-hour execution, 12 optimization strategies and more than 1,000 tool calls.
So the most accurate reading is: the 13-hour case was publicly claimed and repeatedly described, but it has not been made fully reconstructable for outside readers.
To turn a launch example into a robust capability claim, the public record would need to answer questions such as:
The sources currently provide headline figures and narrative details: continuous runtime, tool-call counts, lines changed and the exchange-core example. Those details make the claim traceable, but they are not enough to establish reliability, generality or unattended production readiness.
Even if Kimi K2.6 is better at planning and tool use than earlier models, a long-running coding agent is still a system-engineering problem. VentureBeat, discussing Kimi K2.6 and long-running agents, notes that many orchestration frameworks were designed for agents that run for seconds or minutes, and that longer-running agents expose limits in enterprise orchestration and stateful agent management.
That matters because a 13-hour coding run depends on more than the model. It also depends on the agent framework, tool interfaces, state management, error recovery, test loop, monitoring and execution environment.
Availability is expanding: Cloudflare’s changelog says Moonshot AI Kimi K2.6 is available on Workers AI, while Microsoft Foundry, SiliconFlow and Ollama also provide K2.6-related pages or access points. But being available on developer platforms is not the same as independent verification of a 13-hour autonomous coding capability.
Reasonable statements include:
exchange-core, with public summaries describing a 13-hour run, more than 1,000 tool calls and more than 4,000 lines of code changed. Claims to avoid, unless stronger evidence is published, include:
Kimi K2.6’s “13-hour coding” story should not be dismissed as made up. Public sources do point to a 12–13 hour long-horizon coding case, and K2.6 is clearly being framed as a model for agentic coding and autonomous execution.
But the stronger claim — that Kimi K2.6 has been independently shown to perform stable, unattended 13-hour software development in ordinary real-world projects — is not established by the available evidence. The safest conclusion is: Kimi K2.6 is genuinely being aimed at long-horizon coding agents, but the 13-hour figure should be treated as a reported showcase, not a verified productivity guarantee.