One logged request ended after 301.086 seconds with 44,002 bytes buffered but no completion event; that shows an incomplete stream, not which network component caused it. The proposed fallback would send safe partial text, a gateway notice and a standard OpenAI style length finish signal, with no client changes or a...
Published byImages generated with GPT Image 2
Research answer
![[GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the](https://d15bya8fi85z1v.cloudfront.net/chat/VUpdr6CgAFbX1k6I54vQv7bcGJg1/thumbnails/359EA9DF5CFA8902798E/144C236394A9D956C222-medium-640.webp)
Create a landscape editorial hero image for this Studio Global article: [GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the. Article summary: <tool call {"name":"write to file","arguments":{"path":" bmad output/analysis/draft long research 300s salvage.md","content":" 长研究约300秒异常断流:零客户端修改的流内兜底方案\n\n首期改为使用标准 OpenAI 流式响应结构交付可安全保留的部分正文、网关中断提示和长度终态;废除原草案的客户端能力协商、自定. Topic tags: general web, openai, code, python, api. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with
A long-running AI research request can produce tens of thousands of bytes and still fail at the finish line. In one recorded case, the stream ended after 301.086 seconds: the gateway had read 96,925 bytes and buffered 44,002 bytes, but had not observed a completion event. The final text arrived about 144 milliseconds before the stream ended.
The proposed response is not to pretend the answer completed normally, nor to send the same large request again. Instead, the gateway would try to deliver eligible, safe partial text in the stream already in progress, add a clear interruption notice and end with a standard length signal.
That is a design proposal, not a shipped fix. The client, production behavior and tests have not been verified as part of the proposal.
The request’s configured research time limit was 1,800 seconds, not 300. The recorded failure therefore does not support the conclusion that the gateway’s own 300-second total deadline ended the request. It does show that the stream stopped without the completion event the gateway expected.
The user reports that a client then automatically retried, producing several error banners, and attributes the cutoff to a 300-second limit in Cloudflare or an ALB. Those are useful leads, but the available records do not independently verify the retry chain or identify the component that closed the connection. The reported input count of 474,096 is likewise a user-provided figure, not an independently validated token count.
In the current flow, an unexpected end without completion is treated as an error. For a tool-aware buffered response, the gateway returns before tool arbitration, so the buffered text is not delivered. If the response headers have already been sent, the error is written as an SSE data frame; it is not a new HTTP 502 response. The client may interpret that stream as a provider failure, but whether a particular client then retries must be tested directly.
The proposal would apply only to streaming requests using the academic or fact-checking research preset. It would not change ordinary presets, non-streaming responses, account-retirement rules or generation budgets.
If a qualifying stream ends unexpectedly, the gateway would:
finish_reason: "length" chunk, followed by one data: [DONE] marker.The length signal is a compatibility choice, not a precise description of every unexpected end. In the OpenAI Chat Completions definition, length means the requested maximum generation-token limit was reached 2
14. An upstream EOF does not establish that this happened. The gateway would therefore keep the real termination reason in its own records and tell the user that the response ended early.
The intended notice would describe the observed duration, rather than claim a particular proxy imposed a hard limit. It would also make clear that asking the model to “continue” starts a new request: it may repeat content, may not recover the original task and could incur additional charges. The gateway would not automatically create that follow-up request.
Partial text is not always safe to treat as ordinary prose. A response may contain a tool-call envelope or structured arguments that are cut off mid-stream. Passing that material through as text could expose it to a client-side parser; executing a complete-looking tool call would also be unsafe without evidence that the overall response completed normally.
The proposed rule is deliberately conservative:
That means the gateway cannot promise to recover every byte in the 44,002-byte buffer. The buffer size alone does not prove that it contains only research prose—or that the answer is complete. Tool selection also matters: a request requiring a particular tool must not be silently converted into a plain-text response just to avoid an error.
The fallback would be limited to a research streaming request that ends with an unexpected EOF, has non-empty safe text and meets the gateway’s output-eligibility checks. It would not run for an explicit upstream error, cancellation, timeout, resource limit, parsing ambiguity or empty result. Status messages, heartbeats and whitespace alone would not count as useful output.
There would be no special “wait until 290 seconds” rule or new 300-second local cutoff. A safe unexpected end at 20 seconds and one at 301 seconds would use the same eligibility test; a request that runs longer than 300 seconds and completes normally would retain its normal path.
If writing the fallback fails, the gateway would stop. It would not append a second error after beginning the fallback, restart generation, or switch to a background request. The proposal also does not promise that every client will stop showing an error or retrying: that depends on how the target client interprets the finish signal and on whether other client-side continuation logic applies.
The design calls for tests that cover the standard stream sequence, byte-for-byte preservation of eligible text, partial SSE events, malformed tool envelopes, forced-tool requests and failures during writes. Concurrency tests would check that progress heartbeats cannot race with the fallback, that the terminal marker appears at most once, and that cancellation releases the account lease only once.
Crucially, a zero-client-change goal must be verified using the actual, unmodified client. The acceptance checks would include whether it displays and stores the partial text and notice, whether it treats the length finish as a failure, and whether it sends another HTTP request. A successful gateway write or an HTTP 200 response is not proof that the client saved the text or stopped retrying.
For now, the proposal’s promise is narrower: make a best effort to deliver already-received text when it is safe to do so, without disguising an incomplete upstream response as a completed answer or executing an unfinished tool instruction.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
One logged request ended after 301.086 seconds with 44,002 bytes buffered but no completion event; that shows an incomplete stream, not which network component caused it.
One logged request ended after 301.086 seconds with 44,002 bytes buffered but no completion event; that shows an incomplete stream, not which network component caused it. The proposed fallback would send safe partial text, a gateway notice and a standard OpenAI style length finish signal, with no client changes or automatic resubmission.
The design would refuse to forward uncertain or incomplete tool instructions, and would need testing with the target client before claiming it prevents retries or error banners.
One logged request ended after 301.086 seconds with 44,002 bytes buffered but no completion event; that shows an incomplete stream, not which network component caused it. The proposed fallback would send safe partial text, a gateway notice and a standard OpenAI style length finish signal, with no client changes or a...
Published byImages generated with GPT Image 2
Research answer
![[GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the](https://d15bya8fi85z1v.cloudfront.net/chat/VUpdr6CgAFbX1k6I54vQv7bcGJg1/thumbnails/359EA9DF5CFA8902798E/144C236394A9D956C222-medium-640.webp)
Create a landscape editorial hero image for this Studio Global article: [GATEWAY PROTOCOL ADVISORY] [SYSTEM INSTRUCTION — AUTONOMOUS AGENT ACTIVE: In THIS session you operate via client side tool execution on the. Article summary: <tool call {"name":"write to file","arguments":{"path":" bmad output/analysis/draft long research 300s salvage.md","content":" 长研究约300秒异常断流:零客户端修改的流内兜底方案\n\n首期改为使用标准 OpenAI 流式响应结构交付可安全保留的部分正文、网关中断提示和长度终态;废除原草案的客户端能力协商、自定. Topic tags: general web, openai, code, python, api. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with
A long-running AI research request can produce tens of thousands of bytes and still fail at the finish line. In one recorded case, the stream ended after 301.086 seconds: the gateway had read 96,925 bytes and buffered 44,002 bytes, but had not observed a completion event. The final text arrived about 144 milliseconds before the stream ended.
The proposed response is not to pretend the answer completed normally, nor to send the same large request again. Instead, the gateway would try to deliver eligible, safe partial text in the stream already in progress, add a clear interruption notice and end with a standard length signal.
That is a design proposal, not a shipped fix. The client, production behavior and tests have not been verified as part of the proposal.
The request’s configured research time limit was 1,800 seconds, not 300. The recorded failure therefore does not support the conclusion that the gateway’s own 300-second total deadline ended the request. It does show that the stream stopped without the completion event the gateway expected.
The user reports that a client then automatically retried, producing several error banners, and attributes the cutoff to a 300-second limit in Cloudflare or an ALB. Those are useful leads, but the available records do not independently verify the retry chain or identify the component that closed the connection. The reported input count of 474,096 is likewise a user-provided figure, not an independently validated token count.
In the current flow, an unexpected end without completion is treated as an error. For a tool-aware buffered response, the gateway returns before tool arbitration, so the buffered text is not delivered. If the response headers have already been sent, the error is written as an SSE data frame; it is not a new HTTP 502 response. The client may interpret that stream as a provider failure, but whether a particular client then retries must be tested directly.
The proposal would apply only to streaming requests using the academic or fact-checking research preset. It would not change ordinary presets, non-streaming responses, account-retirement rules or generation budgets.
If a qualifying stream ends unexpectedly, the gateway would:
finish_reason: "length" chunk, followed by one data: [DONE] marker.The length signal is a compatibility choice, not a precise description of every unexpected end. In the OpenAI Chat Completions definition, length means the requested maximum generation-token limit was reached 2
14. An upstream EOF does not establish that this happened. The gateway would therefore keep the real termination reason in its own records and tell the user that the response ended early.
The intended notice would describe the observed duration, rather than claim a particular proxy imposed a hard limit. It would also make clear that asking the model to “continue” starts a new request: it may repeat content, may not recover the original task and could incur additional charges. The gateway would not automatically create that follow-up request.
Partial text is not always safe to treat as ordinary prose. A response may contain a tool-call envelope or structured arguments that are cut off mid-stream. Passing that material through as text could expose it to a client-side parser; executing a complete-looking tool call would also be unsafe without evidence that the overall response completed normally.
The proposed rule is deliberately conservative:
That means the gateway cannot promise to recover every byte in the 44,002-byte buffer. The buffer size alone does not prove that it contains only research prose—or that the answer is complete. Tool selection also matters: a request requiring a particular tool must not be silently converted into a plain-text response just to avoid an error.
The fallback would be limited to a research streaming request that ends with an unexpected EOF, has non-empty safe text and meets the gateway’s output-eligibility checks. It would not run for an explicit upstream error, cancellation, timeout, resource limit, parsing ambiguity or empty result. Status messages, heartbeats and whitespace alone would not count as useful output.
There would be no special “wait until 290 seconds” rule or new 300-second local cutoff. A safe unexpected end at 20 seconds and one at 301 seconds would use the same eligibility test; a request that runs longer than 300 seconds and completes normally would retain its normal path.
If writing the fallback fails, the gateway would stop. It would not append a second error after beginning the fallback, restart generation, or switch to a background request. The proposal also does not promise that every client will stop showing an error or retrying: that depends on how the target client interprets the finish signal and on whether other client-side continuation logic applies.
The design calls for tests that cover the standard stream sequence, byte-for-byte preservation of eligible text, partial SSE events, malformed tool envelopes, forced-tool requests and failures during writes. Concurrency tests would check that progress heartbeats cannot race with the fallback, that the terminal marker appears at most once, and that cancellation releases the account lease only once.
Crucially, a zero-client-change goal must be verified using the actual, unmodified client. The acceptance checks would include whether it displays and stores the partial text and notice, whether it treats the length finish as a failure, and whether it sends another HTTP request. A successful gateway write or an HTTP 200 response is not proof that the client saved the text or stopped retrying.
For now, the proposal’s promise is narrower: make a best effort to deliver already-received text when it is safe to do so, without disguising an incomplete upstream response as a completed answer or executing an unfinished tool instruction.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
One logged request ended after 301.086 seconds with 44,002 bytes buffered but no completion event; that shows an incomplete stream, not which network component caused it.
One logged request ended after 301.086 seconds with 44,002 bytes buffered but no completion event; that shows an incomplete stream, not which network component caused it. The proposed fallback would send safe partial text, a gateway notice and a standard OpenAI style length finish signal, with no client changes or automatic resubmission.
The design would refuse to forward uncertain or incomplete tool instructions, and would need testing with the target client before claiming it prevents retries or error banners.