How do you use Sentry?
Sentry Saas (sentry.io)
Version
2.71.0 (master 8afefe8), openai 3.22.1
Steps to Reproduce
_wrap_synchronous_completions_chunk_iterator (sentry_sdk/integrations/openai.py:1001) finishes the gen_ai.chat span after its for loop, with no finally and no patched Stream.close(). The caller never reaches that code if it closes the stream early, leaves a with block after one chunk, or loses the connection. The span then has no timestamp and the transaction drops it. The Responses API iterator at :1143 and both async iterators have the same gap.
Anthropic finishes the span in a finally (#5643) and in a patched Stream.close() (#5674, #5675). #5645 closed on 2026-09-30, but the openai iterators on master still finish the span only after the loop.
Probe: a httpx.MockTransport serves five SSE chunks, and before_send_transaction records the span ops before the gen_ai split.
$ python l17_sentry_openai_stream_span_lost_on_early_exit.py
sentry_sdk 2.71.0 /var/tmp/p9bL17/sentry-python/sentry_sdk/__init__.py
openai 3.22.1
read to the end spans=['gen_ai.chat', 'subprocess', 'subprocess.communicate', 'subprocess.wait', 'subprocess.wait', 'http.client']
break, then stream.close() spans=['http.client']
with stream: first chunk only spans=['http.client']
connection drops mid-stream spans=['http.client']
The same probe against client.responses.create(stream=True):
$ python l17_sentry_openai_responses_stream_early_exit.py
openai 3.22.1
read to the end spans=['gen_ai.responses', 'subprocess', 'subprocess.communicate', 'subprocess.wait', 'subprocess.wait', 'http.client']
break, then stream.close() spans=['http.client']
Probe source
"""An OpenAI chat stream that is not read to the end never finishes its gen_ai.chat span."""
import json
import httpx
import openai
import sentry_sdk
from sentry_sdk.integrations.openai import OpenAIIntegration
print("sentry_sdk", sentry_sdk.VERSION, sentry_sdk.__file__)
print("openai", openai.__version__)
def sse(n_chunks):
body = ""
for i in range(n_chunks):
chunk = {
"id": "c1", "object": "chat.completion.chunk", "created": 0, "model": "gpt-4o",
"choices": [{"index": 0, "delta": {"content": f"tok{i} "}, "finish_reason": None}],
}
body += f"data: {json.dumps(chunk)}\n\n"
return body + "data: [DONE]\n\n"
transport = httpx.MockTransport(
lambda request: httpx.Response(200, headers={"content-type": "text/event-stream"}, text=sse(5))
)
seen = []
def before_send_transaction(event, hint):
seen.append(event)
return None
sentry_sdk.init(
dsn="https://public@example.com/1",
traces_sample_rate=1.0,
integrations=[OpenAIIntegration()],
before_send_transaction=before_send_transaction,
)
client = openai.OpenAI(api_key="x", http_client=httpx.Client(transport=transport))
def run(label, consume):
seen.clear()
with sentry_sdk.start_transaction(name=label):
stream = client.chat.completions.create(
model="gpt-4o", messages=[{"role": "user", "content": "hi"}], stream=True
)
consume(stream)
ops = [s["op"] for s in seen[0]["spans"]]
print(f"{label:30} spans={ops}")
def read_all(s):
for _ in s:
pass
def break_then_close(s):
for _ in s:
break
s.close()
def with_block_first_chunk(s):
with s:
next(iter(s))
run("read to the end", read_all)
run("break, then stream.close()", break_then_close)
run("with stream: first chunk only", with_block_first_chunk)
class BrokenBody(httpx.SyncByteStream):
def __iter__(self):
yield sse(2).replace("data: [DONE]\n\n", "").encode()
raise httpx.ReadError("connection reset")
broken = openai.OpenAI(
api_key="x",
max_retries=0,
http_client=httpx.Client(
transport=httpx.MockTransport(
lambda request: httpx.Response(
200, headers={"content-type": "text/event-stream"}, stream=BrokenBody()
)
)
),
)
seen.clear()
try:
with sentry_sdk.start_transaction(name="connection drops mid-stream"):
for _ in broken.chat.completions.create(
model="gpt-4o", messages=[{"role": "user", "content": "hi"}], stream=True
):
pass
except openai.APIConnectionError:
pass
print(f"{'connection drops mid-stream':30} spans={[s['op'] for s in seen[0]['spans']]}")
Expected Result
A finished gen_ai.chat span per call, errored on a connection error.
Actual Result
No gen_ai.chat span on any of those three paths, and no gen_ai.responses span after an early close. The http.client span is the only trace of the call.
@alexander-alderman-webb was openai meant to be covered by #5645? If not, would you take a PR that ports both parts to openai.Stream, AsyncStream and the four iterators?
How do you use Sentry?
Sentry Saas (sentry.io)
Version
2.71.0 (master
8afefe8), openai 3.22.1Steps to Reproduce
_wrap_synchronous_completions_chunk_iterator(sentry_sdk/integrations/openai.py:1001) finishes thegen_ai.chatspan after itsforloop, with nofinallyand no patchedStream.close(). The caller never reaches that code if it closes the stream early, leaves awithblock after one chunk, or loses the connection. The span then has no timestamp and the transaction drops it. The Responses API iterator at:1143and both async iterators have the same gap.Anthropic finishes the span in a
finally(#5643) and in a patchedStream.close()(#5674, #5675). #5645 closed on 2026-09-30, but the openai iterators on master still finish the span only after the loop.Probe: a
httpx.MockTransportserves five SSE chunks, andbefore_send_transactionrecords the span ops before the gen_ai split.The same probe against
client.responses.create(stream=True):Probe source
Expected Result
A finished
gen_ai.chatspan per call, errored on a connection error.Actual Result
No
gen_ai.chatspan on any of those three paths, and nogen_ai.responsesspan after an early close. Thehttp.clientspan is the only trace of the call.@alexander-alderman-webb was openai meant to be covered by #5645? If not, would you take a PR that ports both parts to
openai.Stream,AsyncStreamand the four iterators?