Announcements
Where to Troubleshoot
Trick for Balancing Precision and Stability
Hi, thanks for sharing this issue.
We are currently experiencing very similar behavior with GPT-5 Chat where the agent works correctly in Copilot Studio test mode but gives 'Something went wrong' error after publishing to Teams.
Has anyone found a root cause or reliable resolution for this? It would be great if you could share any findings, Microsoft responses, workarounds, or configuration changes that helped resolve the issue.
Thanks in advance!
The fact that GPT-4.1 works consistently in both the test pane and Teams, while GPT-5 Chat works in testing but fails in Teams, is the key clue.
I would not immediately assume the prompt or agent configuration is wrong. This looks more like a model/channel-specific issue.
A few things I would check:
Create a minimal test topic/agent using GPT-5 Chat
Remove knowledge sources, tools, flows, and complex instructions temporarily and test something simple such as:
"What is 2 + 2?"
Publish that to Teams.
If GPT-5 still returns the generic "I'm sorry, I couldn't return an answer" message in Teams while the same minimal agent works in the test pane, you've isolated the issue to the published channel/model path.
Check the published version
Make sure the GPT-5 configuration was saved and the agent was republished after changing the model. Also start a new Teams conversation rather than continuing an older conversation, since conversation state can sometimes make testing misleading.
Compare the same prompt across channels
Test:
GPT-5 Chat → Copilot Studio test GPT-5 Chat → Teams GPT-4.1 → Copilot Studio test GPT-4.1 → Teams
If only GPT-5 + Teams fails, that is valuable evidence for a Microsoft support case.
GPT-5 + Teams
Check whether tools/knowledge are involved
If the minimal GPT-5 agent works in Teams, add the knowledge sources and actions back one at a time. A model may behave differently once orchestration, grounding, or tool calls are involved.
I would also capture the Conversation ID / activity details and timestamp from a failed Teams interaction if available. The generic "I'm sorry" response isn't very diagnostic by itself; the backend telemetry is much more useful.
I wouldn't downgrade permanently to GPT-4.1 based only on this test, but I would use GPT-4.1 as the control case. Since it works with the same agent and Teams channel, that makes it a particularly useful comparison.
If GPT-5 consistently fails with even a brand-new minimal agent in Teams, while GPT-4.1 succeeds, I'd raise it with Microsoft as:
GPT-5 Chat succeeds in Copilot Studio test mode but fails consistently in the Teams published channel; GPT-4.1 succeeds in both channels.
That gives Microsoft a much cleaner reproduction than simply reporting that GPT-5 returns an error.
Under review
Thank you for your reply! To ensure a great experience for everyone, your content is awaiting approval by our Community Managers. Please check back later.
Congratulations to our community stars!
Expanding mentorship, skilling, and AI innovation
These are the community rock stars!
Stay up to date on forum activity by subscribing.
Mohsin Ali 360
Valantis 253 Super User 2026 Season 2
11manish 184 Super User 2026 Season 2