GPT-6 brings Intelligent UI to ChatGPT, generating interactive interfaces and starting web-search answers 44% sooner on average
Related research and updatesSynopsis
The release introduces GPT-6 in ChatGPT with Intelligent UI, letting the model compose text, graphics, buttons, forms, charts and interactive components and interleave thinking with answering; in internal evaluations GPT-6 Extra High begins answering in the same time as GPT-5.6 Medium while scoring better overall than GPT-5.6 Extra High, GPT-6 Instant starts answering 44% sooner on average for web-search questions, and safety training strengthens resistance to multi-turn adaptive jailbreak attempts.
Interpretation
Intelligent UI lets ChatGPT choose and combine text, visuals and interactive elements per question, so responses can include graphics, tappable buttons, forms, charts and interactive experiences usable directly in the conversation. Conversational answers were previously fitted into a single text format; the model now has the freedom to build the most useful interface for each question, and can still return plain text when that is most useful. The text describes the mechanism: a library of native, streamable components plus a compiler that processes the interface as the model generates it, so the interface appears progressively without waiting for the full response; training methods were expanded to evaluate the interfaces it creates for clarity, usefulness and completeness.
GPT-6 can begin answering while it continues to think, interleaving thinking with answering and taking the user's waiting time into account to progressively compose an answer from knowledge and findings up to that point. With reasoning models, users had to wait for the model to finish thinking before getting an answer; answers can now be built across multiple partial responses while staying as cohesive and factual as one written all at once. In an internal evaluation of high-value everyday agentic tasks, GPT-6 Extra High begins answering in the same amount of time as GPT-5.6 Medium while achieving a better overall score than GPT-5.6 Extra High; for web-search questions, GPT-6 Instant starts answering 44% sooner on average.
GPT-6 improves the quality of answers that require web search, making better decisions about when to look something up and more reliably finding information that supports its answer. Relative to GPT-5.6, the model improves both retrieval timing and evidence grounding. In an internal evaluation of difficult problems, GPT-6 correctly addressed the key aspect of the user's question more often than GPT-5.6.
On safety, GPT-6 builds on several of Astra's advances to better resist attempts to bypass its safety training, particularly attacks that adapt across multiple turns, and is better at recognizing and communicating the limits of its own capabilities. Relative to GPT-5.6 Sol, the model makes progress in resisting jailbreaks, following safeguards and being clearer about what it can and cannot do, using conversation history and context to recognize risks not apparent from a single prompt. It showed stronger resistance in adversarial testing; alignment evaluations assessing honesty and responses to safety restrictions ran throughout training; more detail on results and review is available in the system card.
Perspective
The result applies to the Chat experience in ChatGPT: GPT-6 with Intelligent UI first rolls out globally in the Chat tab to Plus, Pro, Business and Enterprise tiers, expands to Free and Go the next day, and enterprise availability depends on workplace admin settings; Plus, Pro, Business and Enterprise are powered by GPT-6 Sol, while Free and Go are powered by GPT-6 Luna, both tuned for everyday conversation. The examples show the capability on planning tasks, such as a Sunday lamb roast menu that scales with headcount, shopping quantities and a timeline, plus interactive calculators, bill splitters and in-conversation games. For researchers and engineers, the value to watch is the streaming interface-generation path of a component library plus compiler, and the training idea of factoring user waiting time into answering strategy.
The key performance numbers are internal evaluations, with no external replication or public benchmark comparison, and the task composition, sample sizes and statistical basis behind the 44% earlier start and the overall-score comparison are not detailed in the text. Safety conclusions rest on adversarial testing and alignment evaluations, so the coverage and long-term robustness against multi-turn adaptive attacks remain to be seen. The text also notes there is still work ahead to improve the model's design judgment and expand what it can create. In addition, the example conversations show interface forms but give no data on interface-generation failure rates or actual user outcomes.
