GPT‑5: OpenAI’s Latest Model, Open Releases, and First Reactions

OpenAI launched GPT‑5 in August 2025, calling it their most advanced model yet. CEO Sam Altman described it as a “PhD-level expert in your pocket,” capable of tackling everything from code and math to health advice and image analysis.

It’s a big upgrade over previous versions. GPT‑5 introduces a new architecture, improved reliability, and better performance across the board. All ChatGPT users got access from day one – with free users seeing limits, Plus subscribers getting more usage, and Pro users gaining access to a special “GPT‑5 Pro” version designed for longer reasoning.


Key Features and Improvements

Smarter Architecture

GPT‑5 runs on a dual-model system: a fast lightweight model handles everyday questions, and a more powerful “thinking mode” model kicks in for harder ones.

An AI router decides in real-time which model to use, based on how complex the task is or what tools are needed. So it can stay quick when you’re asking for a definition – and take its time when you want a multi-step plan or deep analysis.

If you run out of high-quality responses, GPT‑5 switches to a smaller backup model so things don’t just stop working.

Bigger Context, Better Inputs

GPT‑5 supports a massive 400,000-token context window – about 300k input plus up to 128k output. That’s enough to feed in entire codebases, books, or large document sets at once.

It also handles both text and images, letting you drop in screenshots or diagrams as part of your prompt. Vision support isn’t new – GPT‑4 had it – but GPT‑5 is more accurate and better at tying it into the broader conversation.

Stronger Across the Board

On benchmarks, GPT‑5 raised the bar:

  • 74.9% on SWE-bench Verified (coding)
  • 94.6% on AIME 2025 (math)
  • 84.2% on MMMU (multimodal tasks)
  • 46.2% on HealthBench Hard subset (medical Q&A)

These aren’t just good scores – they show up in real-world use. Developers noted that GPT‑5 can build full apps or front-end layouts from a single prompt, often with solid design choices baked in.

Writers found it better at sticking to tone and style – it can carry a poem in unrhymed iambic pentameter or generate free verse that feels natural. And in medical settings, it doesn’t just spit out info – it asks questions, tailors its responses, and engages more like a helpful assistant than a search engine.

Fewer Errors, More Transparency

Hallucinations are down significantly. Compared to GPT‑4o, GPT‑5 gave 45% fewer factual errors in regular use, and about 80% fewer when using its deep thinking mode compared to the older o3 model.

It’s also less likely to make things up just to sound agreeable. OpenAI says it trained GPT‑5 to be more truthful and less sycophantic – a model that doesn’t just try to please, but tries to be right.

Safety-wise, GPT‑5 went through new evaluations and seems more resistant to jailbreaks or risky prompts. There’s also better control: you can now ask it to show its reasoning, and developers can adjust a verbosity setting to control how much detail it outputs.


OpenAI’s Open-Weight Models: GPT‑OSS

Two days before GPT‑5 launched, OpenAI dropped something unexpected: two open-weight models – gpt-oss-120b and gpt-oss-20b.

These models (branded GPT‑OSS) are licensed under Apache 2.0, meaning they’re free to use commercially and can run on your own hardware. They’re not as powerful as GPT‑5, but they’re solid.

  • gpt-oss-120b gets close to GPT‑4-level reasoning and runs on a single 80GB GPU
  • gpt-oss-20b is smaller, can run on 16GB devices, and holds up well for simpler tasks

Despite the size difference, both models perform well on chain-of-thought reasoning and tool use. They even do surprisingly well on tests like HealthBench and TauBench, sometimes beating older closed models like GPT-4o.

These weren’t trained from scratch – OpenAI distilled techniques from their flagship models, and it shows. The trade-off? Knowledge gaps in areas like pop culture or casual reasoning. Some users found the models nailed academic questions but stumbled on simpler stuff. Analysts think the training data may have been heavily filtered – or even partly synthetic – similar to Microsoft’s Phi series.

Still, this marks a big shift. OpenAI had locked down its best models for years. GPT‑OSS is a step back toward openness – or at least, something closer to it. You can’t see the training data or code, but you can download the weights and run them however you want.


GPT‑5 vs the Competition

Here’s how GPT‑5 stacks up:

GPT‑5

  • Unified dual-model system
  • 400K token context
  • State-of-the-art on benchmarks
  • Great at coding, reasoning, creative writing
  • Less hallucination, lower latency
  • Available via ChatGPT and API, with new pricing that beats GPT‑4

GPT‑OSS-120B / 20B

  • Open weights, Apache 2.0 license
  • 120B model reaches o4-mini-level performance
  • Runs locally on 80GB or 16GB hardware
  • Good for devs who want privacy or full control

Claude 4 (Opus & Sonnet)

  • Released May 2025
  • Opus 4: best-in-class coding and long agent sessions
  • Sonnet 4: smaller, faster, more responsive
  • Large context (100K+), tool support, strong safety features
  • Still competitive, though GPT‑5 now edges ahead in tough coding tasks

Other players – like Google’s Gemini 2.5 or open-source Llama variants – are active, but GPT‑5 and Claude 4 dominate most use cases for now.

One big change: GPT‑5’s launch meant the removal of GPT‑4o and earlier models from ChatGPT. This caught users off guard – and didn’t go over well.


Reactions, Praise, and Pushback

The Good

Developers generally liked what they saw. One user on Hacker News said:

“I’ve been testing it against Opus 4.1 [Claude]… and [GPT‑5] has done better… It’s definitely better, at least so far.”

On Qodo’s code review benchmark, GPT‑5’s mid-tier model topped the leaderboard. Others praised how it handles big codebases or asks clarifying questions when things aren’t clear – instead of guessing and getting it wrong.

The price drop also earned points. GPT‑5 is cheaper per token than GPT‑4, which makes high-volume use (like analyzing 200K tokens) more practical for developers.

The Bad

Everyday ChatGPT users weren’t as happy.

Common complaints:

  • GPT‑5 responses feel too short or robotic
  • Less nuanced than GPT‑4o
  • Sometimes refuses tasks GPT‑4 handled
  • No option to switch back to older models

A top Reddit comment read:

“They have completely ruined ChatGPT. It’s slower, gives short replies, and ignores instructions.”

Another user said GPT‑5 “doesn’t have the same vibe as 4o… It’s accurate, but clipped.”

And the sudden removal of GPT‑4o hit hard. “The day they killed GPT-4o, it felt like watching a friend die,” one longtime user wrote. The lack of model choice – especially for paid users – caused a lot of frustration.

The Ugly

Some critics raised bigger concerns.

Charlie Meyer’s blog post, “The GPT‑5 Launch Was Concerning”, called out:

  • GPT‑5 still makes dumb mistakes (e.g., miscounting letters in “blueberry”)
  • OpenAI hyped cosmetic features – like new chat bubble colors – instead of model capabilities
  • The company demoed internal coding tools right before bringing on Cursor’s CEO – a partner building an AI coding editor – which felt like undercutting third-party devs

These moves raised trust issues. Some developers now wonder: will OpenAI copy their product ideas next? And if the platform keeps removing features or changing behavior with no warning, can enterprises rely on it?


Final Thoughts

GPT‑5 is undeniably powerful – and for developers, it opens up new possibilities in reasoning, code generation, and multimodal tasks. The open-weight GPT‑OSS models were a surprise bonus for the AI community and give more control to users who want to host their own models.

But the way OpenAI handled the rollout – removing older models, changing response behavior, tightening control – alienated a lot of core users.

As one commenter put it:

“OpenAI did not blow me away… Rather than show sparks of AGI, the presentation showed sparks of a company that’s starting to wander aimlessly as model progress slows.”

Time will tell if these are just launch bumps – or signs of a bigger shift in how AI is evolving.

Share this post

Twitter
Facebook
LinkedIn
Reddit

Related posts

ChatGPT Live and the New Architecture of Voice AI

OpenAI has introduced GPT-Live, a new generation of voice models that now powers ChatGPT Voice. At first, this may sound like another voice-quality update. The voices have been remastered, ChatGPT should interrupt less often, and it can respond more naturally

Read More »

Node.js
Experts

Learn more at risingstack.com

Node.js Experts