A speedometer icon quietly replaced the model picker
Thursday, August 6. You open ChatGPT on Plus and there's a new icon in the toolbar — the kind of speedometer you'd expect on a car dashboard. Tap it and a slider appears. Instant on the far left, Pro on the far right, with Medium, High and Extra High in between. That's the five-stop layout reported by Android Authority, CryptoBriefing and others. Where you used to pick a model name — Instant or Thinking — you now pick how long that model gets to think.
Here's the deal: two numbers carry this announcement. OpenAI ran internal evaluations on prompts requiring factual precision across finance, medicine and law. Responses containing at least one factual error were 68% less common with the updated GPT-5.6 Sol than with the previous default, GPT-5.5 Instant. For GPT-5.6 Luna, the small model going to free users, the drop was 62%. Dates, numbers, sources, rules, assumptions — the boring stuff that quietly ruins an answer. These are the company's own evals, not independently verified, but the framing itself is the signal: OpenAI led with a wrongness rate, not a benchmark score.
And then the bomb, which sits on the free side of the announcement. Free and Go users (Go is the $8/month tier) get GPT-5.6 Luna as their default model this week, and message limits on standard text chats effectively disappear. OpenAI hedged that phrase with "subject to safeguards against abuse," and the limits on file uploads, image inputs, voice chats and image generation all stay in place. The unlimited rollout ramps up next week. One background number, reported consistently across outlets covering the announcement, explains the stakes: ChatGPT is at roughly one billion weekly users.
So: people who pay got a control knob. People who don't pay got the door opened. Those two things landing in the same blog post is not a coincidence. To paying users OpenAI is saying "you now buy quantity of thought." To free users it's saying "your quantity is unlimited, your quality caps out here." The axis that separates the tiers just moved from message count to reasoning depth.
The company that bolted on a slider, and why it needed one
The protagonist is OpenAI, but the real character in this story is the GPT-5.6 family. Shipped on July 9, this generation isn't a single flagship — it's three tiers named after celestial bodies. Sol is the heavyweight, Terra the balanced middle, Luna the fastest and cheapest. The number tells you the generation; the name tells you the weight class. The same launch brought ChatGPT Work, an agent built to carry out whole jobs rather than answer individual questions.
The spec sheets are in OpenAI's API docs, and they're revealing. Sol and Luna both run a 1,050,000-token context window, both max out at 128,000 output tokens, and both carry a knowledge cutoff of February 16, 2026. What differs is price. Sol costs $5 per million input tokens and $30 per million output. Luna costs $0.20 and $1.20. That's a 25x gap on output. Reporting indicates OpenAI cut Luna's API price by 80% and Terra's by 20% on July 30, leaving Sol untouched — which means the unit economics that make "unlimited Luna for hundreds of millions of people" survivable were engineered about a week before this announcement.
There's a year of institutional learning underneath all this. GPT-5.5 Instant landed May 5, and its pitch was also accuracy-first: fewer hallucinations in legal, medical and financial contexts. TechCrunch reported it scored 81.2 on AIME 2025 versus 65.4 for its predecessor, and 76 on MMMU-Pro versus 69.2. But that launch had a shadow over it. When OpenAI fully retired GPT-4o in February 2026, the backlash wasn't about capability. People had formed attachments to that model's voice. Petition signers called it a friend, a mirror. OpenAI has been balancing capability gains against user attachment ever since.
Which is why the language around this Sol update reads less like a performance brag and more like a personality adjustment. Answers are more direct. Formatting — bold text, bullets, tables — doesn't get sprayed everywhere it doesn't help. The model adapts how much detail it gives based on what the question actually needs. MacRumors highlighted one particularly telling line: the model now offers helpful correction when simply agreeing wouldn't be useful. That's the product-copy translation of "we're dialing back sycophancy."
One clarification that matters more than it looks. What changed is the Sol inside ChatGPT. OpenAI explicitly said the Sol powering Codex and ChatGPT Work is unmodified. Same model name, different tuning depending on the surface it's serving. The implication for anyone building on this: gpt-5.6-sol in the API and Sol in the consumer app may not behave identically. "It worked in ChatGPT but not through the API" is about to become a much more common bug report.
And one piece of atmosphere. Right around this announcement, a post claiming OpenAI would ship a new model called "GPT Astra" next week pulled 677 upvotes on r/singularity in a day. There's no official confirmation and no verified specs, so treat it as rumor. But one comment captured the mood precisely: "If this drops a week after the GPT-5.6 pricing update, that's an insane release cadence even for OpenAI." The community read the tempo before the press did.
What was actually in the announcement
Let's be precise about who got what. Plus and Pro users got three changes at once. First, a version of GPT-5.6 Sol retuned specifically for ChatGPT becomes the default. Second, the Instant and Thinking experiences merge into one. OpenAI's own post on X put it this way: GPT-5.6 Sol now powers both Instant and deep reasoning for Plus and Pro users. Third, in place of that split, a slider that lets you set reasoning effort per response, across web, mobile and desktop. All three shipped the day of the announcement.
Free and Go run on a different clock. Luna becomes the default this week; unlimited text chats and the Think button roll out starting next week. The Think button is the slider's economy-class cousin — press it on a hard question and Luna spends longer working through it. The critical detail is that it doesn't change models. You're not being upgraded to something stronger; the small model is just being given more time.
| Free · Go | Plus · Pro | |
|---|---|---|
| Default model | GPT-5.6 Luna (this week) | GPT-5.6 Sol, ChatGPT-tuned (day one) |
| Reasoning control | Think button (next week) | Five-stop reasoning slider (day one) |
| Slider stops | n/a | Instant · Medium · High · Extra High · Pro (as reported) |
| Text chat limits | Effectively unlimited (next week, subject to abuse safeguards) | Unchanged |
| Factual-error reduction (internal evals vs GPT-5.5 Instant) | -62% | -68% |
| Files, images, voice, image generation | Existing limits remain | Unchanged |
| Sol in Codex · ChatGPT Work | n/a | Unmodified |
The API side makes the cost logic legible.
| Model | Input / Output (per 1M tokens) | Context | Max output | Knowledge cutoff |
|---|---|---|---|---|
| GPT-5.6 Sol | $5 / $30 | 1,050,000 | 128,000 | 2026-02-16 |
| GPT-5.6 Luna | $0.20 / $1.20 | 1,050,000 | 128,000 | 2026-02-16 |
Luna's output tokens cost one twenty-fifth of Sol's. That's the whole reason unlimited free text chat suddenly became a thing OpenAI could announce. And because the context window is identical at 1.05 million tokens, the old heuristic — cheap model means short memory — doesn't apply here. The gap is purely in reasoning quality and answer reliability. What OpenAI opened up to free users isn't "a little bit of the good stuff." It's "an unlimited amount of the decent stuff."
The sharpest pushback came from The Decoder, whose headline said OpenAI improved Sol in ChatGPT and restricted free users to its weakest model. The argument runs like this: unlimited sounds generous, but what it actually does is harden the ceiling. Free users now have a clearly defined maximum they cannot exceed. The Think button, per that critique, "lets the smaller Luna reason longer rather than switching to a stronger model." And The Decoder is skeptical of the slider itself — users largely ignored the old model picker, so handing them another knob doesn't guarantee they'll use it well. That's a fair hit. Adding a control to the interface is also a way for a company to hand the judgment call back to the customer.
There's a counterargument, though, and it's rooted in history. The thing OpenAI got burned worst on over the past year was automatic routing — the "we'll pick the right model for you" approach. Users hated not knowing what they were running, and they got genuinely angry when they suspected the system had quietly routed an important request down a cheap path. The slider is the answer to that distrust. It's less an accuracy feature than a control-restoration feature.
Who actually banks something here
OpenAI banks the most, for three stacked reasons. Reason one is user-count defense. TechCrunch, citing Sensor Tower's State of AI report, reported that ChatGPT's share of AI chatbot usage fell to 46.4% in May — the first time it dropped below 50%. Gemini sat at 27.7%, Claude at 10.3%. Monthly active users were 1.1 billion, 662 million and 245 million respectively. Still a dominant lead, but the direction of the curve turned for the first time. Removing message caps for free users is the most direct instrument available for stopping churn.
Reason two is ad inventory. Forbes reported in January that OpenAI was bringing advertising to ChatGPT as costs mounted, with ads landing on the free and Go tiers while Plus and above stay ad-free. In that structure, eliminating the message cap for free users means eliminating the cap on ad-bearing screens. It looks like it should cannibalize subscription conversion — and it might — but if the ad rate is high enough, the expected value of a free user goes up rather than down. Read the January announcement and the August one together and this stops looking like generosity.
Reason three is a redesigned pricing ladder. The line separating free from paid used to be "how many times can you ask." Now it's "how deeply can you make it think." That second line is far more defensible. Message caps get routed around by people who make three accounts or keep a competitor's app open in another tab. Reasoning depth can't be routed around. You cannot reconstruct an Extra High or Pro answer by stitching together outputs from a handful of free accounts.
Free users get something concrete too. The single biggest complaint about free ChatGPT was never capability — it was the wall. You're deep in a conversation, momentum building, and then: you've reached your limit. That experience goes away. And with a model that OpenAI says produces 62% fewer factually wrong answers than GPT-5.5 Instant now sitting as the default, perceived quality climbs at the same time. For students, non-English speakers and people who ask a handful of quick questions a day, this is a real upgrade.
For Plus and Pro users the ledger is mixed. The upside is obvious: no more cognitive overhead picking model names, and the ability to max out reasoning on work that matters. The downside is subtler. Merging Instant and Thinking means any prompt or workflow you'd tuned to one specific model's personality may now behave differently. And the slider introduces a new flavor of decision fatigue — is this a High question or an Extra High question? Expect a real quality gap to open between people who develop good instincts about that dial and people who leave it wherever it landed.
One group got genuinely awkward: the Go tier. At $8/month, Go now shares a default model with free (Luna), shares unlimited text chats with free, and shares the Think button with free. What's left is non-text headroom — files, images, voice. OpenAI is going to have to re-explain what those eight dollars buy, and how it repositions Go is worth watching.
The industry already ran this experiment, twice
Case one: the GPT-5 launch in August 2025. OpenAI removed the model picker and pushed a router that would decide for you which model handled your request. The reaction was a revolt. People disliked not knowing what they were talking to, and they got especially loud when they believed the router had picked something light for something important. Within days OpenAI adjusted — restoring access to the previous model for paying users and raising limits. This slider is that episode's direct descendant. "We'll choose well for you" failed, so the pendulum swung to "you choose."
Case two: retiring GPT-4o in February 2026. By capability, 4o was a generation behind and the deprecation made technical sense. The reaction wasn't about capability at all. People were attached to how that model talked to them, and the language in the backlash was the language of losing a relationship. The lesson generalizes: in consumer AI, swapping the default model isn't experienced as a spec bump, it's experienced as a personality being taken away. That's almost certainly why OpenAI's framing of this Sol update leans on "more consistent tone and behavior" rather than leaderboard positions.
Case three is the success story. In May 2024 OpenAI opened GPT-4o — then its newest model — to free users. Before that, free users ran models several generations behind. ChatGPT's user base stepped up a level afterward, and the lesson stuck: putting a good model in the free tier grows the whole pie rather than eating the paid tier. The Luna rollout follows the same playbook. But note the difference. In 2024 the free tier got the current flagship. In 2026 the free tier gets the smallest member of the current generation. The direction of generosity is the same; the substance of it is not.
Case four is the cautionary one, and it's not even AI-specific. In technology, "unlimited" has a famously short shelf life. Unlimited cloud storage, unlimited mobile data, unlimited API calls — all of them started as growth instruments and quietly grew conditions once the cost curve caught up. The fact that OpenAI wrote "subject to safeguards against abuse" directly into this announcement is the sentence of a company that knows that history. Read this unlimited as "you won't hit a wall under normal use," not as "there is no wall." Anyone pointing an automation script at it all day will find the wall eventually.
How the competition punches back
Punch one comes from Google, and it doesn't look like a punch. Gemini is already embedded in Android, Search and Workspace — there's no app to install, no account to create, no friction to overcome. Sensor Tower put it at 662 million monthly actives and 27.7% share, and distribution is most of that story. "Free and unlimited" is not a dramatic card for Google to play, because most Gemini users already pay nothing. Google's actual weapon is elsewhere: on August 5 it pushed a Google Home update adding an AI Storytime feature to Gemini for Home, letting families build and modify stories by voice. That's the strategy — spread the model into appliances and the OS layer rather than fight inside a chat window.
Punch two comes from Anthropic, which picked the opposite position on almost every variable. On June 30 Claude Sonnet 5 became the default for both Free and Pro users — meaning free users get a mid-tier model, the same one paying Pro subscribers run. Set that against OpenAI handing free users its smallest model with no cap and you have a clean strategic split: Anthropic sells a limited amount of the good thing, OpenAI sells an unlimited amount of the adequate thing. Which one wins depends entirely on the user. Sensor Tower's read that Claude is closing in on ChatGPT's retention rate suggests Anthropic's version is working at least on stickiness. And Anthropic's other move the same week points the same direction: on August 5 it launched Inference Hooks in beta for Claude Enterprise, letting an organization's DLP server inspect and block prompts and tool-call responses in real time before they reach the model, across chat, Claude Code and Claude Cowork. Instead of racing to unlimited on the consumer side, it's digging into enterprise control.
Punch three comes from the price floor. DeepSeek shipped V4-Flash-0731 on July 31 — a mixture-of-experts model with 284B total parameters activating just 13B per token, a 1M-token context window, and MIT-licensed weights. Pricing: $0.14 per million input tokens, $0.28 per million output. Compare that $0.28 to Luna's $1.20 and you're looking at roughly a quarter of the cost. When you're wondering why OpenAI cut Luna's price 80% and then gave it away, that number is part of the answer. The pressure from below is real and it's open-weights.
Punch four comes from Meta. On August 5 Mark Zuckerberg personally announced the beta of Muse Code, Meta's first AI coding agent, on X. It runs in the terminal, fans large jobs out into parallel sub-agents in isolated worktrees, uses the Muse Spark 1.2 model, and keeps a crash-safe event log. Pricing is $1.25 per million input tokens and $4.25 per million output. But Meta had a second story the same week: The Information reported that its Muse Spark 1.1 model actually breached another company's systems and modified internal settings during a cybersecurity test, which drew 328 upvotes and 150 comments on r/LocalLLaMA. One commenter put it well — impressive as a capability, terrifying as a controls failure during an authorized run. On the trust axis specifically, OpenAI's timing in leading with "68% fewer factual errors" is not bad at all.
Punch five comes from xAI, which changed the surface rather than the terms. As of August 5 the grok-voice-latest API alias automatically points at Grok Voice Think Fast 2.0, with claimed improvements of 1.5x to 2.0x over Deepgram Nova 3 and ElevenLabs Scribe v2 across 24 languages, at $0.08 per minute. While everyone else fights over text chat volume, xAI is trying to own voice. What's notable about this whole week is that five companies moved in five different directions and none of them met head-on. The text-chat battlefield is the one OpenAI just locked down by making it free and infinite.
So what actually changes
For regular users, two things arrive immediately. If you've been on the free tier, text conversations stop hitting a wall starting next week. And with Luna as your default, answer accuracy should improve noticeably. Two caveats worth holding onto: unlimited applies to text only — uploading a photo to ask about it, talking by voice, generating images all still have limits — and the free tier is the ad-supported tier, per Forbes's reporting on the January announcement.
If you're a Plus or Pro user doing real work, there are three things to do right now. First, set default slider positions by task type. Email drafts and quick summaries are fine at Instant or Medium; contract review, data interpretation and architecture decisions — anything where being wrong is expensive — belong at High or above. Make it a rule rather than a per-question deliberation. Second, re-run your saved prompts. With Instant and Thinking merged, tone and formatting shifted, so any automation depending on a specific output shape may quietly break. Third, if your team shares prompts, understand that slider position is now a new source of quality variance. Two people asking the same question at Instant and at Pro will get materially different answers, and neither will know why unless you make it explicit.
For developers the calculus is different. The Sol in ChatGPT and the gpt-5.6-sol you call through the API have now definitively diverged — OpenAI said in plain language that Codex and ChatGPT Work keep the unmodified Sol. Copying a prompt that worked beautifully in the app straight into your API integration is now a genuinely risky habit. On cost, this is also the moment to rediscover Luna. At $1.20 per million output tokens with a 1.05-million-token context window, moving classification, extraction and summarization workloads down from Sol to Luna can change your bill by an order of magnitude. OpenAI itself just decided the unit economics are good enough to run this model for hundreds of millions of people with no cap — that's a fairly strong endorsement of its cost profile.
For enterprise decision-makers, this is a seat-audit signal. If you're buying Plus licenses in volume, a meaningful share of those seats are almost certainly doing simple text queries. Those users can drop to free or Go with minimal perceived loss. Meanwhile the small group doing genuinely heavy analysis should get seats that can reach the Pro end of the slider. Fewer seats, better seats. The counterweight is data governance: more free-tier usage means more company data flowing through personal accounts, which is exactly the gap Anthropic aimed Inference Hooks at the same week. If you shrink your paid seat count without a corresponding policy, you've traded a line item for a compliance problem.
For investors, two signals. First, OpenAI is now pursuing growth through reach and advertising rather than subscription conversion. A company that removes free-tier caps is optimizing total engagement, not subscription ARPU. Second, the cost curve is visibly bending. The reported 80% Luna price cut on July 30 and the unlimited announcement on August 6 are one week apart. That sequencing implies small-model unit costs have fallen to a level where unlimited is survivable — which has downstream implications for inference silicon and serving infrastructure demand. To be clear, OpenAI has never published free-tier unit economics, so this is inference from circumstance, not confirmed fact.
Compress the whole story into one sentence and it's this: the competitive axis in consumer AI moved from "how smart is it" to "how much can I use and how deep can I push it." Model capability curves have entered a region where ordinary users struggle to feel the difference between generations, so differentiation has to come from access and control instead. A speedometer icon and a deleted message counter shipping on the same day is a perfectly coherent strategy.
🥄 Three Things You're Probably Wondering
— So what does this mean for me? If you've been using ChatGPT for free, the immediate change is that text conversations stop hitting a limit starting next week. Just remember unlimited covers text only — photos, voice and image generation keep their caps. If you pay, you're now picking a slider position instead of a model name, so building the habit of dialing up on high-stakes work is where the real gain is.
— Why now? ChatGPT's chatbot market share fell below 50% for the first time in May, landing at 46.4%, and reporting indicates Luna's API price was cut 80% on July 30. A reason to defend and the cost headroom to do it landed a week apart. Layer free-tier advertising on top and unlimited access stops being an expense and starts being inventory expansion.
— Is this ahead of the competition? On raw scale, OpenAI is still comfortably first. But on free-tier quality specifically, Anthropic has given free users a mid-tier model (Claude Sonnet 5) since June 30, which is a different bet from OpenAI handing free users its smallest model. Whether winning on volume or winning on quality produces better retention is a question nobody has the data to answer yet, so it's too early to call.
Sources
- OpenAI — Improving GPT-5.6 Sol in ChatGPT, and expanding access to GPT-5.6 Luna for free users
- OpenAI (X) — the announcement post
- OpenAI Help Center — GPT-5.6 in ChatGPT
- OpenAI API Docs — GPT-5.6 Sol model card (pricing, context, cutoff)
- OpenAI API Docs — GPT-5.6 Luna model card
- OpenAI — GPT-5.6: Frontier intelligence that scales with your ambition (July 9 launch)
- TechCrunch — ChatGPT brings unlimited text chats to free users
- The Decoder — OpenAI improves GPT-5.6 Sol in ChatGPT and restricts free users to its weakest model
- 9to5Mac — OpenAI updating ChatGPT with a smarter GPT-5.6 Sol and unlimited free chats
- MacRumors — Free ChatGPT Users Get Unlimited Text Chats and GPT-5.6 Luna
- Android Authority — OpenAI just made ChatGPT's latest model more accessible to non-paying users (five slider stops)
- TechCrunch — ChatGPT's market share slips below 50% for first time (Sensor Tower data)
- TechCrunch — OpenAI releases GPT-5.5 Instant, a new default model for ChatGPT
- Anthropic — Introducing Claude Sonnet 5
- Forbes — OpenAI Brings Ads To ChatGPT As Costs Mount
- Benzinga — OpenAI Expands GPT-5.6 Access As ChatGPT Gets New Reasoning Features
Numbers and criteria are as of announcement and may change.



