The trade OpenAI just made on GPT-5.6 Sol is a clean window into the economics of serving frontier models at consumer scale. Demand for Sol doubled, the five-hour rolling usage cap became the loudest complaint among paying subscribers, and OpenAI's answer was not more GPUs — it was an engineering trade: shrink the context window from 372K to 272K tokens, reclaim the memory and compute that longer contexts consume, and convert it into roughly 10% more usage per session while pausing the cap for Plus, Business and Pro tiers.
The subtext is that context length — a headline spec in every model announcement — is also one of the most expensive things a provider serves. Attention costs grow with context, and most everyday sessions never approach 272K tokens, let alone 372K. Trimming the ceiling that few users hit to relieve a cap that many users hit is a rational reallocation, even if it means the spec sheet quietly got worse while the experience got better.
With ChatGPT reportedly at six million active users on Sol alone, these knobs — caps, context, routing between model tiers — are becoming the real battleground of consumer AI. For users, the lesson is to watch effective limits, not launch-day specs: what a model offered in its announcement and what it serves six months later under load are increasingly different numbers.