May 31, 2026

Opus 4.8 + Codex Browser

Opus 4.8 matters, but the bigger story is that Codex is turning into a full super app with browser state, mobile control, Windows computer use, and agent-to-agent workflows.

This one starts with Opus 4.8, but the real story is the shift from model updates to product updates. A slightly better model is useful, but a better agent surface can change how you actually work every day.

My take is that we are entering the iPhone-update era of models. The jumps are still real, but they are harder to feel. The platform changes are where the daily workflow starts to change.

What mattered more than the headline

CategorySignalWhat I care about
Model updateOpus 4.8Useful, but hard to feel as a major leap.
Model economicsGPT 5.5 efficiencyBetter score per dollar changes heavy agent work.
Platform updateCodex browser and mobile controlMakes Codex feel like a real operating surface.
Workflow shiftAgent mini appsApps get built for agents to use, not just humans.

Opus 4.8 is good, but not a step change

Anthropic framed Opus 4.8 as a major model release with better coding, reasoning, computer use, and knowledge work. I tested it against Opus 4.7 and did not feel a huge difference in normal use.

That does not mean it is bad. It means we should be honest about what actually changes behavior. If a model release is only marginally better, the app layer starts to matter more.

GPT 5.5 looked more efficient for coding

The benchmark story I talked through was cost, time, output tokens, and score. The important takeaway was that GPT 5.5 looked like it could score higher on some long-horizon coding tasks while costing less.

Opus still feels strong for design and some knowledge-work tasks, but for deep coding workflows I am watching efficiency very closely.

Codex had the more important update

Codex added Windows computer use, mobile remote control through ChatGPT, a browser that stays signed in, better search, and the ability for one Codex thread to spin up more Codex threads.

That is the super app thesis. The app is not just where you chat. It becomes the control room for your browser, computer, tools, files, and background agents.

Vibe coding platforms are getting pressure

People like Replit and Lovable because they make hosting, auth, databases, and viewing the app easy. But a lot of that can become a prompt inside Codex or Claude if the agent has the right plugins.

That is why dedicated vibe coding tools are under pressure. Codex and Claude are not all the way there yet, but they are close enough that people are already moving work over.

Agent mini apps are the big idea

The most important concept here is agent-native apps. These are not apps you open manually first. They are apps your agent can create, update, and render inside its workspace.

That is why the in-app browser matters so much. It gives agents a visual surface where documents, dashboards, internal tools, and workflows can exist right next to the conversation.