The most useful AI news this week is less about who tops the leaderboards and more about price, speed and how people actually work with these tools. Anthropic released Claude Haiku 5.5, a small model built for cheap, high-volume jobs. OpenAI started a 28-day run of daily releases and opened GPT-6 with Intelligent UI to everyone. xAI turned Grok Bot into an assistant that can hand work to other companies’ models. Here is what changed, what it costs and where each release is worth trying.

Claude Haiku 5.5 trades peak intelligence for low cost

Anthropic describes Claude Haiku 5.5 as its fastest and most cost-effective small model. It is aimed at repetitive, high-volume work such as summaries, data compaction, database queries and classification, not at state-of-the-art reasoning. For software engineering, Anthropic recommends running it as a sub-agent, a helper that takes small tasks handed down by a larger model, alongside Sonnet 5.5 or Opus 5.5.

On benchmarks, Haiku 5.5 does not catch the bigger models, but it is a large step up from the previous Haiku.

Benchmark Haiku 5.5 Sonnet 5.5 Haiku 4.5
GDPval-AA v2.1, knowledge work 1620 1840 735
AA-Briefcase v1.1, knowledge work 1578 1824 614
OSWorld 2.1, computer use 72.4% 83.9% 15.7%

OpenAI’s GPT-6 Luna scored 1437 on GDPval-AA v2.1. Haiku 5.5 also posted 45.9% on Humanity’s Last Exam without tools and 57.4% with tools, 39.2% on the agentic coding test Terminal-Bench 4.0, and 46.4% on both FrontierCode 1.1 and the visual reasoning test Chartography.

The more interesting result is accuracy per dollar. Claude lets you set effort, how long the model thinks before answering, from Low to Max. On OSWorld 2.1, Haiku 5.5 at extra-high effort reached 67.6% accuracy at $0.28 per attempt. Sonnet 5.5 at medium effort reached 66.0% at $0.93 per attempt. On that test, letting the small model think longer matched the larger model for less than a third of the cost.

Price per million tokens Haiku 5.5, prompts up to 100K tokens Haiku 5.5, above 100K Haiku 4.5 Sonnet 5.5
Input $0.10 $0.50 $1.00 $2.00
Output $0.50 $2.50 $5.00 $10.00

Tokens are the small chunks of text a model reads and writes, and they are the unit of pricing. Cache reads cost $0.01 per million tokens ($0.05 above 100K), and cache writes cost $0.125 ($0.625 above 100K). On the independent leaderboards from Artificial Analysis, Haiku 5.5 ranks below Kimi k3, GLM 5.3 and Grok 4.7 on the intelligence index, but sits near the bottom on cost per task at about $0.21. It is wordy: only Sonnet 5.5 produces more reasoning and output tokens per task. The low per-token price keeps the total cheap anyway.

Where Haiku 5.5 pays off

People on a paid Claude plan of $20 a month or more will probably keep using Sonnet or Opus in the chat window. Haiku 5.5 makes the most sense where token savings add up: API automations, Claude Code and agent tooling. The setup Anthropic suggests is a larger model doing the planning and judgment while Haiku 5.5 handles quick subtasks such as code lookups or formatting.

A large robot at a planning table handing small task cards to a team of fast small robots that sort and file them

▲ A large model delegating to small sub-agents

Claude adds Dashboards and Motion

Anthropic also announced two tools for turning business data into visuals. Claude Dashboards connects directly to data warehouses and CRM platforms, including Google BigQuery, Databricks, Snowflake and Salesforce, and builds live dashboards from them. It is in beta on paid plans. At the same time, Claude Docs, Slides and Design left beta and are now available on every plan, including Free.

Claude Motion turns prompts and ideas into animated explainers and motion graphics inside Claude. One example turned bike-share departures between two stations into a 38-second animation. For now, Motion is limited to Team and Enterprise plans. Individual Pro subscribers ($17 to $20 a month) and Max subscribers ($100 to $200 a month) cannot use it, which means a $20 Team seat gets a feature that a $200 Max plan, sold partly on early access, does not. That looks like an odd choice for individual power users. Similar animations could already be built with JavaScript or frameworks such as Remotion, but rendering them inside Claude could shorten the process considerably.

OpenAI’s 28-day release run

OpenAI said it would ship a product improvement or a reset every day for 28 days, starting October 4. The first week already brought several changes that users will notice.

GPT-6 and Intelligent UI

On October 7, OpenAI made GPT-6 and Intelligent UI available to everyone, along with a redesigned ChatGPT interface. Intelligent UI replaces walls of text with graphics, tappable buttons, input forms, interactive charts and other controls inside the reply.

  • Bike assembly: an exploded diagram of a 7-speed bicycle that lets you isolate the frame, wheels, drivetrain, brakes or cockpit
  • Dinner planning: a Sunday lamb roast with a meat-weight calculator per guest, ingredient cards and a prep schedule
  • Travel: a seven-day Italian wedding trip with outfits by day, image inspiration cards and an interactive packing checklist
  • Small tools: a bill splitter for a party of five, built inside the chat

With thinking effort set to Instant, a question about how a four-stroke engine works comes back almost immediately as illustrated steps for intake, compression, power and exhaust. An internal OpenAI agent evaluation shows GPT-6 reaching higher accuracy while cutting response time from 215 seconds to 109 seconds.

Speed, approvals and meeting notes

  • October 5: default generation speed for GPT-6 Astra and GPT-6.1 Sol rose by about 50% across subscriptions and partner APIs.
  • October 6: Auto-review arrived free for every signed-in account in ChatGPT Work and Codex. The permission menu now offers three choices: Ask for approval, Approve for me and Full access. With Approve for me, a second agent watches what the working agent does and only asks you when an action looks high-risk or drifts from your intent. The goal is to end the fatigue of approving every step in a long task.
  • Meetings plugin: in beta for Pro and Business users of the ChatGPT desktop app on macOS. It captures system audio and the microphone during Zoom, Google Meet or Slack calls, writes summaries and next steps, and keeps the notes in ChatGPT’s memory so you can draft follow-up emails or update plans later.
  • Developer platform: the Decisions API, which returns fast structured judgments instead of long text, entered public beta. OpenAI also cut its five rate-limit tiers to three: Build, Launch and Grow. Accounts reach Grow after $500 in lifetime API spend and get up to 40 million tokens per minute on the frontier models and 180 million on Luna.

Ultrafast and Instant Steering

On October 8, OpenAI added an Ultrafast mode for GPT-6.1 Sol in Codex, ChatGPT Work and the API. It generates up to eight times faster than standard Sol, at a steep price.

GPT-6.1 Sol speed Speed vs. standard Input per million tokens Output per million tokens
Standard 1x $2 $10
Fast 2x $4 $20
Ultrafast up to 8x $12 $60

Ultrafast costs six times the standard rate, and using it in ChatGPT requires the $500-a-month Pro plan. Unless latency is the real bottleneck, as in live user-facing apps, interactive coding or agents that control a desktop, the premium appears hard to justify.

Instant Steering in ChatGPT Work lets you change direction while a response is still being generated. Ask for a picture of a dog, queue a change to a wolf, press Steer, and the job pivots right away instead of finishing the first request. Correcting course the moment output starts to wander saves both waiting time and compute.

A chat window opening into a bicycle exploded view, a recipe card, a checklist and a slider calculator

▲ ChatGPT answers built from visual interface elements

Grok Bot starts routing work to other companies’ models

Elon Musk announced that Grok Bot will send each task to the best available model, including outside ones such as Claude Opus 5.5, Midjourney and Suno. Instead of staying inside its own models, xAI’s assistant will call a third-party model whenever it is likely to do better. With Anthropic’s models currently leading frontier benchmarks, access to Claude Opus 5.5 looks like a real advantage.

Grok Bot can now also search, read and monitor posts on X. It can track rising themes in a niche, spot posts that are gaining traction and suggest a posting schedule based on engagement. Its design differs from single-thread assistants such as Dot or Muse. It uses specialist agents, such as Email Triage, which connects to Gmail and drafts replies, and Research Bot, which follows industry headlines and new papers. A Chief of Staff agent takes high-level instructions and delegates them to the right specialist. In one example, it analyzed an account’s last 75 posts to find the hours with the strongest engagement and drafted new post ideas. Splitting multi-step admin work across dedicated agents, rather than cramming it into one chat, appears to be the more flexible approach.

Other releases worth knowing

  • Mistral Large 4: an open-weight model, meaning its trained weights are released for others to run, built as a mixture of experts that activates only part of the network for each token. It has 1 trillion total parameters, the learned values inside a model, with 52 billion active per token. Mistral AI says it performs on par with leading open-weight models from Chinese labs.
  • Reflection Beam: an open-weight model with 501 billion total and 23 billion active parameters, pretrained on 25.8 trillion tokens according to Reflection. It is not downloadable yet. Both models are meant for companies running their own servers, not personal computers.
  • Google Playground: an experimental web app that builds playable 2D and 3D games from a text description. It is open to users in the United States who are 18 or older.
  • Google AI Edge Foresight: a macOS app that transcribes and summarizes meetings with on-device models, so audio and notes never leave the computer.
  • Gemini Agent: a Google Cloud enterprise agent that keeps long jobs running for hours or days without losing context and connects to Google Workspace, Microsoft 365, Slack, Salesforce, Jira and GitHub.
  • SynthID Detector: a public tool that checks images, video and audio for hidden digital watermarks. It can only confirm watermarked content from supported providers, not every piece of AI-generated media.
  • Hark Pro: a computer-use agent from Figure founder Brett Adcock that reorders household goods, submits expense receipts and cancels subscriptions by operating the screen directly.
  • Meta Muse: Meta open-sourced hardware designs and code for running Muse on a Raspberry Pi or ESP32 board.
  • Anthropic usage policy: it now bans sustained, needless cruelty toward Claude with no research purpose. Ordinary frustration, red-teaming, which means deliberately probing for weaknesses, and model testing are not the target.
  • Jobs outlook: Nobel laureate economist Daron Acemoglu estimates that AI will replace only about 5% of human work over the next 10 years, in an article shared by Mustafa Suleyman of Microsoft AI. That is far more cautious than common industry predictions of near-total automation.

What to try this week

The common thread is doing the same work cheaper, faster and with fewer interruptions. Computer-use agents are also multiplying across companies, and as their features converge, the choice may come down to which ecosystem you already use.

  • If you run repetitive jobs through the API or Claude Code, move subtasks such as classification, summaries and lookups to Claude Haiku 5.5 and compare the bill.
  • In ChatGPT, ask questions that benefit from visuals, such as calculations, schedules or how something works, with Instant effort and see what Intelligent UI builds.
  • If you hand long tasks to ChatGPT Work or Codex, try Approve for me in Auto-review, knowing that high-risk actions still come back to you.
  • Consider Ultrafast only where latency truly matters, and weigh the sixfold price and the $500 monthly plan first.
  • When choosing a meeting-notes tool, decide by where your data lives: the Meetings plugin keeps notes in ChatGPT, while Google AI Edge Foresight processes everything on your Mac.