Back to Insights
    AI Tools
    9 min read

    GPT-6 Astra Costs 2.5x More. Here Is What Your AI Bill Actually Does

    By Andrew Mudd·

    OpenAI raised frontier pricing 2.5x this week while Anthropic cut cache costs 75%. Here is the Blended Rate Audit, a 30-minute exercise that shows why most service businesses can absorb the hike and still spend less.

    OpenAI shipped GPT-6 Astra on September 3 at $10 per million input tokens and $50 per million output tokens. That is 2.5 times the price of GPT-5.6 Sol, the model it replaces. Two days earlier, Anthropic shipped Claude Fable 5.1 at the same token price as before, with cache reads cut by 75%.

    So in one week, the two most used frontier models moved in opposite directions on price. Here is the answer up front: for a service business, neither headline changes your bill much, because your bill was never mostly about the frontier model. It is about how much of your work you are routing to the expensive tier when a cheap one would do.

    The practical move is a 30-minute exercise I call the Blended Rate Audit. You figure out what you actually pay per task, not per token, and you route work by task value. Most operators who run it find they can absorb a 2.5x price hike on the top tier and still spend less overall. Here is how, with the real numbers.

    What did GPT-6 Astra and Claude Fable 5.1 actually change on price?

    OpenAI raised the frontier price 2.5x, Anthropic held it flat and made repeated context far cheaper, and both moves push you toward the same habit: stop sending everything to the top model.

    GPT-6 Astra lists at $10 per million input tokens and $50 per million output. Cached input drops to $1. Batch mode runs at half price. Fast mode runs at double. It ships with a 1,050,000 token context window and a 128,000 token maximum output.

    The capability jump is real. It scores 72.6% on OSWorld 2.0, the computer-use benchmark, and finishes those tasks in roughly 47% less time than GPT-5.6 Sol. It saturates FrontierMath Tier 4 at 97.6%. For long, multi-step agentic jobs, it is the strongest thing you can rent today.

    Claude Fable 5.1 took a different route. Same token prices as Opus 5. Cache reads dropped 75%, which Anthropic says lowers effective cost by 25% to 45% on typical workloads. On AutomationBench, which measures business workflow completion, it scores 31.4%, up from 26.9% for Opus 5.

    Google also shipped Gemini 3.8 Flash on September 2, and Meta pushed Muse Spark 1.3. So there were four frontier moves in 48 hours.

    Here is the part that matters for you. Frontier pricing now spans from about $1 per million tokens on cached input up to $50 per million on Astra output. That is a 50x range inside a single week's releases. The model you pick for a task now matters more to your bill than which lab you subscribe to.

    Does a 2.5x price increase actually hit a service business?

    For most operators, no, because 80% of your AI work does not need a frontier model, and the 20% that does is worth far more than the token cost.

    Let me put numbers on it. A typical client proposal runs about 3,000 words of output, roughly 4,000 tokens, plus maybe 10,000 tokens of input context: your notes, the discovery call summary, your pricing sheet.

    On GPT-6 Astra that is $0.10 for input and $0.20 for output. Thirty cents. For a proposal that closes a $12,000 engagement. Even at 2.5x the old price, the token cost is noise against the outcome.

    Now flip to volume work. Say your intake automation classifies 1,500 inbound leads a month, each about 800 tokens in and 100 tokens out. On Astra that is $12 per month for input and $7.50 for output, call it $20. Not scary either. But route the same job to a mid-tier model at $1 per million and it costs under $2.

    Alone, neither example hurts. The problem shows up when you multiply. Operators who set up their stack in 2025 often pointed everything at one model because it was simpler. Call summaries, tagging, draft replies, weekly reports, all through the same frontier endpoint. That is the setup a 2.5x increase punishes.

    The evidence that this pattern is common is everywhere. The average company now runs 12 AI agents, and half of them operate alone with no coordination, according to Belitsoft's 2026 report. Uncoordinated agents almost always share one thing: whoever built them picked one model and never revisited it.

    Gartner expects AI agent software spending to jump from about $86 billion in 2025 to about $206 billion in 2026, a 139% increase. The practitioners on r/AI_Agents have a blunt take on that number. A lot of it is spent on agents that should have been simple automations, running on models that should have been cheaper.

    What is the Blended Rate Audit?

    The Blended Rate Audit is a one-page exercise that converts your AI spend from cost per token to cost per task, then sorts every task into one of three tiers.

    Step one, list every recurring AI task in your business. Not the tools, the tasks. Call summary. Lead classification. Proposal draft. Weekly client report. Support reply. Contract review prep. Most service businesses land between 8 and 15.

    Step two, for each task estimate monthly volume and rough tokens per run. You do not need precision. Order of magnitude is fine. A call summary is about 6,000 tokens in and 400 out. A support reply is about 1,500 in and 200 out.

    Step three, multiply by the price of the model currently running it. Now you have a real monthly cost per task. Add them up and divide by total runs. That number is your blended rate, your actual average cost per AI task across the business.

    Most operators land somewhere between $0.02 and $0.08 per task. If yours is above $0.10, you are almost certainly running volume work on a frontier model.

    Step four, sort each task into a tier. Tier one is judgment work, where a mistake costs real money: proposals, pricing analysis, anything a client reads that represents your expertise. Tier two is production work, where quality matters but the stakes are moderate: call summaries, first-draft emails, internal reports. Tier three is volume work: classification, tagging, extraction, routing.

    Tier one goes to the frontier. Tier two goes to a mid-tier model. Tier three goes to the cheapest thing that passes your accuracy check.

    Where should each tier of work actually run right now?

    Frontier for judgment, mid-tier for production, budget for volume, and the whole stack should cost a ten-person shop roughly $300 to $700 per month.

    Tier one, judgment work, is where GPT-6 Astra and Claude Fable 5.1 earn their price. A solo consultant might run 20 proposals and 10 pricing analyses a month through the frontier. At Astra rates that is under $10 in tokens. Through a subscription like ChatGPT Pro or Claude Max at roughly $100 to $200 per month, it is one seat, not a team plan.

    Fable 5.1 has a specific edge here for anyone who reuses context. If your proposals all pull from the same 40-page services doc and case study library, that content sits in cache. At a 75% cache-read discount, the tenth proposal in a month costs a fraction of the first. Astra's $1 cached input works similarly, so both labs are rewarding the same behavior: build a stable context and reuse it.

    Tier two, production work, is where most of your token volume lives. Call summaries, draft follow-ups, internal reports. Mid-tier models from every major lab now handle these tasks well enough that a client cannot tell the difference. Pricing in this tier runs roughly $1 to $4 per million input tokens. A business running 300 call summaries a month spends about $5 to $10 here.

    Tier three, volume work, is where routing saves real money. Classification and extraction do not need reasoning. Gemini 3.8 Flash, GPT mini variants, and open-weight models running through a cheap host all do this for under $0.50 per million input tokens. If you run automations through Make.com at about $9 to $30 per month, or n8n self-hosted for close to nothing, you pick the model per step. There is no reason for a tagging step to call Astra.

    Put together, a ten-person service business with a full stack lands around $300 to $700 per month all in, including subscriptions. The frontier tier is usually less than a third of that, even after this week's price increase.

    What is the biggest mistake operators make with a new frontier model?

    They upgrade the endpoint everything already runs on, which turns a price increase on 20% of their work into a price increase on 100% of it.

    MIT research found that 95% of enterprise generative AI pilots failed to deliver measurable ROI, despite $30 to $40 billion in investment. The failures were rarely about model quality. They were about integration, sloppy data, and nobody owning the process after launch.

    Cost sprawl is the small-business version of the same failure. Nobody decided to spend $400 a month running lead tagging on a frontier model. It happened because the automation was built in an afternoon, it worked, and nobody looked at the invoice line by line.

    A new release like Astra is when that gets expensive. If your Make.com scenarios or your GoHighLevel workflows point at a generic "latest model" setting, a 2.5x price change hits every step at once. Pin your model versions. Route by tier. Then upgrade tier one deliberately and leave the rest alone.

    The other mistake is the reverse: refusing to pay for the frontier at all. If a $50 per million output model writes a proposal that closes one extra $8,000 engagement a quarter, the token cost is irrelevant. Being cheap on judgment work is the most expensive decision in the audit.

    How do you keep your AI bill stable when prices keep moving?

    Run the Blended Rate Audit once a quarter, pin your model versions, and let your automation platform's per-step routing do the rest.

    Prices will keep moving in both directions. This week proved it. One lab went up 2.5x while another cut effective cost by up to 45%. Next quarter it will happen again, and Nvidia's server builders have already signaled that hardware prices rise more than 15% in early 2027, which pushes costs downstream eventually.

    You cannot control that. You can control your routing. When every task is tagged with a tier and pinned to a model, a price change is a 15-minute re-route on a Tuesday, not a panic.

    The audit also gives you something most operators lack: a defensible number. When a client asks what AI costs you, or a partner asks whether you should upgrade, you can answer with a blended rate and a task list instead of a shrug.

    The point of AI in a service business is that your people get to spend more of their time on the work clients pay for. Cheap volume models buy back hours. Frontier models sharpen the work that wins engagements. Neither replaces the person making the call. They just make that person's hours worth more.

    GPT-6 Astra and Claude Fable 5.1 are both worth having in your stack, in the right tier, for the right tasks. Figure out your blended rate this week. The operators who know that number are the ones who never notice a price hike.


    If you want a clear picture of what AI can actually do for your specific operation, book a free AI Clarity Call. Thirty minutes, no pitch, you leave with a real answer.

    If you want to learn alongside other operators and stay current on what is working, join the Abra AI community. That is where I share what I am actually building.

    Subscribe to the newsletter for more breakdowns like this.

    Ready to see what AI can actually do for your business?

    Book Your Free AI Clarity Call