INSIGHTS

    Is your firm overpaying for the wrong AI models?

    AI Dojo Team•October 6, 2026

    The pace of AI model releases over the past month has been remarkable. Major updates from Anthropic, OpenAI and Google have arrived in quick succession, each bringing measurable improvements in capability, speed and running costs.

    For accounting firms, keeping track of these changes can feel exhausting, but the underlying economics are worth paying attention to. The model that offered the best performance for financial analysis or complex tax work six months ago is no longer the clear default, and options that were previously cost-prohibitive for high-volume tasks are now accessible at a fraction of their original price.

    Here is a look at what the latest releases have introduced, and why model selection matters more than ever.

    Anthropic: Claude 5.5 and Haiku 4.5

    Anthropic recently updated its core lineup with Claude Opus 5.5 and Claude Sonnet 5.5, followed by Claude Haiku 4.5.

    The improvements in execution and token efficiency are significant:

    • Claude Sonnet 5.5: While maintaining its base pricing of $2 per million input tokens and $10 per million output tokens, Sonnet 5.5 operates roughly 30% faster and requires fewer tokens to complete equivalent work, lowering per-task costs by up to 30%. On Terminal-Bench 4.0, which evaluates autonomous command-line execution on complex tasks, Sonnet 5.5 scored 70.6%, compared with 10.3% for Sonnet 5.
    • Claude Opus 5.5: Positioned for complex reasoning and deep technical work, Opus 5.5 costs approximately 40% less to run than Opus 5 on typical workloads. Pricing sits at $4 per million input tokens and $20 per million output tokens, with cache read costs reduced by 60% to $0.20 per million tokens. On the FrontierCode evaluation, Opus 5.5 matches or exceeds top-tier frontier performance at roughly one-fifth the per-task cost of earlier flagship models.
    • Claude Haiku 4.5: Designed for fast, high-volume tasks, Haiku 4.5 delivers performance comparable to earlier frontier models like Sonnet 4, but runs at more than twice the speed and one-third the cost ($1 per million input tokens and $5 per million output tokens).
    Artificial Analysis Intelligence Index comparing AI model performance
    Intelligence: a snapshot of model performance across the Artificial Analysis evaluations. Source: Artificial Analysis, supplied chart.

    OpenAI: GPT-6.1 Sol and the GPT-6 family

    OpenAI expanded its GPT-6 family with the release of GPT-6 Sol and Luna, followed closely by GPT-6.1 Sol.

    The focus of these releases has been narrowing the gap between flagship reasoning and everyday operating costs:

    • Price reductions across tiers: GPT-6 Sol halved API pricing relative to GPT-5.6 Sol, dropping to $2 per million input tokens and $10 per million output tokens. GPT-6 Luna sits at $0.10 per million input tokens and $0.50 per million output tokens.
    • GPT-6.1 Sol: This model approaches the benchmark performance of OpenAI's flagship model, GPT-6 Astra, across agentic coding, computer use and professional knowledge work, but at one-fifth the token cost. Cached input pricing is set at $0.10 per million tokens, representing a 95% reduction on standard input rates. On the DeepSWE software engineering benchmark, GPT-6.1 Sol achieves near-Astra performance at roughly 20% of the cost per task.
    Artificial Analysis cost per Intelligence Index task comparing AI models
    Cost: the supplied comparison shows how the cost of completing a task varies across models (USD). Source: Artificial Analysis, supplied chart.

    Google: Gemini 3.1 Pro and Gemini 3.8 Flash

    Google has maintained a rapid release cycle across both its frontier reasoning and lightweight tiers:

    • Gemini 3.1 Pro: Google's updated reasoning model scores 77.1% on ARC-AGI-2, a benchmark evaluating novel pattern logic and problem-solving, more than double the score of Gemini 3 Pro. It features a 1 million token context window, handles text, audio, video and complex documents like financial PDFs, and is priced at $2 per million input tokens and $12 per million output tokens.
    • Gemini 3.8 Flash: Priced at an introductory $0.75 per million input tokens and $3.75 per million output tokens, 3.8 Flash delivers substantial reasoning improvements over 3.7 Flash. On benchmarks evaluating complex multi-step workflows, including DeepSWE and professional domain evaluations, it demonstrates performance that rivals significantly larger, more expensive models.
    Artificial Analysis output speed comparison in tokens per second
    Speed: output tokens per second vary significantly across models. Higher is faster. Source: Artificial Analysis, supplied chart.

    Why locking into one model creates commercial risk

    These releases highlight an important dynamic: no single technology provider holds a permanent lead across every capability.

    One provider might release a model that excels at long-document review and structured data extraction, only for a competitor to release an alternative two weeks later that handles iterative spreadsheet calculations at half the price.

    When a firm standardises on a single closed ecosystem, it accepts two clear disadvantages:

    1. Paying a premium for routine tasks: Using a frontier reasoning model for basic document summarisation, formatting or standard correspondence is commercially inefficient.
    2. Missing performance gains: When a competitor lab makes a genuine technical leap, a firm tied to one tool cannot adopt it without renegotiating vendor agreements, changing software and retraining staff.

    The advantage of model flexibility

    This is the principle behind Dojo. Rather than tying your firm to a single model or provider, Dojo integrates top-tier models from Anthropic, OpenAI, Google and specialised open-weight options like DeepSeek, Qwen and MiniMax within a single accounting workspace.

    As models update, your team benefits immediately without changing how they work:

    • Task-specific selection: You can route complex corporate tax structuring to deep reasoning models, while directing high-volume data extraction or preliminary drafting to fast, economical options.
    • Automated optimisation: Dojo's routing automatically matches the appropriate model to the specific accounting workflow, balancing accuracy, turnaround time and cost.
    • Cost control: You retain full visibility over firm-wide usage and consumption, ensuring spending aligns directly with completed work.

    In a market where model performance and pricing change every month, flexibility is the most practical way to protect your firm's margins and ensure your team is always using the right tool for the job.

    Find the right models for your firm

    To see how Dojo helps firms match the right models to accounting workflows while managing consumption costs, book a short walkthrough with our team or learn more about the platform.