← All posts

Claude vs ChatGPT for Excel: an honest comparison.

The short version

Claude for Excel ChatGPT for Excel
Plans Paid only — Pro ($20/mo) and up All plans, including Free
Where it runs Excel (Microsoft 365 builds) Excel (Microsoft 365 builds) + Google Sheets
Strongest at Model construction, error-finding, financial work Speed, iteration volume, explaining, Sheets crossover
Own-admitted limits No VBA/macros, no data tables; "not for audit-critical work" Complex formulas need refinement; "can make unintended edits"

What each add-in actually is

Claude for Excel

A sidebar add-in, part of Anthropic's Claude-for-Microsoft-365 suite (Word and PowerPoint are also GA; Outlook is in beta). It shipped in October 2025 and went generally available in March 2026. It's included with every paid Claude plan — Pro at $20/month and up — with no separate fee, and no free tier. You can switch between Opus and Sonnet models from the sidebar.

The capability set is aimed squarely at finance work: it navigates multi-tab workbooks and answers with cell-level citations that jump to the referenced cells, updates assumptions while preserving formula dependencies, debugs #REF! / #VALUE! / circular references, and applies native Excel operations directly — sorting, filtering, editing pivot tables and charts, conditional formatting, data validation. It also has data connectors to S&P Global, LSEG, Daloopa, PitchBook, Moody's, and FactSet. Anthropic's own docs are unusually direct about the limits: no macros or VBA, no data tables, and it's "not recommended" for audit-critical calculations or final client deliverables without human review.

ChatGPT for Excel

Also a sidebar add-in — for Excel and Google Sheets, which Claude's isn't. It went GA on May 5, 2026 on the GPT-5.5 family, and the headline difference is distribution: it's available on every ChatGPT plan, including Free. It builds spreadsheets from a prompt, summarizes across tabs, links its answers to the cells it referenced, and asks permission before making changes. OpenAI shipped financial data integrations alongside it (FactSet, Dow Jones Factiva, LSEG, Daloopa, S&P Global).

OpenAI's docs are similarly candid where it counts: complex formulas and edge cases "may require manual refinement," VBA "may not be fully supported," the add-in's chats don't sync with your chatgpt.com history and have no memory, and — their words — it "can make mistakes, including making unintended edits or deleting content if a request is unclear."

What the head-to-head tests found

Most pages comparing these two products are feature tables. Two authors actually ran the same tasks through both, and their results agree with each other:

One more data point made the rounds in May: a widely shared test of file generation where Claude produced a valid .xlsx in one shot while ChatGPT produced a corrupt file four attempts running. Whatever it says about the models, it says something real about how much of this product category is the unglamorous file-format layer, not the intelligence. (We've measured that layer ourselves — it's where tools quietly lose your data.)

What practitioners report

The test articles are single authors on single days. The longer record is the practitioner threads — accountants, FP&A analysts, and data folks comparing notes over months. Read enough of them and the pattern is remarkably stable:

Where both of them fail

The most useful review of Claude's add-in we found wasn't a rave — it was a CPA's careful post-mortem of building a pipeline-driven ARR forecast purely by prompting. The model produced a plausible-looking build with drivers too coarse to test assumptions against, static formulas where dynamic ones were needed, and a core logic gap — it ignored open deals outstanding per month — that would have materially misstated the forecast. His conclusion: getting deliverable quality from a single prompt required so much prompt detail that the time saved evaporated.

That matches both vendors' own disclaimers, and it matches everything we've written before: these tools produce output at the quality bar of "plausible," and the review pass is not optional. Claude being ahead on spreadsheet work right now doesn't change the shape of the job — it changes how often the review pass finds something.

What about Copilot?

Copilot in Excel deserves one honest paragraph. Its Agent Mode went GA on desktop in January and can now run Claude models under the hood — you can literally pick Anthropic from Microsoft's model switcher, which tells you how the capability race is going. It costs roughly $30/user/month on the commercial license, wants your files in OneDrive or SharePoint, isn't available in the EU or UK, and applies its changes to your workbook directly, with no review step. It places last in every capability test we found — and it will still be the right answer for plenty of teams, because it's the only one of the three that keeps everything inside the Microsoft tenancy your IT department already trusts. Weakest tool, strongest compliance story.

The wall neither launch post mentions

Both add-ins require a Microsoft 365 subscription build of Excel — the web app, or current Windows/Mac desktop builds. If you're on a perpetual license (Excel 2016, 2019, or an LTSC build your company standardized on precisely to keep AI away from its data), neither add-in will install. Same if you don't run Excel at all. The entire comparison above is moot for that crowd, and it's a bigger crowd than either vendor acknowledges.

So which should you use?

And hold the ranking loosely. These two products went GA eight weeks apart and both are shipping monthly. This is the snapshot, not the destination.