Claude vs ChatGPT for Excel: an honest comparison.
Within eight weeks this spring, Anthropic and OpenAI both took their Excel add-ins to general availability — Claude for Excel in March, ChatGPT for Excel in May. "Claude for Excel" is now the top rising search next to "ChatGPT for Excel," which means a lot of people are standing at this exact fork. We build GridPath, a spreadsheet agent that runs on either subscription, so we genuinely don't have a horse in the model race. This is the comparison we'd want to read: what each add-in actually is, what the published head-to-head tests found, what practitioners report once the demo glow wears off — and the wall both products share that nobody puts in the launch post.
The short version
| Claude for Excel | ChatGPT for Excel | |
|---|---|---|
| Plans | Paid only — Pro ($20/mo) and up | All plans, including Free |
| Where it runs | Excel (Microsoft 365 builds) | Excel (Microsoft 365 builds) + Google Sheets |
| Strongest at | Model construction, error-finding, financial work | Speed, iteration volume, explaining, Sheets crossover |
| Own-admitted limits | No VBA/macros, no data tables; "not for audit-critical work" | Complex formulas need refinement; "can make unintended edits" |
What each add-in actually is
Claude for Excel
A sidebar add-in, part of Anthropic's Claude-for-Microsoft-365 suite (Word and PowerPoint are also GA; Outlook is in beta). It shipped in October 2025 and went generally available in March 2026. It's included with every paid Claude plan — Pro at $20/month and up — with no separate fee, and no free tier. You can switch between Opus and Sonnet models from the sidebar.
The capability set is aimed squarely at finance work: it navigates
multi-tab workbooks and answers with cell-level
citations that jump to the referenced cells, updates
assumptions while preserving formula dependencies, debugs
#REF! / #VALUE! / circular references, and
applies native Excel operations directly — sorting, filtering,
editing pivot tables and charts, conditional formatting, data
validation. It also has data connectors to S&P Global, LSEG,
Daloopa, PitchBook, Moody's, and FactSet.
Anthropic's
own docs are unusually direct about the limits: no macros or
VBA, no data tables, and it's "not recommended" for audit-critical
calculations or final client deliverables without human review.
ChatGPT for Excel
Also a sidebar add-in — for Excel and Google Sheets, which Claude's isn't. It went GA on May 5, 2026 on the GPT-5.5 family, and the headline difference is distribution: it's available on every ChatGPT plan, including Free. It builds spreadsheets from a prompt, summarizes across tabs, links its answers to the cells it referenced, and asks permission before making changes. OpenAI shipped financial data integrations alongside it (FactSet, Dow Jones Factiva, LSEG, Daloopa, S&P Global).
OpenAI's docs are similarly candid where it counts: complex formulas and edge cases "may require manual refinement," VBA "may not be fully supported," the add-in's chats don't sync with your chatgpt.com history and have no memory, and — their words — it "can make mistakes, including making unintended edits or deleting content if a request is unclear."
What the head-to-head tests found
Most pages comparing these two products are feature tables. Two authors actually ran the same tasks through both, and their results agree with each other:
- F9 Finance ran the same FP&A dataset and prompts through Claude, ChatGPT, and Copilot. Claude finished first "and it wasn't close" — roughly 90% usable on the first pass. ChatGPT came last, with a forecast model the author called structurally broken and missing a line item that was sitting in the source data; his conclusion was that he "would not use the ChatGPT for Excel plugin for core FP&A work right now."
- FindSkill ran four tasks side-by-side and split the verdict: Claude for serious financial modeling, ChatGPT for speed, polish, and Google Sheets crossover.
- A third test planted three errors in a model — a circular reference, an inconsistent growth assumption, a hardcoded value — and asked each tool to find them. Claude caught all three and explained why each mattered; ChatGPT caught two; Copilot caught only the circular reference, which Excel itself already flags.
One more data point made the rounds in May: a widely shared test of
file generation where Claude produced a valid
.xlsx in one shot while ChatGPT produced a corrupt
file four attempts running. Whatever it says about the models, it
says something real about how much of this product category is the
unglamorous file-format layer, not the intelligence. (We've
measured that layer
ourselves — it's where tools quietly lose your data.)
What practitioners report
The test articles are single authors on single days. The longer record is the practitioner threads — accountants, FP&A analysts, and data folks comparing notes over months. Read enough of them and the pattern is remarkably stable:
- Claude owns the workbook-manipulation job. A 189-upvote r/Accounting thread on Claude's add-in is full of working formulas, multi-step tasks that hold together, and messy client files handled without hand-holding. The most-quoted line: "Finding errors in thousands of lines of GL data, setting analytics or underwriting models is a good prompt away." The r/dataanalysis consensus is blunter still — "I don't think you can get any better than Claude for excel."
- ChatGPT owns volume and versatility. The same threads credit it with long-form reasoning, pattern-heavy VBA generation, trial-and-error debugging, explaining code, and executive communication. FP&A users who live in it describe it as stronger for "analytical and structured work" — and it doesn't hit a paid-plan wall, which matters if your whole team is on Free.
- The variable that matters most isn't the model. Practitioners who get great results from either tool report the same habits: well-organized files, context provided up front (some maintain a dedicated notes sheet describing the workbook's structure), and very specific prompts. People with chaotic inherited workbooks report both tools struggling — Claude's most consistent reported weakness is exactly the "highly complex, poorly constructed file."
- Nobody defends Copilot. Across every thread we read, Microsoft's tool is the consensus last place — refused macros, partial snippets, drift on multi-step tasks. More on it below, because it still wins one thing.
Where both of them fail
The most useful review of Claude's add-in we found wasn't a rave — it was a CPA's careful post-mortem of building a pipeline-driven ARR forecast purely by prompting. The model produced a plausible-looking build with drivers too coarse to test assumptions against, static formulas where dynamic ones were needed, and a core logic gap — it ignored open deals outstanding per month — that would have materially misstated the forecast. His conclusion: getting deliverable quality from a single prompt required so much prompt detail that the time saved evaporated.
That matches both vendors' own disclaimers, and it matches everything we've written before: these tools produce output at the quality bar of "plausible," and the review pass is not optional. Claude being ahead on spreadsheet work right now doesn't change the shape of the job — it changes how often the review pass finds something.
What about Copilot?
Copilot in Excel deserves one honest paragraph. Its Agent Mode went GA on desktop in January and can now run Claude models under the hood — you can literally pick Anthropic from Microsoft's model switcher, which tells you how the capability race is going. It costs roughly $30/user/month on the commercial license, wants your files in OneDrive or SharePoint, isn't available in the EU or UK, and applies its changes to your workbook directly, with no review step. It places last in every capability test we found — and it will still be the right answer for plenty of teams, because it's the only one of the three that keeps everything inside the Microsoft tenancy your IT department already trusts. Weakest tool, strongest compliance story.
The wall neither launch post mentions
Both add-ins require a Microsoft 365 subscription build of Excel — the web app, or current Windows/Mac desktop builds. If you're on a perpetual license (Excel 2016, 2019, or an LTSC build your company standardized on precisely to keep AI away from its data), neither add-in will install. Same if you don't run Excel at all. The entire comparison above is moot for that crowd, and it's a bigger crowd than either vendor acknowledges.
So which should you use?
- Financial modeling, error-hunting, inherited multi-tab workbooks — Claude for Excel. Every test and every practitioner thread points the same direction, and the cell-level citations make its answers checkable, which is the point.
- High-volume everyday work, VBA drafting, explaining, Google Sheets — ChatGPT for Excel. Faster iteration, no paid-plan wall, and the only one of the two that follows you into Sheets.
- Locked-down enterprise tenancy — Copilot, eyes open. You're trading capability for compliance, and that's a legitimate trade.
- Whichever you pick — organize the file, describe the structure, be specific, and review everything. The practitioner record says those habits move results more than the logo on the sidebar does.
And hold the ranking loosely. These two products went GA eight weeks apart and both are shipping monthly. This is the snapshot, not the destination.
Where GridPath fits
GridPath is the third path the add-in comparison skips: a local
desktop agent that edits your real .xlsx directly —
no Microsoft 365 requirement, no upload, every change a
reviewable diff you accept or reject —
running on the same Claude or ChatGPT subscription you'd buy for
either add-in. If you're already paying the $20, you have both
paths. Download for Mac or Windows, or
read why we built it.