Blog
5 September 2026/10 min read

Artificial Analysis Intelligence Index: What It Means for GEO

Three frontier models shipped within 48 hours in September 2026, and OpenAI's own benchmark table does not agree with the independent Artificial Analysis Intelligence Index. Here is what the gap means if you track AI visibility for a living.

Abdelmoghit Idhsaine
Author:Abdelmoghit Idhsaine,Content Strategist
Artificial Analysis Intelligence Index: What It Means for GEO

Three frontier AI labs shipped new flagship models within 48 hours of each other in early September 2026: Google's Gemini 3.8 Flash on September 2, then OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 on the same day, September 3. On the Artificial Analysis Intelligence Index, the one cross-vendor benchmark not run by any of the three labs themselves, Astra scored 61. Claude Fable 5.1 scored 66. That gap, on the one scoreboard nobody involved gets to grade themselves on, is the actual story.

OpenAI's own launch page told a different story. Two hours after its planned publish time, the page went live with revised numbers, including Astra's stated hallucination rate falling from 4.2% to 2% in a version saved 80 minutes later (Source: Fortune, September 2026). If you track which AI model gets cited, recommended, or trusted for your brand's category, a week where three labs ship at once and self-reported numbers move after launch is not just tech news. It is a measurement problem for anyone doing GEO or AI-visibility work.

What is the Artificial Analysis Intelligence Index?

The Artificial Analysis Intelligence Index is an independent, cross-vendor benchmark that combines results from roughly a dozen individual evaluations (math, reasoning, coding, knowledge tests) into a single comparative score, run by a third party rather than self-reported by the model's own lab. That independence is the entire point: OpenAI, Anthropic, and Google all publish their own launch-day benchmark tables, and every one of those tables is produced by the vendor being graded.

As of September 2026, the index places Claude Fable 5.1 at the top of the three new launches, with GPT-6 Astra five points behind at 61, tied with the prior OpenAI flagship GPT-5.6 Sol (Source: Artificial Analysis). Astra is not weak everywhere. It leads decisively on FrontierMath, cybersecurity exploitation benchmarks, and desktop/computer-use tasks, and it does agentic coding work at roughly half the token cost of Fable 5.1 for a comparable score. It simply does not lead on the one index built specifically so no single vendor controls the scoreboard.

Warning: a benchmark table published by the same company whose model is being scored is not an independent measurement. Cross-check any launch-day number against a third-party index before repeating it in client-facing material.

Why did OpenAI's own benchmark numbers change after launch?

OpenAI revised several of Astra's evaluation metrics in the hours after its blog post went live, including lowering the stated hallucination rate and, in the same revision pass, showing worse numbers for rival Anthropic models than the original version had (Source: Fortune, September 2026). OpenAI's own explanation to Fortune was that "evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run," and that the changes reflected the company's "best estimate of available model performance."

That explanation may be true and the numbers still worth treating with caution. A benchmark that moves after publication, in the vendor's own favor, without an independent re-run to confirm it, is not a number a marketing or growth team should cite without a caveat. It is the same failure mode covered in our comparison of AI visibility audit services: a single, undisclosed-methodology number from the party being measured is a claim, not a fact.

61 vs. 66
Astra vs. Fable 5.1
On the independent Artificial Analysis Intelligence Index, September 2026
48 hours
Between 3 flagship launches
Gemini 3.8 Flash (Sep 2), GPT-6 Astra + Claude Fable 5.1 (Sep 3)
4.2% → 2%
Astra's hallucination rate, revised
Changed in a version of OpenAI's launch page saved 80 minutes later (Fortune)

How do the three new models actually compare?

No single model swept every benchmark, which is itself the more useful finding than picking a winner.

BenchmarkGPT-6 AstraClaude Fable 5.1Claude Opus 5What it measures
Artificial Analysis Intelligence Index616663Independent, cross-vendor composite score
FrontierMath Tier 4 v297.6%87.8%73.2%Advanced mathematical reasoning
Cybersecurity exploitation (ExploitBench)100%not reportednot reportedFinding and exploiting security flaws
Agentic coding cost-efficiencyComparable score at roughly half the token costBaselinenot reportedSame task outcome, cost per run
Price (standard API)$10/M input, $50/M outputMatches Astra~$5/M inputPer-token cost

Three uneven bar heights against a shared benchmark line, illustrating how independent scores diverge from vendor-reported claimsThree uneven bar heights against a shared benchmark line, illustrating how independent scores diverge from vendor-reported claims

Read the table by task, not by a single row. A team choosing a model for cybersecurity research or high-volume agentic coding has a real reason to pick Astra. A team optimizing for general reasoning quality, the kind that shapes how a model paraphrases or summarizes a brand in an AI-generated answer, has a real reason to prefer Fable 5.1 or Opus 5, per the Intelligence Index.

Key Takeaway

The Artificial Analysis Intelligence Index exists because vendor-published benchmarks are not independent measurements. When a launch-day number moves after publication, as Astra's did, treat the vendor's own table as a claim and the third-party index as the check on that claim.

What does a three-way frontier launch mean for AI visibility tracking?

Here is the part that matters for GEO rather than for AI news generally: ChatGPT, Perplexity, Gemini, and Google AI Overviews do not necessarily switch to a lab's newest model on the day it ships to the API. Consumer-facing products often lag behind an API release by days or weeks, run the new model alongside older ones for a subset of traffic, or route different query types to different model versions entirely. A citation rate you measured last week reflects whichever model mix was actually answering questions at the time, not whatever a lab announced this week.

That has a direct, practical consequence for anyone running AI citation tracking: a GEO score or visibility baseline captured before a major model swap can go stale the moment the underlying answer engine changes which model it is running, even if nothing about your content changed at all. Three labs shipping within 48 hours is a real-world stress test of exactly that gap. If your tracking tool re-baselines slowly, or does not log which model version it queried, you cannot tell whether a citation-rate move next month reflects your content or the model underneath the product you are tracking.

Not sure which model is actually answering the questions that matter to your brand?
Our AI visibility audit checks how your brand is cited across ChatGPT, Perplexity, Gemini, and Google AI Overviews, and logs which model version answered each run so a re-baseline after a launch like this one is a documented fact, not a guess.
Get Your Free Audit

Should you change your GEO strategy every time a new model ships?

No, and the three-launch week is a good argument against it. Astra, Fable 5.1, and Gemini 3.8 Flash each specialize differently (agentic cost-efficiency, general reasoning and frontend quality, and price, respectively), so betting your entire content and entity-optimization strategy on whichever lab had the loudest launch post this week is a way to whiplash your own roadmap on someone else's release schedule.

Worth doing after a major model launch
  • Re-run your citation-tracking prompt set to confirm which model version is actually live in the product you're measuring
  • Check whether your entity clarity and structured data hold up under a genuinely different model's retrieval behavior, not just the one you tested against last quarter
  • Note the launch date in your tracking log so a score change has a documented, testable explanation
Not worth doing after every launch
  • Rewriting content strategy around one lab's self-reported benchmark table before an independent index confirms it
  • Treating a single week's Intelligence Index snapshot as a permanent ranking rather than one data point in a fast-moving field
  • Chasing whichever model is loudest on social media instead of the one actually serving your target queries

How should GEO teams respond to frontier model volatility?

Start by separating the news cycle from your measurement cycle. The generative engine optimization fundamentals, entity clarity, retrievability, third-party corroboration, do not change because a new model shipped. What can change is which model is doing the retrieving, and that is worth checking, not assuming.

Second, treat any vendor's own launch-day benchmark table the way you would treat a single-run GEO score from an unverified tool: informative, but not something to repeat to a client without a caveat. The Fortune reporting on Astra's post-launch metric changes is a useful reminder that even a frontier lab's own numbers can shift after publication, independent of anyone's marketing intent.

Third, if your AI visibility monitoring setup does not log which model version answered a given prompt run, that is a gap worth closing before the next launch week, not after. A visibility audit that records the model version alongside the citation result is the difference between explaining a score change and guessing at one.

Our GEO optimization services are built around that same principle: model churn is a constant in this category, and a citation strategy that depends on one lab staying still is not a strategy, it is a bet on someone else's release calendar.

Frequently asked questions

What is the Artificial Analysis Intelligence Index?

The Artificial Analysis Intelligence Index is an independent, cross-vendor benchmark that combines roughly a dozen individual AI model evaluations into one comparative score, run by a third party rather than the labs whose models it grades. As of September 2026, it places Claude Fable 5.1 ahead of GPT-6 Astra (66 vs. 61) despite OpenAI's own launch materials emphasizing different, self-reported benchmarks.

Why does GPT-6 Astra score lower than Claude Fable 5.1 on the Intelligence Index but win on other benchmarks?

Astra and Fable 5.1 are specialized differently rather than one being universally stronger: Astra leads on math (FrontierMath), cybersecurity exploitation, and cost-efficient agentic coding, while Fable 5.1 scores higher on the Intelligence Index's broader composite of general reasoning and knowledge tasks. A single index score does not capture task-specific strengths.

Did OpenAI change GPT-6 Astra's benchmark numbers after launch?

Yes. OpenAI revised several evaluation metrics for Astra in the hours following its blog post, including lowering the stated hallucination rate from 4.2% to 2% in a version saved roughly 80 minutes after the page's delayed publication (Source: Fortune, September 2026). OpenAI attributed the changes to normal noise across evaluation checkpoints and runs.

How fast do AI answer engines like ChatGPT actually adopt a newly launched model?

There is no fixed timeline; consumer-facing products frequently run a new model alongside older versions, roll it out gradually, or route only certain query types to it, rather than switching entirely on the API launch date. This is why an AI SEO strategy built on logging the model version behind each result matters more after a major launch than at any other time.

Should a GEO or content team switch strategy every time a new frontier model launches?

No. Individual model launches change which model may be doing the retrieving behind an AI answer engine, but the underlying GEO fundamentals (entity clarity, structured and extractable content, third-party corroboration) do not change, so a re-baseline of tracking is more useful than a strategy rewrite.

How many frontier models launched in September 2026, and how close together?

Three: Google's Gemini 3.8 Flash on September 2, 2026, followed by both OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 on September 3, 2026, a 48-hour window across all three labs.

Is a vendor's own launch-day benchmark table a reliable number to cite?

Treat it as a claim rather than an independent measurement, since the lab being graded produced it. Cross-checking against a third-party index like Artificial Analysis, and noting whether the vendor has revised the number post-launch, is the more defensible approach before repeating a figure in client-facing reporting.


Sources: Benchmarking GPT-6 Astra (Artificial Analysis, 2026), OpenAI quietly boosts some of Astra's evaluation metrics (Fortune, September 2026), GPT-6 Astra vs Claude Fable 5.1 model comparison (Artificial Analysis)

This post is part of our GEO Optimization guide. Related reading: AI SEO Checklist for 2026, ChatGPT SEO, GEO Implementation Guide for 2026.

About the Author
Abdelmoghit Idhsaine
Abdelmoghit Idhsaine
Content Strategist

Abdelmoghit drives the content engine at AY Rank. He researches keywords, plans content clusters, and produces citation-optimized articles that rank in both Google and AI search engines.

Full Bio →
More From the Blog
Product Pages Beat Blogs for B2B AI Citations

Product Pages Beat Blogs for B2B AI Citations

A 768,000-citation study found product pages make up nearly 56% of B2B AI citations, while blogs, PR, and educational content trail well behind. Here is what the data means for where B2B content teams should actually invest.

Read article →
1 in 10 AI Citations Are From Self-Promo Listicles

1 in 10 AI Citations Are From Self-Promo Listicles

A 232,000-citation study found roughly 1 in 10 AI search citations come from a self-promotional listicle, a vendor's own "best tools" post ranking itself first. Here is what the data actually shows, and the long-term risk that comes with the tactic.

Read article →
Domain Authority Doesn't Predict AI Citations

Domain Authority Doesn't Predict AI Citations

An analysis of 5 million AI citations found domain authority correlates with citation frequency close enough to zero to be noise. A separate study found over a third of citations came from sites with Domain Authority under 40. Here is what actually predicts a citation instead.

Read article →