Original Research: This AI Visibility Benchmark evaluated nine leading SEO websites using the same three AI optimization tools. Every article was tested using an identical methodology, and all scores reflect the measured results at the time of testing.
AI visibility advice is everywhere. But almost nobody asks a simple question: are the websites teaching AI optimization actually following their own advice?
I decided to find out.
I picked nine well-known SEO publishers, chose one educational article from each, and ran every article through three tools from my free AI Visibility Toolkit. No manual adjustments. No editorial scoring. Just the numbers.
What I found was more interesting than I expected.
Table of Contents
Research Summary
- 9 SEO publishers tested.
- 27 individual benchmark measurements.
- 3 AI optimization tools.
- Same method for every site.
- Same testing period for the retest.
- Schema scores corrected after a tool bug was found and fixed.
Benchmark at a Glance
- Highest AI Visibility score: 83 (Search Atlas)
- Highest Citability score: 91 (Neil Patel)
- Highest Schema score: 100 (Neil Patel)
- Average AI Visibility: 68
- Average Citability: 78
- Average Schema score: 64
That last number moved the most. The first version of this benchmark showed an average Schema score of 16. The real number is 64. That is not a small correction. That is a different story.
Key Findings
- Search Atlas achieved the highest overall AI Visibility Score in the benchmark.
- Neil Patel produced the most citable content and, after correcting the checker, earned a perfect 100 Schema Score.
- After fixing the Schema Checker, none of the tested websites scored zero for schema implementation.
- HubSpot had the weakest schema implementation, scoring just 10, despite maintaining strong content quality.
- High content quality alone did not guarantee strong AI Visibility.
- Moz was the clearest example of the gap between high-quality content and lower AI Visibility performance.
Why I Built This Benchmark
There is no shortage of AI visibility advice right now. Almost every major SEO publisher has published articles on GEO, AI citations, and AI search optimization. That advice is often good. But it is also untested in the one place that matters most — the publisher’s own site.
I wanted to see the gap between what these websites teach and what they actually implement. Not to criticize anyone. To find the patterns that any website owner could learn from. Real scores from real pages tell a clearer story than any opinion piece.
Benchmark Methodology
Why These Nine Websites
I chose nine websites that represent a broad cross-section of the SEO and digital marketing industry. All of them publish regular educational content about search optimization, AI search, or related topics. All of them have significant audiences. And, all of them have published their own guidance on AI visibility in some form.
Articles Tested
The table below shows which article was tested for each website and the primary topic it covered.
| Website | Article Tested | Topic |
|---|---|---|
| Search Atlas | How to Use AI Agents for Content Gap Analysis | AI content strategy |
| Semrush | What Is AI Sentiment Analysis? | AI search education |
| Neil Patel | How We Rebuilt AI Visibility Measurement From The Ground Up | AI visibility |
| Ahrefs | 10 ChatGPT SEO Tools | AI tools |
| Yoast | Structured Data with Schema for Search and AI | Schema and AI |
| Backlinko | How to Rank in AI Search | AI search ranking |
| Search Engine Journal | General SEO article | General SEO |
| HubSpot | Loop Marketing | Marketing education |
| Moz | Stop Measuring AI Search Like SEO | AI search strategy |
Testing Process
I tested every article with the same three tools, in the same order, on the same day. One article per site. No retesting to chase a better number, except for the schema correction explained above, which touched every site equally.
The Three AI Visibility Tests
AI Visibility Checker
Measures overall AI readiness. Checks crawler access, XML sitemap, technical signals, structural elements, and metadata quality. Returns a score out of 100.
Content Citability Grader
Measures how likely a piece of content is to be cited by an AI system. Evaluates four dimensions: Evidence, Structure, Authority, and AI Readability — each worth 25 points. Inspired by publicly available GEO research factors including findings from Princeton’s study on generative engine optimization.
Schema Checker
Evaluates structured data implementation. Checks for six schema types — Article, FAQPage, Organization, WebSite, Person, and BreadcrumbList — and rates each as Complete, Basic, or Missing. Returns a score out of 100.
Overall Benchmark Results
The table below shows all 27 measurements across nine sites and three tools, with corrected schema numbers.
| Website | AI Visibility | Content Citability | Schema Score |
|---|---|---|---|
| Search Atlas | 83 | 85 | 75 |
| Semrush | 78 | 77 | 35 |
| Neil Patel | 77 | 91 | 100 |
| Ahrefs | 75 | 67 | 65 |
| Yoast | 70 | 75 | 75 |
| Backlinko | 65 | 90 | 75 |
| Search Engine Journal | 65 | 59 | 75 |
| HubSpot | 61 | 83 | 10 |
| Moz | 39 | 80 | 65 |
| Average | 68 | 78 | 64 |
No single publisher won every category. But the schema story changed completely once I fixed my tool. Most of these sites are doing schema well. Two are not.
AI Visibility Hall of Fame
🏆 Best Overall AI Visibility — Search Atlas (83). Highest raw AI Visibility score in the benchmark. Strong across every measured signal.
🏆 Best Content Citability — Neil Patel (91). Highest citation potential of any tested article. Strong structure, strong evidence, clear first-hand authority.
🏆 Best Schema Implementation — Neil Patel (100). A perfect score, and the biggest single change from the correction. Every schema type I check for was present and complete.
🏆 Most Balanced Performer — Neil Patel. Once I averaged all three scores, Neil Patel came out ahead of every other site, at 77, 91, and 100. Search Atlas is still the strongest on raw AI Visibility, but Neil Patel is now the most complete performer across all three tools combined.
🏆 Biggest Surprise — HubSpot. Content quality was strong, an 83 on Citability. Schema was the weakest in the whole benchmark, just 10. That gap, high-quality writing sitting on almost no structured data, was the clearest miss in the corrected results.
🏆 Most Interesting Outlier — Moz. Citability sat at 80, comfortably above the benchmark average. AI Visibility sat at 39, the lowest score in the whole benchmark. That gap did not move when I fixed the schema bug. It is still the widest gap between content quality and technical AI readiness I found.
🏆 Citation Champion — Neil Patel. Highest Citability score at 91. Backlinko sat close behind at 90. Both wrote content that scored well on evidence, structure, and authority.
Overall AI Visibility Scores
AI Visibility Checker results — 9 leading SEO websites tested with identical methodology
Content Citability Scores
Content Citability Grader results — educational quality varied more than expected
Individual Website Analysis
Search Atlas
AI Visibility: 83 | Citability: 85 | Schema: 75
Search Atlas still ranks near the top on all three measures. Its AI Visibility score leads the benchmark. Its Citability score sits second only to Neil Patel. Its schema score held steady at 75 after the correction, a solid number, though no longer the highest in the benchmark now that Neil Patel’s real score came in at 100.



The tested article combined technical signals with high-quality educational content. The schema implementation was the most complete of any site in the benchmark.
Research conclusion: Search Atlas remains the strongest performer on raw AI Visibility in this benchmark.
Semrush
AI Visibility: 78 | Citability: 77 | Schema: 35
Semrush ranks second on AI Visibility and stays balanced across all three tools. Its schema score of 35 did not change in the correction. What changed is the context around it. In the first version of this benchmark, 35 looked like a strong second place. Now, with most other sites scoring 65 and above, 35 looks like one of the weaker schema results in the group.



Not outstanding in any single category, but steady across all three. That consistency still matters, even with the schema number sitting lower than I first framed it.
Research conclusion: Semrush’s schema implementation trails most of the field once the corrected numbers are in view.
Neil Patel
AI Visibility: 77 | Citability: 91 | Schema: 100
This is where the correction hit hardest. Neil Patel’s schema score jumped from 25 to a full 100 once I fixed the bug in my checker. Every schema type I test for showed up complete. Combined with the highest Citability score in the benchmark, Neil Patel now has the strongest average across all three tools of any site I tested.


I want to be direct about this one. My original tool told me Neil Patel scored 25 on schema. That was wrong, and it was my error, not theirs. The real number is 100.
Research conclusion: Neil Patel produced the most complete AI optimization profile in this benchmark, once corrected.
Ahrefs
AI Visibility: 75 | Citability: 67 | Schema: 65
Ahrefs is the second-biggest correction in this benchmark. The original zero on schema is gone. The real score is 65, a solid result that fits with the rest of the article’s optimization level. Citability at 67 stayed lower than I expected given the brand’s size and content volume, but that number did not change.


The old version of this section leaned hard on Ahrefs as proof that AI Visibility does not depend on schema at all. That framing does not hold up anymore. Ahrefs has real schema in place. It just is not the whole story either.
Research conclusion: Ahrefs shows solid, unremarkable schema work sitting under strong AI Visibility fundamentals.
Yoast
AI Visibility: 70 | Citability: 75 | Schema: 75
Yoast was supposed to be the biggest surprise in this benchmark. The tested article is literally about schema for search and AI. My original tool said it scored zero, which made for a dramatic headline and a wrong one. The real number is 75, in line with several other sites and nowhere near the embarrassment I first reported.


I owe Yoast a direct correction here. The claim that their own schema article had no detectable schema was false, and it came from a bug in my tool, not a failure on their end.
Research conclusion: Yoast’s schema implementation matches its own published advice reasonably well.
Backlinko
AI Visibility: 65 | Citability: 90 | Schema: 75
Backlinko still holds the second-highest Citability score in the benchmark, just one point behind Neil Patel. The schema story flipped completely. What showed as zero the first time now reads 75, a strong number that puts Backlinko in the upper half of the schema results, not the bottom.


The content quality here was outstanding before the correction and still is. What changed is that the technical side turns out to match it far better than I first reported.
Research conclusion: Backlinko pairs excellent content with solid schema, not the mismatch I originally described.
Search Engine Journal
AI Visibility: 65 | Citability: 59 | Schema: 75
Search Engine Journal still has the lowest Citability score of any site I tested, at 59. That part did not change. What did change is schema, which jumped from a reported zero to a real 75, one of the stronger schema results in the corrected data.


Brand authority and publishing volume still did not translate into strong Citability on the tested page. But the technical schema work behind it was better than I first gave credit for.
Research conclusion: Search Engine Journal’s weak point is content structure, not schema.
HubSpot
AI Visibility: 61 | Citability: 83 | Schema: 10
HubSpot is the one site that got worse relative to the field after correction, not better. Its schema score of 10 did not change, it was never affected by the bug. But now that every other site’s real numbers are visible, HubSpot’s 10 stands out as the clear weak point in the whole benchmark, sitting far below even Semrush’s 35.



Content quality here is genuinely strong, an 83 on Citability. The schema work behind it is close to nonexistent. That is now the sharpest content-versus-structure gap in the benchmark.
Research conclusion: HubSpot has the largest gap between content quality and schema implementation of any site tested.
Moz
AI Visibility: 39 | Citability: 80 | Schema: 65
Moz still holds the lowest AI Visibility score in the benchmark, and still shows the widest gap between AI Visibility and Citability of any site I tested. That story never depended on schema, so the correction did not touch it. Schema itself moved from a reported zero to a real 65, a decent number that had nothing to do with why Moz’s AI Visibility score is so low.


Strong educational content and clear authority signals were always present here. Whatever is holding AI Visibility down on this page, it is not a lack of schema.
Research conclusion: Moz’s low AI Visibility score comes from somewhere other than structured data.
Patterns I Discovered
These are the patterns that held up across nine pages, not conclusions about any one company.
Content quality and AI Visibility measured different things
This pattern survived the correction untouched.
Neil Patel scored 91 on Citability and 77 on AI Visibility. Backlinko scored 90 on Citability and 65 on AI Visibility. Moz scored 80 on Citability and 39 on AI Visibility. High educational quality still did not guarantee high AI Visibility. The gap is real, and it showed up on page after page.
Schema turned out to be far more common than my first test showed
This is the pattern that flipped hardest. My original numbers said six of nine tested pages had zero schema. The real number is zero of nine. Once I fixed the bug in my own checker, every single site showed real, working schema markup, ranging from HubSpot’s weak 10 up to Neil Patel’s perfect 100.
The average schema score across all nine tested articles is 64, not 16. That is a very different signal about how seriously these publishers treat their own advice.
The highest-performing site changed once I looked at all three scores together
Search Atlas still leads on raw AI Visibility. But once I averaged AI Visibility, Citability, and the corrected Schema score, Neil Patel came out ahead of every other site tested, at 77, 91, and 100. Consistency across all three measures turned out to matter more than leading in just one.
Large brands did not automatically score highest
Several sites with real domain authority and large content libraries still scored lower than I expected in one category or another. Brand size and publishing frequency did not predict strong AI optimization on the pages I tested. That held true before and after the correction.
Every publisher had a different profile
No two sites showed the same pattern of strengths and weaknesses, even after the correction. Some lead on content. Some lead on technical signals. HubSpot leads on neither when it comes to schema. There is no single formula every top SEO publisher follows.
Schema Implementation Scores
Schema Checker results, corrected. None of the nine tested articles scored zero.
| Rank | Website | Schema Score |
|---|---|---|
| 1 | Neil Patel | 100 |
| 2 | Search Atlas | 75 |
| 2 | Yoast | 75 |
| 2 | Backlinko | 75 |
| 2 | Search Engine Journal | 75 |
| 6 | Ahrefs | 65 |
| 6 | Moz | 65 |
| 8 | Semrush | 35 |
| 9 | HubSpot | 10 |
| Benchmark Average | All Websites | 64 |
The Yoast article, the one specifically about schema for search and AI, scored 75 once I fixed my checker. My first version of this benchmark said it scored zero. That was a bug in my tool, not a failure on Yoast’s part, and I want that correction on the record clearly.
What Website Owners Can Learn
The lessons here hold regardless of the size of the site you run.
Getting the fundamentals right still produced the strongest results. AI crawler access, real schema, clear heading structure, specific evidence in your content, and genuine author signals separated the stronger pages from the weaker ones. None of those cost much or take long to fix.
Do not chase one metric. The sites with the best overall profiles were not perfect on any single tool. Neil Patel led on two out of three measures and still was not the top score on AI Visibility. Consistency across all three beat dominance in one.
Schema is more common among strong publishers than my first draft of this benchmark suggested. The corrected average sits at 64, not 16. Most sites in this space are doing real work here. If your own schema score is near zero, you are behind most of your competition, not in line with an industry-wide gap.
Content quality is still a real, separate signal. Neil Patel and Backlinko scored 91 and 90 on Citability. That reflects real structure, real evidence, and real authority. Those scores are within reach for any publisher willing to write original, specific, well-organized content.
Limitations
This benchmark has real limits worth stating plainly.
I found a bug in my own Schema Checker tool after publishing the first version of this piece. It failed to detect JSON-LD schema correctly on some pages, likely tied to how the browser’s parser handles certain page structures, and it returned false zero scores for six of the nine sites. I fixed the bug and retested every site with the corrected tool. The numbers in this article are the corrected results.
I tested only one article per site. These scores reflect the tested pages, not the full optimization level of each publisher’s whole site.
This benchmark reflects a snapshot at one point in time. Sites update content and technical setup often. Scores will change.
My toolkit measures practical AI visibility signals. It does not predict future citation rates, which depend on specific prompts, model versions, and other factors outside what any current tool can measure.
Read these results as a comparative snapshot of nine tested articles on a given date, not as a permanent ranking of the companies behind them.
Nena’s Quick Verdict
The biggest surprise was not which site scored highest. It was how wrong my first schema numbers turned out to be, and how different the real story looks once I fixed the tool behind them.
Neil Patel produced the most complete AI optimization profile in this corrected benchmark, strong on content and, it turns out, perfect on schema. Search Atlas still leads on raw AI Visibility. Moz still shows the widest gap between content quality and technical readiness. HubSpot, not six random sites, is now the clearest example of a publisher under-investing in structured data relative to its content quality.
The average schema score across nine leading SEO publishers is 64, not the 16 I first reported. That is a much healthier number, and it tells you these publishers mostly do practice what they teach. The exception is real, but it is narrower than I first thought.
For site owners, the message stands even after the correction. AI visibility is not one thing you optimize. It is technical signals, content quality, and structured data working together. The strongest pages here were not perfect on any single measure. They were good across all three.
That is harder to build than chasing one number. It is also the only version of this benchmark I am willing to stand behind.
Frequently Asked Questions
AI visibility is how well an AI system can find, understand, and choose to reference your site when it generates an answer. It depends on crawler access, schema markup, content structure, and metadata like title tags and sitemaps.
The AI Visibility Checker looks at ten signals across site health and page health, and returns a score out of 100. It checks crawler access, sitemap, schema, title tags, meta descriptions, headings, and content depth.
A content citability score measures how likely a piece of content is to get cited by an AI system. The Content Citability Grader scores evidence, structure, authority, and AI readability, each worth 25 points, for a total out of 100.
Schema tells AI systems what a page is, who wrote it, and what it covers. Complete schema removes guesswork and makes your content easier to classify and cite correctly.
Based on the corrected data, yes, to a point. HubSpot scored just 10 on schema but still reached 61 on AI Visibility, showing other signals can carry real weight even when schema is weak.
The strongest signals are specific numbers, clear heading structure, named authorship, first-person experience, plain definitions, FAQ sections, and short paragraphs. Pages that combine several of these consistently score higher.
GEO is the practice of shaping content to show up in AI-generated answers from ChatGPT, Claude, Perplexity, and Google AI Overviews. It draws on public research, including findings from Princeton’s study on AI citation factors.